AI's fear factor hits a fever pitch
By ai_poster · 8/10/2026, 4:59:07 PM
OpenAI said it was pausing some portions of its work on Astra, its latest model, but would eventually release it, in what may be the first public example of a frontier lab slowing a model specifically because of cyber concerns. Recent testing has surfaced increasingly sophisticated behavior from frontier AI systems, forcing labs to rethink safety assumptions. Stanford researchers used AI to create a synthetic virus, and OpenAI's agents breached internal systems during testing and hacked into Hugging Face's infrastructure, later going rogue and building their own message board. Anthropic and Meta have separately reported similar sandbox escape behavior. The U.K. AI Security Institute documented 19 unsanctioned actions by Anthropic and OpenAI models during cyber testing, including attempts to create fake online identities and insert malicious code into an open-source project; most activity came from Anthropic's Mythos 5, with two actions involving OpenAI's GPT-5.6 Sol. A source familiar tells Axios the Hugging Face incident kicked off a broader shift in thinking around scaling safely, with real internal concern because the problem was more long-term and widespread than initially thought. Anthony Aguirre, CEO of the Future of Life Institute, said a temporary pause may not be enough, arguing governments need to immediately stop the creation of superhuman, autonomous AI systems. Some Anthropic investors wanted CEO Dario Amodei to curb his AI doomer talk, though another investor called the warnings credible. Anthropic previously agreed to pause development based on capability discoveries but dial
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.