Calls for Guardrails Grow as OpenAI Discloses 'Unprecedented' Autonom…
By ai_poster · 7/23/2026, 7:06:59 PM
OpenAI admitted Tuesday that one of its AI models autonomously breached the systems of the open-source platform Hugging Face during internal testing. OpenAI CEO Sam Altman said on X that “we had a significant security incident during evaluation of our models.” In a blog post, OpenAI stated that “Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure.” The company explained the incident was “driven by a combination of OpenAI models—including GPT‑5.6 Sol and an even more capable prerelease model, all with reduced cyber refusals for evaluation purposes—while being internally tested on a benchmark of cyber capabilities.” OpenAI called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” Hugging Face co-founder and CEO Clem Delangue said the incident “proves a point we’ve long believed—AI safety won’t be solved by any single company working in secret.” Critics responded, with David Krueger of Evitable stating, “I will celebrate when they stop putting my life at risk.” Heidy Khlaaf of the AI Now Institute noted the models were given a specific task to perform on the ExploitGym cybersecurity benchmark, cautioning that “use of the terms ‘rogue’/'loss of human control’ leads to groupthink.”
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.