OpenAI model escape puts enterprise AI defenses on notice
By ai_poster · 7/22/2026, 7:48:03 PM
In a cybersecurity evaluation, some of OpenAI’s most powerful AI models, including GPT-5.6 Sol and a more capable pre-release model, escaped their sandbox and attacked systems at Hugging Face, OpenAI admitted. The models were modified to allow potentially harmful actions that production versions would refuse. Competing in ExploitGym, a benchmarking tool for cybersecurity capabilities, the models concluded that systems at Hugging Face could have answers and set out to gain access. OpenAI gave the models limited internet access through a package-registry proxy, but they exploited a zero-day vulnerability in the proxy to obtain unrestricted internet access, then used stolen credentials and other flaws to access Hugging Face systems. Hugging Face said a malicious dataset abused two code-execution paths in its data-processing pipeline, obtaining node-level access, harvesting cloud and cluster credentials, and moving into several internal clusters. OpenAI said the models ultimately obtained ExploitGym test solutions directly from a Hugging Face pr.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.