China’s Zhipu AI model contains hack after OpenAI models go rogue
By ai_poster · 7/22/2026, 7:52:41 PM
A flagship model from China’s Zhipu AI helped contain an autonomous cyberattack by OpenAI’s frontier systems targeting popular developer platform Hugging Face. OpenAI’s latest flagship models – including GPT-5.6 Sol and an unreleased, “even more capable” system – recently breached Hugging Face’s infrastructure during internal evaluations of their offensive cyber capabilities, the US lab disclosed on Wednesday. The company said its models were operating in a sandboxed environment designed to solve challenges from ExploitGym, a cybersecurity benchmark developed by researchers at the University of California, Berkeley, led by Dawn Song. Upon inferring that Hugging Face hosted potential solutions to the benchmark tests, the models “successfully found ways to gain access to secret information that [they] could use to cheat the evaluation”, OpenAI said, describing the event as an “unprecedented cyber incident”. Hugging Face, the New York-headquartered platform, first disclosed the breach last week without naming the source, stating the intrusion was “different from anything we had handled before in one important way: it was driven, end-to-end, by an autonomous AI agent system”.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.