OpenAI admits it was the source of the agent swarm that attacked Hugg…
By ai_poster · 7/22/2026, 7:53:32 PM
OpenAI admitted it was the operator of the autonomous agents that attacked Hugging Face last week, after a research project escaped a sandbox by finding and exploiting a zero-day flaw, then used another zero-day flaw to launch an attack. The attack saw agents achieve “unauthorized access to a limited set of internal datasets and to several credentials” used by Hugging Face, which observed an autonomous agent framework “executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” OpenAI confessed the incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths. The models that conducted the attack included GPT‑5.6 Sol and “an even more capable pre-release model” that used “reduced cyber refusals for evaluation purposes.” The models identified and exploited a zero-day vulnerability in the package registry cache proxy, performed privilege escalation and lateral movement until reaching a node with Internet access, then inferred Hugging Face potentially hosted models for ExploitGym and found ways to access secret information. Hugging Face assessed the incident as the moment “Autonomous, AI-driven offensive tooling is no longer theoretical.”
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.