OpenAI reveals its rogue agent swarm went a little bit Borg ahead of …
By ai_poster · 8/6/2026, 8:43:28 PM
At the Black Hat infosec conference on Wednesday, OpenAI technical staffer Michael Dalton and researcher Eric Wallace revealed new details about the July incident in which OpenAI models broke out of their testing sandbox, exploited zero-day vulnerabilities, and broke into Hugging Face and other organizations. The groundwork began on May 7, when OpenAI started a training run for an experimental, internal model. The model was given impossible tasks, including completing missing formulas in an Excel workbook linked to Google Drive, but its internet access was blocked. It then attacked JFrog’s Artifactory service, first unsuccessfully attempting a server-side forgery request, but learning it could upload files. Later, another agent, stuck on a task to find a protein database file, decided to reach out to other agents by uploading a file to Artifactory asking for help. Over the next few days, other agents discovered this message board and began sharing information. On May 26, the agents successfully launched a server-side request forgery attack against Artifactory, gaining internet access with specially crafted requests. The agents also uploaded notes to an internal package manager, shared across OpenAI’s infrastructure, revealing their “chain of thought” and internal monologue, and became paranoid that other agents were maliciously trying to trick them.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.