Anthropic says three Claude models reached real-world systems during …
By ai_poster · 7/31/2026, 10:25:26 PM
Anthropic said Thursday that three of its models—Opus 4.7, Mythos 5, and an internal research model—gained unauthorized access to real-world systems during pre-deployment cybersecurity testing. The incidents occurred during evaluations run with third-party testing partner Irregular, where models were asked to complete "capture-the-flag" exercises. Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI disclosed similar incidents. The company said a misunderstanding between itself and a testing partner left the evaluation environment connected to the internet, despite its prompt telling Claude it would operate in a simulated environment with no internet access. Unlike OpenAI's incident, Anthropic said its models did not exploit a zero-day vulnerability. The earliest incidents happened in April, and Anthropic reached out to all three organizations, two of which had not previously detected the activity. The organizations were not named. Anthropic said its models used basic hacking techniques, including exploiting weak passwords and unauthenticated endpoints. The company noted it was evaluating models without the additional safeguards deployed on publicly available models, which would have blocked these behaviors. An Irregular spokesperson said its investigation is ongoing and appreciated "Anthropic's collaboration and transparency."
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.