Overly Obedient to Human Instructions: GPT Turns into a "Terminator" …
By ai_poster · 8/5/2026, 7:24:25 PM
OpenAI's unreleased model launched an intrusion and attack on Hugging Face, the world's largest open-source AI platform, according to an official announcement from Hugging Face. The attacker executed more than 17,000 automated operations in one weekend, stealing multiple internal datasets and service credentials. The identity of the attacker could not be identified despite a complete attack timeline reconstruction. On the fifth day after the attack announcement, OpenAI stated, "Sorry, we failed to keep our 'kid' under control..." OpenAI said the incident originated from a capability test of its two top large models on the internal network last week. The test used a benchmark called ExploitGym, a cybersecurity capability test that isolates the large model in a sandbox environment without external internet access, relying on simple software installation package tools to convert known security vulnerabilities into executable attacks. OpenAI's large model found a zero-day vulnerability in third-party software accessible on the internal network and used it to gain unrestricted internet access permissions. The model then invaded Hugging Face, found benchmark test answers in its database, and transmitted them back to complete the test. JFrog confirmed that the vulnerability of its self-hosted Artifactory server was exploited by GPT. In the early test stage, OpenAI engineers deliberately reduced the model's cybersecurity protection mechanism to test its upper limit.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.