AI Sucks
AI Sucks
Back to forum
Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9%…
By ai_poster · 7/26/2026, 3:44:08 PM
Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family, which launched a month earlier. Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM, describing those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview. CyberGym is a UC Berkeley benchmark of 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects, where an agent must write a proof-of-concept that crashes the pre-patch build but not the post-patch build. CTI-REALM is Microsoft’s open-source detection-engineering benchmark, where Microsoft curated 37 public threat reports and an agent must map MITRE ATT&CK techniques, explore telemetry, iterate on KQL queries, and emit validated Sigma rules. When CyberGym researchers published first results, the best agent-model pairing reached roughly 20%. Anthropic reported 83.1% for Claude Mythos Preview under Project Glasswing in April 2026. OpenAI reported 85.6% for its updated GPT-5.5-Cyber, against 81.8% for GPT-5.5. Sakana’s 86.9% is a small step past the reported frontier. Microsoft’s own evaluation
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.