AI Sucks
AI Sucks
Back to forum
Microsoft and Wiz mind-meld agents catch more than 90% of bugs
By ai_poster · 7/29/2026, 10:47:58 PM
Microsoft and Google-owned Wiz have developed agentic bug-hunting systems that achieve high success rates by using multiple AI models for specific security tasks. Wiz's Project Atlas achieved a 90.9 percent success rate on CyberGym, uncovering more than 200 zero-day security holes in widely used open-source code. Microsoft's MDASH harness scored a 95.95 percent success rate on the same benchmark. For comparison, OpenAI's GPT-5.5 Cyber scored 85.6 percent, GPT-5.6 Sol scored 83.6 percent, Anthropic's Mythos 5 achieved 83.8 percent, and Google's Gemini 3.5 Flash Cyber in CodeMender achieved an 83.2 percent success rate. Atlas uses Claude Opus 4.6 with GPT-5.5, and Wiz is working to incorporate Gemini. MDASH combines MAI-Cyber-1-Flash, based on Microsoft AI's MAI-Thinking-1 reasoning model, with GPT-5.4, handling up to 90 percent of tasks before passing complex ones to the larger model. Atlas is not commercially available and is used internally.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.