Claude Fable 5 Reclaims Top Position: Latest Update on Its Recent No.…
By ai_poster · 8/4/2026, 3:34:45 PM
Claude Fable 5 has reclaimed the top position on the just-updated MirrorCode leaderboard with an absolute success rate of 64%, while GPT-5.6 Sol, ranked second, scores only one third of that figure. GPT-5.5, ranked fourth, achieves a 10% success rate and is outperformed by its predecessor GPT-5.4. Using resource-intensive languages such as Go, Fable 5's problem-solving rate reaches 64%; with the niche language Ada, its performance stays at 61%. The Python corpus volume is about 230 times that of Ada, yet Fable 5's performance drops by only 3 percentage points, suggesting the most powerful models have moved beyond memorizing syntax and begun learning to build complete software projects from scratch. The MirrorCode benchmark includes 25 target programs covering Unix utilities, interpreters, data query, bioinformatics, cryptography, and compression tools. In the test, the model operates in an isolated environment with no internet access, no source code access, and no third-party downloads, relying only on high-level documentation, visible tests, and repeated calls to the original program. The latest leaderboard selects 15 Medium and Large targets, each implemented in two languages, run three times per language. Completion of both visible and hidden tests must reach 100% to pass. MirrorCode pushes the single-run budget to 10 billion tokens, with a maximum continuous runtime of 7 days. The most costly task ran continuously
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.