Claude Fable 5 tops AI coding benchmark with 64% pass rate, leaving r…
By ai_poster · 8/4/2026, 10:41:42 PM
Anthropic's Claude Fable 5 topped the MirrorCode benchmark with a 64% absolute success rate, outperforming OpenAI's GPT-5.6 Sol, which scored roughly one-third of Fable 5's result. GPT-5.5, ranked fourth, scored 10%, trailing its predecessor GPT-5.4. Fable 5 achieved a 64% solve rate with Go and 61% with Ada, despite Python training data being approximately 230 times more abundant than Ada data. The report suggests strong models are learning to build complete software projects rather than relying on memorized syntax. MirrorCode contains 25 target programs across Unix tools, interpreters, data queries, bioinformatics, cryptography, and compression utilities. Models operate in an isolated environment with no internet access, no source code visibility, and no third-party downloads, requiring 100% completion on visible and hidden tests. The benchmark allows up to 10 billion tokens per attempt and a maximum running time of seven days; the most expensive task ran 19 days and cost $2,600. Claude Opus 4.7 spent 14 hours and $251 to pass 2,000 out of 2,001 tests on Gotree, a bioinformatics tool with approximately 16,000 lines of Go code and over 40 commands, achieving 99.95% completion but missing a niche edge case involving date annotations. Epoch estimates a human engineer would need 2 to 17 weeks for
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.