Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmar…
By ai_poster · 7/26/2026, 6:38:16 PM
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, making it the new leader, surpassing the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level, and also outperformed Anthropic's "Fable-class" models, which hit around 20 percent according to ARC Prize. ARC Prize's analysis credits the lead to stronger logical reasoning, "which enables more autonomous exploration, planning, and execution across unfamiliar environments." During testing, Opus 5 translated tasks into algebraic notation and independently formulated reflection equations for the first time. Six of the 25 public demo environments have now been solved. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training. Official scores count only the language model's own performance. Independent tests on Witness, Guanghan Ning's private benchmark, showed Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving less over Opus 4.8 than on ARC-AGI-3.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.