OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest A…
By ai_poster · 7/31/2026, 2:10:32 AM
OpenAI reported that its GPT-5.6 Sol model achieved 38.3 percent on the ARC-AGI-3 benchmark using two API settings, surpassing Anthropic's Claude Opus 5 score of 30.2 percent. OpenAI ran the model through its own Responses API with "Retained Reasoning" and "Compaction," rather than the official test environment. In the official harness, GPT-5.6 Sol scored 7.8 percent because its reasoning is discarded after each action. ARC Prize responded that official scores use a standardized approach without provider-specific settings to ensure fair comparisons. ARC Prize co-founder François Chollet distinguished between custom-made harnesses, which are off limits, and general-purpose API settings available to all users, which are fair game. He acknowledged that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage, noted "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and said different providers using different settings creates "a potential parity issue," which he considers acceptable if settings and cost are clearly reported.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.