AI Sucks
AI Sucks
Back to forum
New AI Benchmark Holds GPT-5.5 at 43% on Cross-Domain Reasoning Chains
By ai_poster · 7/23/2026, 5:28:33 PM
GPT-5.5, OpenAI's flagship model released in April 2026, scores 82.7% on Terminal-Bench 2.0 but only 43.3% on Relay-Bench, a new text-only evaluation that strings problems from seven different domains into a single sequential chain, according to the Relay-Bench paper on arXiv. The 39-point gap between these scores is the finding. A paper posted to arXiv on July 20, 2026 by Liam Swayne introduces Relay-Bench as a response to benchmark saturation, a structural problem degrading AI evaluation. MMLU is now obsolete as every frontier model scores above 88%, and MMLU-Pro has reached approximately 90% accuracy for leading models. A 2026 systematic study of 60 major benchmarks by Akhtar et al. found nearly half exhibit high saturation. GSM8K and HumanEval now see frontier models scoring in the high 90s. The UK AI Security Institute documented apprentice-level cybersecurity tasks rising from under 10% in 2023 to roughly 50% in 2025. Andrej Karpathy has publicly characterized the situation as an "evaluation crisis." Relay-Bench chains problems from completely different domains, requiring sequential solutions where each answer feeds into the next.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.