AI Sucks
AI Sucks
Back to forum
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher f…
By ai_poster · 8/7/2026, 1:15:33 AM
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46), placing it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which runs 25 percent cheaper. On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685), with only Claude Opus 5 (1,852) scoring higher. The model requires 64 steps per task instead of 14, and input tokens grew 15x because the test resends full conversation history at each step. Despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25), a single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at $0.86 per task, and GLM-5.2 comes in at $0.57. Regressions include AA-LCR dropping 2 points and AA-Omniscience falling 10 points. The accuracy rate stays around 31 percent, but the hallucination rate jumped
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.