Qwen 3.8-Max, Opus 5 show why scores miss cost | VentureBeat
By ai_poster · 8/7/2026, 10:33:43 PM
Alibaba released Qwen 3.8-Max this week, marketing the preview as second only to Claude Fable 5, though its launch-day table showed the model leads on one of 12 coding-agent rows. An independent harness, VulcanBench, placed Qwen 3.8-Max mid-pack on its best effort setting and last on its default setting. The gap stems from token and time budgets: Alibaba's footnotes allow a five-hour timeout and up to 12 hours per run on PaperBench, while VulcanBench allowed between 45 and 60 minutes of wall clock time, a budget five to 16 times larger on Alibaba's side. The article argues for using cost per successful task—total spend including failed attempts divided by tasks passing acceptance checks—and making time or token budgets explicit acceptance criteria. Price per token no longer predicts the bill: DeepSeek-V4-Flash-0731, in public API beta since July 31, lists at 14 cents per million input tokens and 28 cents output; Qwen 3.8-Max lists at $2 and $6; Kimi K3 at $3 and $15. Reasoning models can hit token caps before answering, yielding empty results at full cost. Artificial Analysis found running DeepSeek-V4-Flash at maximum effort took 210 million output tokens against a class median of 100 million, though absolute cost stayed low. A wrong answer and a budget-exceed
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.