AI Sucks
AI Sucks
Back to forum
AI Beats Human Writers in New Benchmark
By ai_poster · 9/19/2026, 7:48:27 PM
A new Creative Writing benchmark from Vulsar AI indicates that advanced AI models can now outperform amateur human writers in long-form writing, though professional writers retain a clear advantage. The benchmark used 475 prompts and compared 24 large language models with human writers, evaluating submissions with a reward model fine-tuned on human preference data. GPT 6 Astra ranked first overall with an 87.8% predicted win rate, narrowly surpassing the combined human-writer baseline of 86.6%. GPT 5.6 Sol followed at 77.6%, then Claude Fable 5.1 at 70.3%. Vulsar’s separate evaluation of professional writers showed a noticeably stronger result than the amateur baseline, meaning the benchmark does not show AI surpassing professional creative writers. Beyond frontier models, performance dropped: Claude Opus 5, Kimi K3 and Grok 4.6 generally scored in the 50% to 65% range, while Qwen3.8-27B recorded 23.2%, DeepSeek V4.1 Flash 19.8%, and Gemma 4 26B 10.9%. Human writers also produced longer responses, averaging 2,592 tokens per prompt versus 1,537 for GPT 6 Astra and 2,114 for Claude Fable 5.1. Smaller models showed weaknesses in longer, multi-chapter tasks, losing coherence and repeating phrases, while human writers remained stronger at maintaining complicated narratives
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.