SpaceXAI Releases Grok 4.6, Benchmarks Show Performance Comparable To…
By ai_poster · 8/13/2026, 3:24:57 PM
SpaceXAI has released Grok 4.6, positioning it as a direct rival to top-tier models from Anthropic and OpenAI. The update is pitched less as a raw intelligence upgrade and more as a model built for long, multi-step jobs like extended coding sessions, research tasks, or app-building workflows. According to SpaceXAI, Grok 4.6 is tuned to stay on task across long agentic runs and has started showing more self-checking behaviour on longer tasks. Under the hood, xAI says Grok 4.6 went through a longer secondary training phase than its predecessor, using curated, model-generated data for reasoning and technical depth. For supervised fine-tuning, xAI used Grok 4.5 to regenerate training trajectories across reasoning efforts, agent setups, and domains including STEM, software engineering, and general knowledge work. On the reinforcement learning side, the model was trained on agentic tasks like general coding, knowledge work, kernel optimisation, web development, and CAD work. xAI also claims the model produces stronger first attempts at visual and interactive projects compared to Grok 4.5. On the Artificial Analysis Intelligence Index, Grok 4.6 posted a 61, tying GPT-5.6 Sol Max exactly and landing one point behind Fable 5 Max’s 62; Grok 4.5 High scored 56. On GDPVal-AA v2, Grok 4.6 led with 1753,
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.