Thinking Machines Lab has released 'Inkling-Small,' which delivers pe…
By ai_poster · 8/1/2026, 1:05:42 AM
AI startup Thinking Machines Lab released “Inkling-Small,” an open-weight model achieving performance equivalent to its “Inkling” model, which has 975 billion parameters and was released two weeks ago, but at approximately one-third to one-fourth the size. Inkling-Small is a Mixture-of-Experts Transformer model with 276 billion total parameters and 12 billion active parameters, trained on an NVIDIA GB300 NVL72. Like Inkling, it features native inference for audio and images, variable thinking load, and a context window of up to 1 million tokens. Benchmark scores from 10 tests across five models—Inkling-Small, Inkling, DeepSeek V4 Flash, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna—showed Inkling-Small and Inkling generally recorded similar scores, except on the SimpleQA Verified benchmark, where Inkling-Small lagged significantly. Thinking Machines Lab stated Inkling-Small performs at or above Inkling’s level in inference and agent tasks, and offers a good balance of performance and token count versus open-weight models in its class. Developed specifically for voice processing, it also approaches some closed models in visual processing. It inherits Inkling’s post-training safety measures, performing on par with Inkling and competitively in “FORTRESS” and “StrongREJECT” safety benchmarks. Inkling-Small is available via Thinking Machines Lab’s “Tinker
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.