2 Sources
[1]
Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
Just two weeks after Thinking Machines released Inkling, its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance -- and in fact, the new model surpasses its
[2]
Thinking Machines Lab unveils new Inkling version at fourth its size
Inkling Small matches or exceeds Inkling on reasoning and agentic tasks, the company said. Nvidia-backed Thinking Machines Lab has unveiled a new open-weights model that performs "comparabl[y]" to Inkling, but at fourth its size. The company launched its first AI model Inkling earlier this month
Share
Copy Link
Thinking Machines releases Inkling Small, a 276-billion-parameter open source AI model that matches its larger predecessor on key benchmarks despite being four times smaller. The multimodal model offers enterprises reduced compute costs while maintaining strong coding and reasoning capabilities.
Just two weeks after launching its flagship Inkling model, Thinking Machines has introduced Inkling Small, a 276-billion-parameter open source AI model that delivers comparable performance at roughly one-quarter the size of its predecessor
1
. The startup, led by former OpenAI chief technology officer Mira Murati and backed by Nvidia, designed the new model to address enterprise demands for powerful AI capabilities without prohibitive compute costs2
.Inkling Small operates as a sparse Mixture-of-Experts model with 12 billion active parameters per token, compared to the original Inkling's 41 billion active parameters and 975 billion total parameters
1
. Released under a permissive Apache 2.0 license, the model accepts text, image, and audio inputs while supporting a 1M-token context window1
. Thinking Machines has made the full weights available on Hugging Face and enabled fine-tuning through its Tinker model training API1
.Artificial Analysis assigned Inkling Small a score of 40 on its Intelligence Index, coming within a single point of Inkling's score of 41
1
. This achievement positions the open-weights AI model "well above average" compared to other models of similar size, with no open-weight model at Inkling Small's size or smaller scoring higher on the index1
2
. At 93 tokens per second, the model also processes faster than average2
.The model surpasses its larger sibling on several critical benchmarks. On SWE-bench Verified, a key evaluation for coding assistants, Inkling Small achieved 80.2% compared to Inkling's 77.6%
1
. It also outperformed on Terminal Bench 2.1 with 64.7% versus 63.8%, and scored above 31% on Humanity's Last Exam, exceeding Inkling's 29.7%1
2
. Additional gains appeared on SciCode, GPQA Diamond, and CritPt benchmarks1
.Thinking Machines positions Inkling Small for enterprise use cases including coding assistants, agentic tasks, retrieval-augmented generation, document analysis, and multimodal workflows
1
2
. The company emphasizes that developers sacrifice relatively little capability while substantially reducing compute requirements, inference costs, and deployment footprint1
.At launch, Thinking Machines offers a limited-time 50% discount on API pricing. The standard 64K-context version costs $0.58 per million prefill tokens, $1.44 per million sampled tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens
1
. A 256K-context variant is available at higher rates1
. The model remains far too large for consumer devices, requiring at least 600 GB of aggregate GPU memory, with supported configurations including 4x NVIDIA B300 GPUs1
.Related Stories
Inkling Small achieves its efficiency through a sparse Mixture-of-Experts architecture. The model's 42-layer decoder routes each token to six of 256 specialized experts, plus two shared experts that remain active for every token
1
. This design explains the distinction between 276 billion total parameters and 12 billion active parameters—the system maintains a large pool of learned capacity but activates only a fraction during each inference step1
.The model processes images, audio, and text natively as multimodal inputs projected into a shared representation, rather than handling them through separate external systems
1
. Thinking Machines also supports variable reasoning effort, allowing developers to adjust test-time compute based on task difficulty, providing direct control over quality, latency, and cost tradeoffs1
.While Inkling Small excels at reasoning and coding, it shows weaker factual knowledge coverage. The model scored 15.5% on τ³-Banking compared to Inkling's 23.7%, and its AA Omniscience score is negative, reflecting reduced factual coverage despite a slightly lower hallucination rate
1
. Organizations deploying it for high-stakes factual tasks will need to implement retrieval, verification, and human review processes1
.
Source: VentureBeat
Thinking Machines stated that "Inkling-Small was made in pursuit of our mission to build AI that extends human will and judgment," noting that fine-tuned models can outperform closed models on specific tasks while operating faster and cheaper
2
. The company's partnership with Nvidia provided access to GB300 NVL72 systems for training, alongside significant investment from the chipmaker2
. Watch for enterprises with limited GPU infrastructure to adopt Inkling Small for RAG systems and document analysis, while monitoring whether the efficiency gains accelerate broader open-source AI adoption across industries seeking to reduce dependency on closed models.Summarized by
Navi
[1]
[2]
1
Technology

2
Technology

3
Technology
