Thinking Machines Unveils Inkling Small: Compact AI Model Delivers Near-Identical Performance at Quarter Size

2 Sources

Share

Thinking Machines releases Inkling Small, a 276-billion-parameter open source AI model that matches its larger predecessor on key benchmarks despite being four times smaller. The multimodal model offers enterprises reduced compute costs while maintaining strong coding and reasoning capabilities.

Thinking Machines Releases Compact Open Source AI Model

Just two weeks after launching its flagship Inkling model, Thinking Machines has introduced Inkling Small, a 276-billion-parameter open source AI model that delivers comparable performance at roughly one-quarter the size of its predecessor

1

. The startup, led by former OpenAI chief technology officer Mira Murati and backed by Nvidia, designed the new model to address enterprise demands for powerful AI capabilities without prohibitive compute costs

2

.

Inkling Small operates as a sparse Mixture-of-Experts model with 12 billion active parameters per token, compared to the original Inkling's 41 billion active parameters and 975 billion total parameters

1

. Released under a permissive Apache 2.0 license, the model accepts text, image, and audio inputs while supporting a 1M-token context window

1

. Thinking Machines has made the full weights available on Hugging Face and enabled fine-tuning through its Tinker model training API

1

.

Matching Performance on Artificial Analysis Intelligence Index

Artificial Analysis assigned Inkling Small a score of 40 on its Intelligence Index, coming within a single point of Inkling's score of 41

1

. This achievement positions the open-weights AI model "well above average" compared to other models of similar size, with no open-weight model at Inkling Small's size or smaller scoring higher on the index

1

2

. At 93 tokens per second, the model also processes faster than average

2

.

The model surpasses its larger sibling on several critical benchmarks. On SWE-bench Verified, a key evaluation for coding assistants, Inkling Small achieved 80.2% compared to Inkling's 77.6%

1

. It also outperformed on Terminal Bench 2.1 with 64.7% versus 63.8%, and scored above 31% on Humanity's Last Exam, exceeding Inkling's 29.7%

1

2

. Additional gains appeared on SciCode, GPQA Diamond, and CritPt benchmarks

1

.

Enterprise Use Cases and Cost Advantages

Thinking Machines positions Inkling Small for enterprise use cases including coding assistants, agentic tasks, retrieval-augmented generation, document analysis, and multimodal workflows

1

2

. The company emphasizes that developers sacrifice relatively little capability while substantially reducing compute requirements, inference costs, and deployment footprint

1

.

At launch, Thinking Machines offers a limited-time 50% discount on API pricing. The standard 64K-context version costs $0.58 per million prefill tokens, $1.44 per million sampled tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens

1

. A 256K-context variant is available at higher rates

1

. The model remains far too large for consumer devices, requiring at least 600 GB of aggregate GPU memory, with supported configurations including 4x NVIDIA B300 GPUs

1

.

Sparse Architecture Enables Efficiency Gains

Inkling Small achieves its efficiency through a sparse Mixture-of-Experts architecture. The model's 42-layer decoder routes each token to six of 256 specialized experts, plus two shared experts that remain active for every token

1

. This design explains the distinction between 276 billion total parameters and 12 billion active parameters—the system maintains a large pool of learned capacity but activates only a fraction during each inference step

1

.

The model processes images, audio, and text natively as multimodal inputs projected into a shared representation, rather than handling them through separate external systems

1

. Thinking Machines also supports variable reasoning effort, allowing developers to adjust test-time compute based on task difficulty, providing direct control over quality, latency, and cost tradeoffs

1

.

Performance Tradeoffs and Future Implications

While Inkling Small excels at reasoning and coding, it shows weaker factual knowledge coverage. The model scored 15.5% on τ³-Banking compared to Inkling's 23.7%, and its AA Omniscience score is negative, reflecting reduced factual coverage despite a slightly lower hallucination rate

1

. Organizations deploying it for high-stakes factual tasks will need to implement retrieval, verification, and human review processes

1

.

Source: VentureBeat

Source: VentureBeat

Thinking Machines stated that "Inkling-Small was made in pursuit of our mission to build AI that extends human will and judgment," noting that fine-tuned models can outperform closed models on specific tasks while operating faster and cheaper

2

. The company's partnership with Nvidia provided access to GB300 NVL72 systems for training, alongside significant investment from the chipmaker

2

. Watch for enterprises with limited GPU infrastructure to adopt Inkling Small for RAG systems and document analysis, while monitoring whether the efficiency gains accelerate broader open-source AI adoption across industries seeking to reduce dependency on closed models.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved