3 Sources
[1]
Pathway says architecture can matter as much as model scale
* A 150M model reached 29.5% while costing just $0.0007 per task * ChatGPT scored higher, yet its comparable reasoning runs cost substantially more * BDH-CQ performs reasoning internally instead of generating lengthy intermediate text Pathway, an AI lab focused on building Post-Transformer
[2]
Did we build the engine before we worked out the physics? What's the pathway to take after the LLM transformer?
The AI boom was built on top of the transformer model first discovered in 2017 for better translation. The stock valuations, planned multi-gigawatt data centers, and the whole AI economy now exist because transformers unlocked the ability to automatically train AI models on human language at scale.
[3]
Pathway Claims Major AI Cost-Efficiency Breakthrough with New 150M-Parameter Model
The cost gap comes from a structural difference. Many transformer-based reasoning systems externalize their work as a chain-of-thought, generating extra tokens sequentially and feeding them back into later steps. Pathway, an AI lab building a Post-Transformer architecture and models, today
Share
Copy Link
Pathway released BDH-CQ, a 150-million-parameter model scoring 29.5% on ARC-AGI-1 at just $0.0007 per task—11 times cheaper than GPT 5.6 Luna. The breakthrough stems from Post-Transformer architecture that reasons internally rather than generating token-heavy intermediate text, challenging the assumption that intelligence requires scale.
Pathway, an AI lab building Post-Transformer architecture, has released benchmark results that challenge fundamental assumptions about AI cost-efficiency and model design. Their BDH-CQ reasoning model, with just 150-million-parameter, scored 29.5% pass@2 on the public ARC-AGI-1 benchmark at a computed inference cost of $0.0007 per task
1
. This performance costs approximately 11 times less than OpenAI's GPT 5.6 Luna (Low) model, which scores only marginally higher at 34.2% while costing $0.008 per task—even after OpenAI's 80% price cut implemented on July 30th3
.The cost gap widens dramatically when compared to frontier models. Claude Opus 5 and Gemini 3.1 Pro achieve 97-98% accuracy but cost around $0.5-$0.6 per task, meaning top-tier performance costs nearly a thousand times more than BDH-CQ for the highest scores
1
. Even budget alternatives like Qwen3 235B cost over three times more than BDH-CQ while delivering inferior performance, eliminating any real competition on either price or capability metrics."Today's AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence," said Zuzanna Stamirowska, CEO and co-founder of Pathway
3
. The efficiency breakthrough stems from a fundamental structural difference in how BDH-CQ performs reasoning during inference computations compared to transformer models.Most reasoning AI systems today generate intermediate text through chain-of-thought processes, adding tokens sequentially before producing final answers. As these written reasoning traces grow longer, both inference costs and latency increase substantially
1
. BDH-CQ takes a radically different approach—it performs reasoning in latent space, solving problems internally within its recurrent state rather than externalizing work as visible text3
.Łukasz Kaiser, co-author of the original 2017 Transformer paper, validated the results and stated: "Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning"
1
. The findings were independently replicated by Kaiser himself and Richard Zhong, an NYU researcher focused on model evaluation and benchmark robustness3
.Stamirowska's background in complex systems rather than deep learning informed Pathway's approach to AI for unstructured data. "With brute force scaling of the Transformer, the AI community built steam engines for thought before developing the thermodynamics of intelligence," she explained
2
.
Source: diginomica
Her doctoral work on network evolution and global trade forecasting enabled her to address missing theoretical foundations in AI by rethinking architecture from first principles.
The Baby Dragon Hatchling model (BDH) draws inspiration from biological intelligence—local interactions, sparse activations, and the brain as a physical network system. Pathway sought to identify a small set of laws, similar to statistical mechanics, that derive how systems reason from how their parts interact
2
. This approach aims to make safety a design property that can be analyzed and engineered before deployment, rather than assessed only after systems are released.Early experiments confirm that standard Transformer-like scaling laws apply during pretraining at scales from 1B to 600B parameters, while preserving the latent reasoning capabilities specific to BDH-CQ
3
. This suggests the efficiency advantages could extend across larger model sizes without sacrificing the architectural benefits.Related Stories
Amazon Web Services recognized BDH-CQ's potential for production deployments. "Customers are increasingly exploring how to move advanced reasoning from experimentation into production, where performance, efficiency, and scalability all matter," said Nicolas Tarducci of AWS
1
. The cost advantages become critical when deploying AI at scale across real-world applications.Pathway plans to extend this approach toward harder benchmarks, including mathematical reasoning, ARC-AGI-2, and eventually full ARC-AGI-3 evaluations
1
. When these efficiency and state-tracking capabilities extend to those domains, BDH will support applications requiring reliable reasoning as information and constraints change—from cybersecurity incident response to real-time industrial operations3
.The limitations of transformer models extend beyond cost. Current systems require extensive scaffolding through context engineering and RAG to remember previous conversations, corrections, or quarterly data. This surrounding stack must reconstruct context on every call, treating AI as a tool requiring periodic updates rather than a system whose useful experience compounds over time
2
.Stamirowska frames the alternative as the difference between consuming context and accumulating experience: "If learning can happen safely during inference, the model can start to accumulate experience, not merely consume context"
2
. This shift could reduce work spent on complicated context management and instead bake organizational knowledge natively into models.Pathway's recent $500 million valuation reflects investor confidence in this Post-Transformer vision
2
. If efficiency gains hold across larger and more difficult tasks, cost rather than raw capability could increasingly separate rival reasoning systems, reshaping competitive dynamics in enterprise AI deployment.Summarized by
Navi
[2]
28 Jan 2025•Technology

10 Jul 2026•Business and Economy

21 Sept 2026•Technology

1
Policy and Regulation

2
Technology

3
Technology
