AMD Acquires Taalas to Etch AI Models Into Silicon for Faster Inference

Reviewed byNidhi Govil

11 Sources

Share

AMD acquired Toronto-based Taalas, an AI chip startup that hardwires model weights directly into silicon. The deal strengthens AMD's AI inference capabilities with technology that promises 48x faster performance than traditional GPUs, though chips remain locked to specific models.

AMD Acquires Taalas to Challenge Nvidia in AI Inference

AMD acquires Taalas, a Toronto-based startup that etches AI models into silicon, marking the chip giant's latest move to challenge Nvidia's dominance in AI hardware

1

2

. The deal, announced Thursday after market close, focuses on making high-performance AI inference services faster and cheaper to run. While AMD didn't disclose financial terms, the acquisition brings in a company that raised $219 million since its 2023 founding

3

5

. The move mirrors Nvidia's $20 billion Groq licensing deal from December, signaling how seriously chipmakers now take specialized inference technology

1

. Vamsi Boppana, AMD's senior vice president of AI, stated that "AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload"

2

5

.

Source: Benzinga

Source: Benzinga

How Model-Specific Integrated Circuits Work

Taalas developed what it calls model-specific integrated circuits that hardwire AI models into silicon rather than loading weights from memory

1

4

. The chips eliminate compute and memory bottlenecks by merging memory and compute functions directly into the silicon

2

4

. According to Taalas co-founder Ljubisa Bajic, this approach eliminates the need for high-bandwidth memory, 3D stacking, or liquid cooling

4

. The processors contain two main regions: a mask-ROM recall fabric where model weights are etched, and an SRAM recall fabric for KV caches and fine-tuning adapters

1

. Taalas manufactured its first test chip, the HC1, on TSMC's 6nm process technology in February

1

.

Source: The Register

Source: The Register

Breakthrough Inference Performance Numbers

The HC1 chip delivered remarkable inference performance when running Meta Llama 3.1 8B, achieving 16,960 tokens per second—48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at the time of announcement

1

3

. Taalas claims its technology can produce output for specific models thousands of times faster than traditional GPUs, though this comes at the cost of flexibility

3

. The company's second-generation HC2 chip, due this summer, aims to support 20 billion parameters per chip

1

. For larger models, weights can be distributed across multiple accelerators using pipeline parallelism—meaning just 50 accelerators would support a trillion-parameter model

1

.

The Trade-Off: Speed Versus Flexibility

The technology's primary limitation is that once deployed, chips remain locked to a single model

3

4

. Any significant change beyond LoRA adapters requires a silicon re-spin, which adds cost and time

1

. However, Taalas mitigates this constraint by requiring changes to only two metal layers rather than starting from scratch, allowing new models to be etched in about two months instead of six

1

4

. Taalas notes that fewer than a handful of a chip's hundred-plus layers need alteration from one design to another

5

. The company claims etching model weights into silicon costs 100x less than training a frontier model

1

. This positions the technology for AI model developers, infrastructure providers, and inference services rather than general enterprise customers

1

.

Source: Wccftech

Source: Wccftech

Integration Into AMD's AI Hardware Portfolio

AMD plans to integrate Taalas technology into its AI accelerator roadmap and develop system-level solutions using AMD Instinct GPUs, EPYC processors, and Helios rack-scale platform

2

5

. The company appears positioned to pair Instinct-based Helios racks with Taalas chips in a disaggregated architecture where compute-heavy prompt processing runs on GPUs while token generation offloads to Taalas accelerators

1

. AMD could adopt a deployment strategy where customers initially validate models on Instinct accelerators before transitioning to Taalas accelerators for production

1

. This approach offers far better space and power efficiency than Nvidia's LPX systems, which would need dozens of GPUs and at least 2,000 Groq LPUs to serve the same trillion-parameter model

1

.

Strategic Context in AMD's Acquisition Spree

The Taalas deal continues AMD's aggressive expansion of its AI capabilities through acquisitions

2

3

. AMD bought MK1, an AI software startup specializing in high-speed inference, in November

2

. The company purchased MEXT in June and added FastFlowLM to its Artificial Intelligence Group in July

2

. In 2024, AMD paid $665 million for Silo AI and acquired ZT Systems for $4.9 billion, which provided the technical foundation for its rack-scale products

3

5

. AMD reported record revenue of $11.5 billion in its most recent quarter, with its data center segment more than doubling year-over-year to $6.7 billion

5

. The company has secured Helios commitments from major customers including Meta, Microsoft, OpenAI, and Anthropic

1

5

. AMD stock rose approximately 1.5% following the announcement

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved