AMD acquires Taalas to etch AI models into silicon, achieving 17,000 tokens per second

Reviewed byNidhi Govil

6 Sources

Share

AMD acquired Toronto-based startup Taalas to boost AI inference performance by hardwiring models directly into silicon. Taalas' HC1 chip serves Llama 3.1 at 17,000 tokens per second, 48x faster than Nvidia GPUs. The technology will integrate with AMD Instinct GPUs and Helios racks to compete in the growing AI inference market.

AMD acquires Taalas to challenge Nvidia in AI inference

AMD acquires Taalas, a Toronto-based startup that etches AI models directly into silicon, marking the chipmaker's latest move to compete with Nvidia in the rapidly expanding AI inference market

1

2

. The deal, announced at market close on Thursday, positions AMD to deliver what it calls "differentiated inference performance and efficiency" as companies race to cut computing costs for real-time AI deployment

2

. While AMD didn't disclose financial terms, Taalas had raised $219 million in total funding since its 2023 founding, including a $169 million round in February

2

4

. The acquisition represents AMD's third AI chip deal in nine months, following MK1 in November and Mext in June

4

.

Source: The Register

Source: The Register

How hardwiring AI models into silicon delivers breakthrough speed

Taalas builds model-specific integrated circuits that permanently etch AI model weights directly into silicon transistors, eliminating the need for high-bandwidth memory and solving the memory wall problem where GPUs sit idle waiting for data

1

5

. The startup's HC1 chip, fabricated on TSMC's 6-nanometer process, achieved 16,960 tokens per second when serving Meta's Llama 3.1 8B model—48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at one-tenth the power

1

4

. The technology uses a proprietary digital architecture where each transistor stores 4 bits of data and acts as a physical router rather than performing calculations, with shared hardware blocks pre-computing all possible mathematical outcomes

5

. Taalas chips comprise two main regions: a mask-ROM recall fabric where model weights are permanently etched, and an SRAM recall fabric for KV caches and fine-tuning adapters

1

.

Source: Wccftech

Source: Wccftech

Integration with AMD Instinct GPUs and Helios racks

AMD plans to integrate Taalas technology into its AI accelerator roadmap and develop system-level solutions combining Taalas chips with AMD Instinct GPUs and Epyc processors within Helios rack-scale systems

2

4

. Vamsi Boppana, senior vice president of AMD's Artificial Intelligence Group, stated that "AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload"

1

2

. The likely architecture will feature disaggregated systems where compute-intensive prompt processing happens on AMD Instinct GPUs while token generation offloads to Taalas-based accelerators

1

. This approach proves far more space and power efficient than Nvidia's LPX systems, which would require dozens of GPUs and at least 2,000 Groq LPUs to serve the same trillion-parameter model that Taalas could handle with just 50 accelerators at 20 billion parameters each

1

.

The tradeoff: speed versus flexibility in AI hardware

The technology comes with a significant constraint—once deployed, Taalas chips run only the specific model they were built for, requiring new silicon for different AI models

1

3

. However, Taalas claims the situation isn't as restrictive as it appears—only two metal layers need changing for new models rather than starting from scratch, and the company can tape out new designs in roughly two months using proprietary tools

1

4

. AMD CEO Lisa Su acknowledged in July that "there's no one-size-fits-all as it comes to chips," noting that GPUs will still dominate the AI chip market due to their flexibility with newly developed models

3

. The technology appears particularly suited for AI model developers, infrastructure providers, and inference services where etching weights into silicon costs 100x less than training frontier models

1

.

Why this acquisition matters for AI workloads

The Taalas deal arrives as specialized inference chips become critical for semiconductor makers responding to AI's shift from training to real-time, high-volume deployment

2

. The acquisition directly counters Nvidia's $20 billion licensing deal with Groq announced in December, which also targeted high-performance inference services for AI agents and code assistants

1

. Alternative chips like Taalas prove especially valuable for low-latency applications where time to first response matters most

3

. AMD's major customers including OpenAI, Anthropic, and Meta already deploy Instinct accelerators, positioning the company to negotiate deals where models like GPT or Claude run on combined Taalas and Instinct systems

1

. Watch for AMD's upcoming HC2 chip targeting 20 billion parameters this summer and potential announcements from major AI companies adopting this architecture for production workloads.

Source: SiliconANGLE

Source: SiliconANGLE

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved