6 Sources
[1]
AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon
In AMD's latest bid to upset Nvidia's dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more. The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia's $20 billion licensing deal with Groq last December: make high-performance "premium" inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn't disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire. Founded in 2023 and based in Toronto, Taalas' approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras' waferscale accelerators. A model-specific integrated circuit The startup's chips don't rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas' chips are really model-specific integrated circuits or MSICs. Perhaps more importantly, Taalas' tech isn't just conceptual. In February, the startup revealed its first test chip fabbed on TSMC's 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta's Llama 3.1 8B at a blistering 16,960 tokens a second -- when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators. While Llama 3.1 is ancient by today's standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept. Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored. For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism. At 20 billion parameters per chip, you'd need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that. That's quite a bit more space and power efficient than Nvidia's recently unveiled LPX systems, which would need a few dozen GPUs and at least 2,000 Groq LPUs to serve the same model. From what we understand, AMD intends to pair its Instinct-based Helios racks with chips based on Taalas' tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators. It's also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they're satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here's what AMD's SVP of AI, Vamsi Boppana, had to say about it in a canned statement: "AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload." You better really love that model While the tech is blazing fast, if you hadn't already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you're stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming. Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas' tech, AMD's customers are going to have to be really sure about their choice of models, which will be easier for some than others. However, if the startup is to be believed, the situation isn't quite as bad as it sounds. While new models will require a re-spin, it doesn't require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming. With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site The Next Platform in February, the company suggested that etching a model's weights into silicon is 100x less expensive than training a frontier model. AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn't be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators. The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to "think" for longer before responding. One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. If AMD's Taalas buy can drive down the cost per token and boost output speeds by 10x or 20x, model devs may opt to extend the reasoning time even further. In any case, we may not have to wait long to see just how Taalas fits into AMD's broader vision. Subject to regulatory approval, the deal is expected to close in the fourth quarter. ®
[2]
AMD deepens AI inference bet with Taalas deal as chip race heats up
Aug 6 (Reuters) - Advanced Micro Devices (AMD.O), opens new tab said on Thursday it would buy chip startup Taalas for an undisclosed amount, strengthening its technology as it taps into the rapidly growing market for AI inference workloads. Toronto-based Taalas develops specialized silicon designed to reduce computing and memory bottlenecks in AI inference, the process of running trained AI models to generate responses or predictions. Specialized inference chips have become a critical focus for semiconductor makers as AI shifts from training to real-time, high-volume deployment, and companies race to cut computing costs. Larger rival Nvidia (NVDA.O), opens new tab in March unveiled a new central processor and AI system built on technology from Groq, a chip startup specializing in inference. AMD plans to integrate Taalas' technology into its accelerator roadmap and develop system-level solutions using AMD Instinct graphics processing units (GPUs). "AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload," Vamsi Boppana, senior vice president of AMD's Artificial Intelligence Group, said in a statement. Taalas' technology and engineering team strengthens AMD's AI portfolio by "delivering differentiated inference performance and efficiency," Boppana added. Taalas, founded in 2023, had raised $169 million in February to support the development of chips optimized for AI models, bringing its total funding to around $219 million. AMD has been bolstering its AI business through a string of acquisitions. In November, the company said it bought MK1, an AI software startup specializing in high-speed inference. The company purchased MEXT to expand its AI portfolio in June and added FastFlowLM to its Artificial Intelligence Group in July. Reporting by Juby Babu in Mexico City; Editing by Sriraj Kalluvila Our Standards: The Thomson Reuters Trust Principles., opens new tab
[3]
AMD buys chip startup that hardwires AI models into its silicon
Advanced Micro Devices is counting on its graphics processing units to drive the bulk of its data center growth as cloud companies snap up all the advanced AI chips they can find. But as the generative artificial intelligence boom approaches its fourth anniversary, it's becoming clear that GPUs don't do everything. On Thursday, AMD said it's entered into an agreement to acquire Taalas, a Toronto-based startup that makes chips for inference. Taalas' accelerators are customized, or hard-wired for a single AI model, rather than being general purpose. In exchange for that loss of flexibility, Taalas' technology promises a less-expensive chip that it says can produce output for specific models thousands of times faster than a traditional GPU. An AMD representative declined to provide a purchase price for the transaction. Taalas has raised a total of $219 million in venture funding since its 2023 founding. The deal comes a little over seven months after Nvidia spent $20 billion buying assets from Groq, a designer of high-performance AI chips. It was Nvidia's largest transaction on record. Taalas' current chip runs a small version of Meta's Llama 3.1 model, though the company is working on chips for bigger and more advanced models. It's manufactured using an older TSMC process, and uses speedy SRAM memory on the chip itself. Taalas CEO Ljubisa Bajic says on the startup's website that the company "developed a platform for transforming any AI model into custom silicon." "From the moment a previously unseen model is received, it can be realized in hardware in only two months," Bajic wrote. Alternative chips like those from Taalas and Groq are particularly important for "low-latency" applications, where time to first response from an AI model is important. "I'm a big believer that there's no one-size-fits-all as it comes to chips," AMD CEO Lisa Su said at a product launch in July. She added that AMD still expects GPUs to make up the majority of the AI chip market because they're flexible enough to support newly developed AI models. Demand for GPUs has turned Nvidia into the world's most valuable company with a market cap of over $5 trillion. The acquisition also reflects the rising importance for leading GPU makers to offer integrated systems with several different components and chips instead of just processors. AMD recently started to ship Helios, its first rack-scale rival to Nvidia's integrated server racks, to customers including Meta and Microsoft. AMD said it would integrate the Taalas chips and technology into its roadmap, including in systems with its central processors and Instinct GPUs. It's not the only complimentary accelerator that AMD is supporting: In July, AMD announced a partnership with Cerebras to integrate its AI chips into its systems later this year. AMD has been on a buying spree to fill out some of the components and technologies it needs to build its Helios racks. In 2024, it paid $665 million for Silo AI, which develops AI models, and purchased ZT Systems, which provided the technical basis of its rack-scale products, for $4.9 billion. Last year, AMD bought several smaller AI companies, including MK1, which made software for inference.
[4]
AMD acquires Taalas to hardwire AI models into silicon
Advanced Micro Devices Inc. said today it has agreed to buy Taalas Inc., a Toronto startup that hardwires artificial intelligence models directly into silicon, in a deal that pushes the chipmaker deeper into the market for AI inference. Terms were not disclosed. AMD shares rose about 1.5% following the announcement. Founded in 2023, Taalas builds what it calls model-specific integrated circuits, chips that cast a model's weights and dataflow into transistors instead of shuttling them in and out of high-bandwidth memory. Its first test chip, HC1, was built on Taiwan Semiconductor Manufacturing Co. Ltd.'s 6-nanometer process and served Meta Platforms Inc.'s Llama 3.1 8B at close to 17,000 tokens per second, a rate Taalas said in February was 73 times that of Nvidia Corp.'s H200 at one-tenth the power. A second chip, HC2, targets models of about 20 billion parameters. The tradeoff is flexibility. A finished Taalas part runs the model it was built for and nothing else, so a new model means new silicon. Taalas says that is cheaper than it sounds. Only two of the chip's 100-plus layers change from one design to the next, and the company puts its tape-out time at roughly two months using in-house tools. The technology is headed for AMD's accelerator roadmap. The larger plan is systems, not just parts. AMD wants Taalas chips working next to its Instinct GPUs, wired into the Helios racks and Epyc processors it sells and programmed through its ROCm software. Vamsi Boppana, senior vice president of AMD's Artificial Intelligence Group, said the aim is choice, giving customers "the right compute solutions for every AI workload." Taalas brings what he called differentiated inference performance and efficiency. Taalas co-founder and Chief Executive Ljubisa Bajic is on his second AI chip company. He ran Tenstorrent Inc. before leaving to start Taalas. The startup was founded "to rethink AI inference from the ground up by building the hardware around the model," he said. His Toronto team joins Boppana's group when the deal closes. The purchase lands in a market that has reorganized itself around inference over the past year. Nvidia agreed in December to license inference technology from Groq Inc. in a reported $20 billion deal, then put it to work in the Groq 3 language processing unit unveiled at its GTC conference in March. Taalas is AMD's third AI acquisition in nine months, following MK1 in November and memory optimization startup Mext in June. The FastFlowLM team joined the company in July. Taalas raised $169 million in February from Quiet Capital, Fidelity and semiconductor investor Pierre Lamond, taking its total to about $219 million. The acquisition is subject to customary closing conditions and regulatory approvals.
[5]
AMD Snaps Up Taalas Weeks After Cerebras Deal, Chasing Chips That Bake AI Models Into Silicon
Just a few weeks after AMD inked a deal with Cerebras to boost the inference capabilities of the upcoming Helios rack-scale system, it has acquired Taalas, a startup that designs chips specifically for inference-related workloads. Taalas bakes AI models directly into silicon chips, and AMD has just inked a definitive agreement to acquire it AMD has just signed an agreement to acquire Taalas in a move designed to complement its Helios systems. For the benefit of those might not be aware, Taalas eliminates the so-called memory wall - an increasingly common situation where the GPUs sit idle as they wait for the appropriate data to arrive to then start crunching numbers - by hardwiring a given AI model into individual transistors, courtesy of a proprietary digital architecture that is capable of storing 4 bits of data and executing math operations using a single transistor. Basically, Taalas HC1 chip features shared hardware blocks that permanently and constantly pre-compute all 16 possible products for a quantized 4-bit weight. Because every possible math outcome is already happening live across the chip, the individual transistor doesn't actually do any math but acts as a physical router. Here, each specific weight of the AI model is represented by a single Mask ROM transistor - during the manufacturing process, a microscopic physical wire (a metal layer mask) is etched to physically connect that transistor to one of the 16 pre-computed product lines. When data flows through the chip, the transistor simply selects the correct pre-calculated mathematical channel and passes it to the adder. Therefore, Taalas' HC1 chip is capable of generating an astonishing 16,000 to 17,000 tokens per second per user by using this setup. This approach is somewhat similar to the Groq LPU, which uses small isolated pools of the SRAM, and where AI model weights are hard-baked directly into the SRAM, completely bypassing the concept of a memory cache. AMD's decision to acquire Taalas complements its integration of the Helios system with Cerebras' Wafer-Scale Engine, which places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, and where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks. Follow Wccftech on Google to get more of our news coverage in your feeds.
[6]
AMD to Buy AI Chip Startup Taalas
AMD has reached a deal to buy the AI chip startup Taalas at undisclosed terms. AMD said that the deal would strengthen its long-term AI ambitions by adding a differentiated inference technology and engineering expertise. Founded in 2023 and based in Toronto, Taalas optimizes inference dataflows, helping to reduce compute and memory bottlenext associated with general-purpose architectures and enabling highly optimized AI inference capabilities. AMD said that Taalas' technology will complement AMD's full-stack AI platform. It eventually plans to integrate the technology into its accelerator roadmap and develop system-level solutions with AMD Instinct GPUs. The deal is subject to customary closing conditions and regulatory approvals.
Share
Copy Link
AMD acquired Toronto-based startup Taalas to boost AI inference performance by hardwiring models directly into silicon. Taalas' HC1 chip serves Llama 3.1 at 17,000 tokens per second, 48x faster than Nvidia GPUs. The technology will integrate with AMD Instinct GPUs and Helios racks to compete in the growing AI inference market.
AMD acquires Taalas, a Toronto-based startup that etches AI models directly into silicon, marking the chipmaker's latest move to compete with Nvidia in the rapidly expanding AI inference market
1
2
. The deal, announced at market close on Thursday, positions AMD to deliver what it calls "differentiated inference performance and efficiency" as companies race to cut computing costs for real-time AI deployment2
. While AMD didn't disclose financial terms, Taalas had raised $219 million in total funding since its 2023 founding, including a $169 million round in February2
4
. The acquisition represents AMD's third AI chip deal in nine months, following MK1 in November and Mext in June4
.
Source: The Register
Taalas builds model-specific integrated circuits that permanently etch AI model weights directly into silicon transistors, eliminating the need for high-bandwidth memory and solving the memory wall problem where GPUs sit idle waiting for data
1
5
. The startup's HC1 chip, fabricated on TSMC's 6-nanometer process, achieved 16,960 tokens per second when serving Meta's Llama 3.1 8B model—48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at one-tenth the power1
4
. The technology uses a proprietary digital architecture where each transistor stores 4 bits of data and acts as a physical router rather than performing calculations, with shared hardware blocks pre-computing all possible mathematical outcomes5
. Taalas chips comprise two main regions: a mask-ROM recall fabric where model weights are permanently etched, and an SRAM recall fabric for KV caches and fine-tuning adapters1
.
Source: Wccftech
AMD plans to integrate Taalas technology into its AI accelerator roadmap and develop system-level solutions combining Taalas chips with AMD Instinct GPUs and Epyc processors within Helios rack-scale systems
2
4
. Vamsi Boppana, senior vice president of AMD's Artificial Intelligence Group, stated that "AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload"1
2
. The likely architecture will feature disaggregated systems where compute-intensive prompt processing happens on AMD Instinct GPUs while token generation offloads to Taalas-based accelerators1
. This approach proves far more space and power efficient than Nvidia's LPX systems, which would require dozens of GPUs and at least 2,000 Groq LPUs to serve the same trillion-parameter model that Taalas could handle with just 50 accelerators at 20 billion parameters each1
.Related Stories
The technology comes with a significant constraint—once deployed, Taalas chips run only the specific model they were built for, requiring new silicon for different AI models
1
3
. However, Taalas claims the situation isn't as restrictive as it appears—only two metal layers need changing for new models rather than starting from scratch, and the company can tape out new designs in roughly two months using proprietary tools1
4
. AMD CEO Lisa Su acknowledged in July that "there's no one-size-fits-all as it comes to chips," noting that GPUs will still dominate the AI chip market due to their flexibility with newly developed models3
. The technology appears particularly suited for AI model developers, infrastructure providers, and inference services where etching weights into silicon costs 100x less than training frontier models1
.The Taalas deal arrives as specialized inference chips become critical for semiconductor makers responding to AI's shift from training to real-time, high-volume deployment
2
. The acquisition directly counters Nvidia's $20 billion licensing deal with Groq announced in December, which also targeted high-performance inference services for AI agents and code assistants1
. Alternative chips like Taalas prove especially valuable for low-latency applications where time to first response matters most3
. AMD's major customers including OpenAI, Anthropic, and Meta already deploy Instinct accelerators, positioning the company to negotiate deals where models like GPT or Claude run on combined Taalas and Instinct systems1
. Watch for AMD's upcoming HC2 chip targeting 20 billion parameters this summer and potential announcements from major AI companies adopting this architecture for production workloads.
Source: SiliconANGLE
Summarized by
Navi
[1]
[4]
05 Jun 2025•Technology

11 Nov 2025•Business and Economy

05 Oct 2024•Technology

1
Technology

2
Technology

3
Science and Research
