14 Sources
[1]
Nvidia's Groq 3 LPU Signals AI's Inferencing Era
This week, over 30,000 people are descending upon San Jose, Calif., to attend Nvidia GTC, the so-called Superbowl of AI -- a nickname that may or may not have been coined by Nvidia. At the main event Jensen Huang, Nvidia CEO, took the stage to announce (among other things) a new line of next
[2]
GTC 2026: Ian Buck press Q&A transcript -- VP of Hyperscale and HPC speaks out on shelving CPX and shipping LPU decode this year
We sat down at GTC 2026 with Nvidia's VP of Hyperscale and HPC, Ian Buck. Following Nvidia's GTC 2026 keynote, where CEO Jensen Huang laid out the company's Vera Rubin architecture and the Groq 3 LPU acquisition, Nvidia VP of Hyperscale and HPC Ian Buck sat down with press for a Q&A session in San
[3]
Nvidia To Upgrade AI Chatbot Performance WIth New 'LPU' Chip
When he's not battling bugs and robots in Helldivers 2, Michael is reporting on AI, satellites, cybersecurity, PCs, and tech policy. To improve chatbot performance, Nvidia is going to sell a new kind of processor called an LPU, which has been optimized to run large language models. The "Nvidia
[4]
A closer look at Nvidia's Groq-powered LPX rack systems
From LPUs and GPUs to CPUs and switches, everything you need to know about Nvidia's latest kit GTC DEEP DIVE At Nvidia's GTC conference this week, CEO Jensen Huang finally addressed a $20 billion question he's dodged for months: Why spend so much to license AI chip startup Groq's tech and hire
[5]
Nvidia Groq 3 LPU and Groq LPX racks join Rubin platform at GTC -- SRAM-packed accelerator boosts 'every layer of the AI model on every token'
Nvidia's Vera Rubin platform is poised to massively power up the next generation of AI data centers, or "factories," as CEO Jensen Huang calls them, when those systems start arriving later this year. Today, during his GTC keynote, Huang revealed how Nvidia is using the IP it acquired from Groq last
[6]
Nvidia's next AI chip may move beyond the all-purpose GPU
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. Forward-looking: Nvidia is poised to move beyond the graphics processors that powered its rise to dominance. Nvidia GTC kicks off this week - now branded more explicitly as an "AI conference" -
[7]
Nvidia slaps Groq into new LPX racks for faster AI response
GPUzilla's $20B acquihire paves to way to AI agents that halucinate faster than ever GTC Nvidia will use Groq's language processing units (LPUs), a technology it paid $20 billion for, to boost the inference performance of its newly-announced Vera Rubin rack systems, CEO Jensen Huang revealed
[8]
NVIDIA Vera Rubin Opens Agentic AI Frontier
Seven New Chips in Full Production to Scale the World's Largest AI Factories With Configurable AI Infrastructure Optimized for Every Phase of AI, From Pretraining, Post-Training and Test-Time Scaling to Agentic Inference News Summary: The NVIDIA Vera Rubin platform is opening the next AI frontier
[9]
Nvidia introduces Vera Rubin, a seven-chip AI platform with OpenAI, Anthropic and Meta on board
Nvidia on Monday took the wraps off Vera Rubin, a sweeping new computing platform built from seven chips now in full production -- and backed by an extraordinary lineup of customers that includes Anthropic, OpenAI, Meta and Mistral AI, along with every major cloud provider. The message to the AI
[10]
Nvidia debuts the Groq 3 language processing unit, a dedicated inference chip for multiagent workloads - SiliconANGLE
Nvidia debuts the Groq 3 language processing unit, a dedicated inference chip for multiagent workloads Nvidia Corp. kicked off its annual GTC 2026 developer conference in San Jose, California today, announcing a number of new chips and computing platforms aimed at data center operators. Most of
[11]
Nvidia ups the stakes for AI infra with turbocharged Vera Rubin platform launch - SiliconANGLE
Nvidia ups the stakes for AI infra with turbocharged Vera Rubin platform launch Nvidia Corp. is throwing down the gauntlet to the rest of the artificial intelligence chip industry with the launch of its next-generation Vera Rubin platform. Announced at GTC 2026 today, it consists of no less than
[12]
NVIDIA Unveils Vera Rubin With Groq's LPX to Break Into Inference, a Market Where It Has Never Been First
NVIDIA's Groq partnership is now formalizing, as Jensen unveils a hybrid compute tray featuring Groq's third-generation LPU units in a Rubin rack. The debate over what NVIDIA would do with Groq has been ongoing for quite some time, and we have maintained a key lead on developments. At GTC 2026,
[13]
Nvidia Puts Groq LPU, Vera CPU And Bluefield-4 DPU Into New Data Center Racks
Announced at Nvidia's GTC 2026 event, the AI infrastructure giant's new Groq-based inference server rack, called the Nvidia Groq 3 LPX, will be available alongside the Vera Rubin NVL72 rack, Vera CPU rack and BlueField-4 STX storage rack in the second half of the year. Nvidia said Monday that it's
[14]
Nvidia unveils Vera Rubin AI platform with seven new chips By Investing.com
SAN JOSE, Calif. - Nvidia (NASDAQ:NVDA) announced today the Vera Rubin platform, featuring seven chips now in full production designed for AI infrastructure deployment, according to a company press release. The $4.43 trillion semiconductor giant continues to dominate the AI infrastructure market
Share
Copy Link
Nvidia introduced the Groq 3 LPU at GTC 2026, a specialized chip designed to accelerate AI inference using SRAM memory instead of traditional HBM. The chip stems from Nvidia's $20 billion acquisition of Groq's intellectual property and promises to deliver faster token generation for chatbots and AI agents through 256-chip LPX rack systems paired with Rubin GPUs.

Nvidia CEO Jensen Huang took the stage at GTC 2026 in San Jose to announce the Groq 3 LPU, a language processing unit that marks a strategic shift for the GPU giant toward specialized AI inference hardware
1
. The chip incorporates intellectual property Nvidia licensed from startup Groq last Christmas Eve for $20 billion, addressing what Huang called "the inflection point of inference"1
. The Groq 3 LPU joins Nvidia's Vera Rubin platform alongside the Rubin GPU, Vera CPU, and networking components to form a comprehensive data center architecture5
.The acquisition reflects mounting pressure on Nvidia to maintain dominance as competitors like AMD close gaps and Amazon Web Services deploys alternative inference solutions combining Trainium accelerators with Cerebras wafer-scale systems
4
. Ian Buck, Nvidia's VP of Hyperscale and HPC, explained the urgency: "We've pulled CPX" to focus resources on optimizing decode performance with the LPU this year2
.The Groq 3 LPU distinguishes itself through an SRAM-based architecture that prioritizes memory bandwidth over capacity. Each chip contains 500 MB of SRAM memory delivering 150 TB/s of bandwidth—seven times faster than the 22 TB/s offered by Rubin GPU's 288 GB of HBM4 memory
1
. The chip achieves 1.2 petaFLOPS of FP8 compute, though support for 4-bit block floating point data types won't arrive until the LP35 generation next year4
.Groq's approach interleaves processing units with memory units directly on the chip, eliminating the need for data to travel off-chip to HBM and back. "The data actually flows directly through the SRAM," said Mark Heaps, formerly Groq's chief technology evangelist and now Nvidia's director of developer marketing
1
. This linear data flow enables the extreme low-latency token generation required for interactive AI chatbot performance and reasoning models that run inference many times before users see output1
.Nvidia will deploy these chips in LPX rack systems containing 256 LPUs spread across 32 compute trays, each with eight LPUs plus fabric expansion logic, DRAM, a host CPU, and a BlueField-4 data processing unit
4
. A complete rack offers 128 GB of on-chip SRAM with 640 TB/s of scale-up bandwidth3
.Nvidia's strategy pairs LPX rack systems with Vera Rubin NVL72 server units to enable what Buck calls disaggregated inference. The decode phase splits between LPU and GPU on a layer-by-layer basis, with computations benefiting from fast SRAM running on the LPU while attention math, softmax, routing, and KV cache calculations execute on GPUs
2
. "We can focus and run the computations that benefit from the fast SRAM of the LPU over here in one layer, and literally the next layer, we can send the intermediate activation state over to the GPUs," Buck explained2
.This hybrid approach allows only the LPUs to store model weights while per-query state and KV cache data—which can grow quite large—remain in GPU HBM
2
. The combined system promises up to 35x throughput increases when running large language models reaching 1 trillion parameters, according to Nvidia benchmarks3
. Buck positioned this capability for multi-agent systems requiring interactive performance while inferencing trillion-parameter models with context windows of millions of tokens5
.Related Stories
The Groq 3 chip is based on Groq's second-generation LPU technology with last-minute tweaks before manufacturing at Samsung fabs
4
. Notably, it lacks NVLink interconnect, NVFP4 hardware support, and CUDA compatibility at launch—indicating Nvidia prioritized speed over deep integration4
. The $20 billion represented an opportunity cost to ship products this year rather than build from scratch4
.Huang indicated the chip will ship in Q3, with one analyst projecting 4 to 5 million LPU shipments through 2026 and 2027
3
. The systems target major AI companies including OpenAI, Anthropic, and Meta, potentially powering chatbot queries and image generation requests3
. Huang suggested high-performance, low-latency inference providers could eventually charge as much as $150 per million tokens for this capability4
.The move addresses a fundamental shift in AI economics as computational load transitions from building larger models to deploying them at scale. D-Matrix CEO Sid Sheth noted that "winning systems will combine different types of silicon and fit easily into existing data centers alongside GPUs"
1
. For AI agents communicating with other AIs rather than humans, Buck envisions moving from 100 tokens per second to 1,500 TPS or more5
.Summarized by
Navi
[2]
[3]
[4]
03 Jan 2026•Technology

24 Aug 2026•Technology

19 Mar 2025•Technology

1
Science and Research

2
Policy and Regulation

3
Technology