9 Sources
[1]
What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble
Nvidia's $20 billion bet on Groq's LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse of just how big a speedup its Groq 3-based LPX racks will provide. In an independent benchmark conducted by Artificial Analysis, Nvidia's LPX rack systems managed to
[2]
Nvidia says Groq racks will be online this year following $20 billion purchase
* Nvidia announced that its Groq 3 LPX chip is in full production, marking the commercialization of technology from the company's largest-ever acquisition. * Nvidia's race to manufacture the chip and make it available to customers highlights the importance of low-latency inference that's needed to
[3]
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
The next era of AI inference won't be defined by a single breakthrough chip, network or system. It'll be defined by how every layer of the AI factory works together. That's why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera
[4]
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72. According to OpenRouter data, agentic AI workloads consume 15x more tokens
[5]
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
Groq 3 LPX Extends the Vera Rubin Platform, Delivering Ultrafast Token Generation for the Next Generation of Agentic AI, With Nebius the First to Adopt * In Artificial Analysis benchmarking, Groq 3 LPX showcased world-class speed for agentic coding and other latency-sensitive workloads. * NVIDIA
[6]
Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI
[7]
NVIDIA Vera Rubin Records A Massive 30x Increase In Throughput Per Watt Than Blackwell While Offering a 35x Token Cost Reduction Across Agentic AI Workloads
NVIDIA's Vera Rubin platform offers disruptive token throughput at a much lower cost than Blackwell, showcasing its Agentic AI prowess. NVIDIA Software Stack Optimizations Continue To Propel Blackwell, But Vera Rubin Sits In A Whole Different League For Agentic AI Use Cases The latest on-silicon
[8]
NVIDIA Enters Full Production of Groq 3 LPX AI Inference Accelerator Chips, Supercharging Vera Rubin With The Fastest Token Generation Speeds Ever Recorded
NVIDIA's Groq 3 LPX is now in full production, offering big token generation speedups for Vera Rubin platforms in the Agentic AI space. NVIDIA Dials Up Vera Rubin NVL72 Token Generation Capabilities, Recording 3,400 TPS With Groq 3 LPX AI Inference Accelerators As part of its Hot Chips 2026
[9]
Nvidia readies the launch of its Groq racks dedicated to inference
Each LPX rack brings together 256 Groq 3 chips, manufactured by Samsung, each integrating 500 megabytes of ultra-high-speed SRAM. According to Nvidia, the system can reach 3,400 tokens per second, performance designed to enable cloud providers to offer premium low-latency services. These
Share
Copy Link
NVIDIA announced its Groq 3 LPX is in full production, delivering record 3,400 tokens per second in benchmarks—4x faster than nearest alternatives. The AI inference accelerator extends the Vera Rubin platform for ultrafast token generation in agentic AI workloads. Netherlands-based Nebius becomes the first AI cloud to deploy the technology following NVIDIA's $20 billion Groq acquisition.
NVIDIA announced Monday that its Groq 3 LPX rack system is in full production, marking the commercialization of technology from the company's largest acquisition on record
2
. The GPU giant acquired assets from chip startup Groq for $20 billion in December, betting that low-latency inference would become critical as AI shifts toward agentic workloads2
. The Groq 3 LPX extends the Vera Rubin platform by dramatically increasing token generation rates for highly responsive agentic systems5
.
Source: NVIDIA
NVIDIA senior director Dion Harris told reporters the Groq rack will be deployed alongside Vera central processors and Rubin graphics processors at neocloud Nebius, with systems coming online later this year
2
. Nebius becomes the first AI cloud to adopt NVIDIA Groq 3 LPX, giving developers access to extreme token generation speed through Nebius Token Factory5
.In an independent benchmark conducted by Artificial Analysis, NVIDIA's Groq 3 LPX rack systems delivered 3,400 tokens per second with a 100,000-token input sequence running Google's Gemma 4 31B model
1
. According to NVIDIA, this makes the interactive AI inference accelerator 4x faster than the nearest alternative platform, which appears to be Cerebras at 882 tokens per second under identical conditions1
.The speed advantage stems from Groq's SRAM-heavy dataflow architecture designed specifically for high-performance inference serving
1
. Unlike traditional datacenter GPUs that rely on GDDR7 and HBM4 memory, Groq chips use entirely on-die SRAM that's orders of magnitude faster than even the best HBM stacks available today1
. The third-generation chips boast 150 TB/s of memory bandwidth compared to around 2.75 TB/s for top HBM stacks1
.The faster token generation directly addresses emerging demands from agentic AI workloads. According to OpenRouter data cited by NVIDIA, agentic AI workloads consume 15x more tokens than a simple chat request
4
. When an AI agent researches a company for investment decisions, it queries financial databases, searches news and filings, invokes sub-agents to run peer comparisons and model valuations, then synthesizes everything into recommendations4
.
Source: NVIDIA
Harris explained that faster token generation unlocks the ability for cloud companies to offer premium tiers of service for customers demanding the most latency-sensitive service agreements
2
. The faster models can generate tokens, the longer they can reason, the more turns agents can take, and the more information they can process or actions they can take in the same window of time1
. Groq 3 LPX enables agentic tasks such as coding in minutes versus hours5
.Each Groq 3 LPU contains just 500 MB of on-die memory—576x less than NVIDIA's top-specced Rubin GPU with 288 GB
1
. SRAM consumes significant die area, limiting total capacity per chip. NVIDIA packages 256 individual Groq 3 chips into its LPX racks, manufactured by Samsung, providing 128 GB of high-bandwidth SRAM total2
1
.NVIDIA's architecture uses Ethernet to distribute models across multiple accelerators through distributed model execution
1
. The Gemma 4 31B model running at FP8 precision requires just over 31 GB, fitting neatly into a single LPX rack with enough capacity to potentially hold four copies using pipeline parallelism and data parallelism for higher concurrency1
. However, serving larger models like DeepSeek V3 at 671 billion parameters would require 1,342 accelerators or just over 5 LPX racks1
.Related Stories
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72
4
. NVIDIA measured this using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions with actual context growth, tool calls and sub-agent spawning preserved4
.
Source: The Register
For power-constrained AI factories, throughput per megawatt determines revenue while cost per million tokens determines profit margin
4
. NVIDIA DSX MaxLPS technologies manage power across GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget4
.Jensen Huang emphasized that breakthrough performance comes from extreme codesign across every layer of the stack rather than optimizing individual components in isolation
3
. Through extreme codesign across seven chips and five purpose-built racks, NVIDIA Vera Rubin represents the most extensive AI factory platform5
.Industry partners are rapidly adopting Vera Rubin platform solutions. SpaceXAI announced NVIDIA Vera CPUs will power its next generation of agentic AI from data centers on Earth to orbital satellites, deploying them to accelerate CPU-intensive work including orchestration, tool use, code execution, data processing and simulation
3
. CoreWeave has deployed Spectrum-X Multiplane into production, connecting NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks3
.Modern inference optimization for agentic AI workloads spans disaggregated serving that separates context processing from response generation, rate matching to synchronize prefill and decode speeds, large-scale expert parallelism for mixture-of-experts models, and distributed KV-caching that extends memory across the scale-up GPU domain
4
. Harris clarified this isn't about replacing GPUs but using the right processor for the right part of the workload2
. At the Vera Rubin and Groq 3 LPX unveiling in March, Huang projected $1 trillion in cumulative sales between current-generation Blackwell chips and new Vera Rubin systems through 20272
.Summarized by
Navi
[1]
[4]
28 Feb 2026•Technology

17 Jul 2026•Technology

16 Mar 2026•Technology

1
Science and Research

2
Policy and Regulation

3
Technology