NVIDIA's Groq 3 LPX Achieves 3,400 Tokens/Second, 4x Faster Than Competitors in Agentic AI Push

8 Sources

Share

NVIDIA announced its Groq 3 LPX is in full production, delivering record 3,400 tokens per second in benchmarks—4x faster than nearest alternatives. The AI inference accelerator extends the Vera Rubin platform for ultrafast token generation in agentic AI workloads. Netherlands-based Nebius becomes the first AI cloud to deploy the technology following NVIDIA's $20 billion Groq acquisition.

NVIDIA Ships Groq 3 LPX After $20 Billion Acquisition

NVIDIA announced Monday that its Groq 3 LPX rack system is in full production, marking the commercialization of technology from the company's largest acquisition on record

2

. The GPU giant acquired assets from chip startup Groq for $20 billion in December, betting that low-latency inference would become critical as AI shifts toward agentic workloads

2

. The Groq 3 LPX extends the Vera Rubin platform by dramatically increasing token generation rates for highly responsive agentic systems

5

.

Source: NVIDIA

Source: NVIDIA

NVIDIA senior director Dion Harris told reporters the Groq rack will be deployed alongside Vera central processors and Rubin graphics processors at neocloud Nebius, with systems coming online later this year

2

. Nebius becomes the first AI cloud to adopt NVIDIA Groq 3 LPX, giving developers access to extreme token generation speed through Nebius Token Factory

5

.

Record-Breaking Performance: 3,400 Tokens Per Second

In an independent benchmark conducted by Artificial Analysis, NVIDIA's Groq 3 LPX rack systems delivered 3,400 tokens per second with a 100,000-token input sequence running Google's Gemma 4 31B model

1

. According to NVIDIA, this makes the interactive AI inference accelerator 4x faster than the nearest alternative platform, which appears to be Cerebras at 882 tokens per second under identical conditions

1

.

The speed advantage stems from Groq's SRAM-heavy dataflow architecture designed specifically for high-performance inference serving

1

. Unlike traditional datacenter GPUs that rely on GDDR7 and HBM4 memory, Groq chips use entirely on-die SRAM that's orders of magnitude faster than even the best HBM stacks available today

1

. The third-generation chips boast 150 TB/s of memory bandwidth compared to around 2.75 TB/s for top HBM stacks

1

.

Why Ultrafast Token Generation Matters for Agentic AI

The faster token generation directly addresses emerging demands from agentic AI workloads. According to OpenRouter data cited by NVIDIA, agentic AI workloads consume 15x more tokens than a simple chat request

4

. When an AI agent researches a company for investment decisions, it queries financial databases, searches news and filings, invokes sub-agents to run peer comparisons and model valuations, then synthesizes everything into recommendations

4

.

Source: NVIDIA

Source: NVIDIA

Harris explained that faster token generation unlocks the ability for cloud companies to offer premium tiers of service for customers demanding the most latency-sensitive service agreements

2

. The faster models can generate tokens, the longer they can reason, the more turns agents can take, and the more information they can process or actions they can take in the same window of time

1

. Groq 3 LPX enables agentic tasks such as coding in minutes versus hours

5

.

Technical Architecture and Scale Challenges

Each Groq 3 LPU contains just 500 MB of on-die memory—576x less than NVIDIA's top-specced Rubin GPU with 288 GB

1

. SRAM consumes significant die area, limiting total capacity per chip. NVIDIA packages 256 individual Groq 3 chips into its LPX racks, manufactured by Samsung, providing 128 GB of high-bandwidth SRAM total

2

1

.

NVIDIA's architecture uses Ethernet to distribute models across multiple accelerators through distributed model execution

1

. The Gemma 4 31B model running at FP8 precision requires just over 31 GB, fitting neatly into a single LPX rack with enough capacity to potentially hold four copies using pipeline parallelism and data parallelism for higher concurrency

1

. However, serving larger models like DeepSeek V3 at 671 billion parameters would require 1,342 accelerators or just over 5 LPX racks

1

.

Vera Rubin Platform Shows Massive Efficiency Gains

New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72

4

. NVIDIA measured this using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions with actual context growth, tool calls and sub-agent spawning preserved

4

.

Source: The Register

Source: The Register

For power-constrained AI factories, throughput per megawatt determines revenue while cost per million tokens determines profit margin

4

. NVIDIA DSX MaxLPS technologies manage power across GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget

4

.

Extreme Codesign Strategy and Industry Partnerships

Jensen Huang emphasized that breakthrough performance comes from extreme codesign across every layer of the stack rather than optimizing individual components in isolation

3

. Through extreme codesign across seven chips and five purpose-built racks, NVIDIA Vera Rubin represents the most extensive AI factory platform

5

.

Industry partners are rapidly adopting Vera Rubin platform solutions. SpaceXAI announced NVIDIA Vera CPUs will power its next generation of agentic AI from data centers on Earth to orbital satellites, deploying them to accelerate CPU-intensive work including orchestration, tool use, code execution, data processing and simulation

3

. CoreWeave has deployed Spectrum-X Multiplane into production, connecting NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks

3

.

Modern inference optimization for agentic AI workloads spans disaggregated serving that separates context processing from response generation, rate matching to synchronize prefill and decode speeds, large-scale expert parallelism for mixture-of-experts models, and distributed KV-caching that extends memory across the scale-up GPU domain

4

. Harris clarified this isn't about replacing GPUs but using the right processor for the right part of the workload

2

. At the Vera Rubin and Groq 3 LPX unveiling in March, Huang projected $1 trillion in cumulative sales between current-generation Blackwell chips and new Vera Rubin systems through 2027

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved