Nvidia Vera Rubin enters full production with 10x efficiency gains for AI inference workloads

11 Sources

Share

Nvidia's Vera Rubin platform has reached full production, with CoreWeave reporting 10 times more tokens per watt compared to the previous GB200 NVL72 system. Major customers including OpenAI, Microsoft Azure, Google Cloud, and Meta are deploying the next-generation AI infrastructure, which combines 72 Rubin GPUs with custom Vera CPUs to handle agentic AI workloads more efficiently.

Nvidia Vera Rubin Reaches Full Production With Massive Efficiency Gains

Nvidia Vera Rubin has officially entered full production, marking a significant milestone in the company's push to dominate next-generation AI infrastructure. Ian Buck, vice president of accelerated computing at Nvidia, confirmed that systems are now shipping to major customers including OpenAI, CoreWeave, Microsoft Azure, Google Cloud, Meta, Oracle Cloud Infrastructure, and Dell

4

. OpenAI plans to deploy the AI platform at scale during the third quarter, according to reports from the company's headquarters briefing

4

.

Source: Wccftech

Source: Wccftech

The production ramp represents the culmination of extreme hardware-software co-design across seven chips and five rack trays, all engineered as a single system rather than assembled from separate components

3

. This approach targets the specific demands of agentic AI workloads, which can consume up to 15 times more tokens than traditional AI applications

3

.

CoreWeave Reports 10x Tokens Per Watt on DeepSeek-R1

The most striking performance benchmark comes from CoreWeave, which reported achieving 10 times more tokens per watt when running DeepSeek-R1 on the Vera Rubin NVL72 system compared to the previous generation GB200 NVL72 based on Blackwell

3

. This dramatic improvement in performance per watt directly addresses the power constraints that define modern AI factories, where delivered performance within a fixed watt data center determines profitability

2

.

CoreWeave also demonstrated that optimizations developed for Vera Rubin can be applied retroactively to GB200 NVL72 systems, improving throughput per megawatt by more than four times over a three-month period

5

. These gains were validated across 250,000-plus distinct configurations and more than 1.4 million GPU hours of testing

5

.

Rubin GPU Architecture Optimized for AI Inference at Scale

The Rubin GPU sits at the heart of the platform's inference capabilities. Each chip joins two compute dies using the Nvidia High Bandwidth Interface, delivering 224 Streaming Multiprocessors containing 896 Tensor Cores alongside 288GB of HBM4 memory with 22 TB/s of memory bandwidth

1

. As an inference-focused accelerator, Nvidia highlights Rubin's 50 sparse PFLOPS of NVFP4 inference throughput as its headline performance figure

1

.

Source: The Register

Source: The Register

Key architectural improvements target the specific challenges of modern transformer-based LLMs and mixture-of-experts models. The Tensor Memory Accelerator now supports GPU kernels that maintain a single unified MoE descriptor directly in the instruction at runtime, reducing computation overhead as expert counts grow

1

. Rubin also doubles the K-dimension throughput in Tensor Cores, cutting the number of loop iterations required for matrix multiplication in half compared to Blackwell .

Softmax performance, essential for attention calculations in long-context models, receives up to a 4x boost versus Blackwell through enhancements to the Special Function Unit

1

. These improvements matter as advanced models now support context lengths of up to one million tokens.

Vera CPU Built for Agentic AI Workloads

The Vera CPU represents Nvidia's bet that faster cores beat more cores for agentic AI workloads. Built on the custom Olympus core microarchitecture, the processor delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs

3

. In recent benchmarks, Nvidia demonstrated 1.9 times faster agentic performance and a six-fold improvement in latency over x86-based alternatives

5

.

The company also claimed the Vera CPU surpasses AMD's EPYC Turin chip by nearly 100% on selected industry benchmarks, particularly Python workloads that dominate the software stack for AI inference

4

. Anthropic and OpenAI are among the first labs to receive the processor, alongside Perplexity, SpaceX, and Oracle

4

.

NVLink 6 and Spectrum-X Networking Accelerate AI Factories

Networking infrastructure plays a critical role in the platform's efficiency. The sixth-generation NVLink 6 scale-up fabric delivers more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet

3

. For scale-out networking, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs, enabling 1.6x higher RDMA bandwidth than standard Ethernet

3

.

Source: Tom's Hardware

Source: Tom's Hardware

Leading AI infrastructure builders including CoreWeave, Microsoft, SpaceXAI, and Tesla are among the first to deploy Spectrum-6 switches in their data centers

3

. The sixth-generation NVLink delivered 2.3 times higher simulated decode throughput for massive language models compared to Ethernet-based networks

5

.

Rapid Assembly and Water Savings Transform Data Center Operations

Three generations of rack-scale co-design produced a Vera Rubin NVL72 system with no cables, fans, or hoses in the compute tray, cutting assembly time from 90 minutes to just one minute—a 90x improvement

2

. Andrew Bell, senior vice president of hardware engineering, demonstrated the streamlined installation process during a lab tour at Nvidia's Sunnyvale facility

2

.

The platform's 45-degree Celsius liquid cooling inlet temperature design enables chiller-free dry-cooler operation, saving approximately 4 million gallons of water per megawatt annually compared to standard cooling methods

3

. By dynamically optimizing the full infrastructure and energy stack, Nvidia can deploy 40% more GPUs within the same power envelope

5

.

Global Supply Chain Supports Gigascale Deployment

Vera Rubin production spans 350-plus factory sites across 30 countries, representing what Nvidia calls the largest and most mature rack-scale supply chain ever assembled

3

. The platform underpins a newly expanded partnership between Microsoft and Mistral that brings frontier AI to Europe, with Mistral adding GPU capacity drawing on thousands of Rubin GPUs

3

.

Jensen Huang first declared the platform in full production at Computex in early June, naming Anthropic, OpenAI, SpaceX, and Oracle as early recipients

4

. The recent headquarters briefing added performance numbers and customer testimonials that transform the Computex announcement into measurable benchmarks

4

.

The timing proves awkward for Nvidia's stock performance. While the company's shares have risen 9% this year, the broader chip index has climbed 66% over the same period, with Intel, ARM, and AMD all more than doubling

4

. Analysts project Nvidia's revenue will increase 82% to roughly $393 billion for the fiscal year, though this growth hasn't translated into the share price momentum seen during the Blackwell cycle

4

.

The business model centers on token economics. Buck emphasized that AI factory revenue depends on how many tokens can be generated in a fixed watt data center, with infrastructure rented for dollars per hour generating revenue through profitable tokens in the hundreds of dollars per hour

2

. As token supply increases through more efficient hardware and competing providers, downward pressure on token cost creates uncertainty about whether demand will grow fast enough to justify current capital expenditures

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved