NVIDIA outlined its AI factory strategy focused on maximizing return on investment through three pillars: productivity, durability, and fungibility. With each megawatt factory costing $60 million, the company's Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt and up to 45x lower cost per million tokens compared to GB300 NVL72.

NVIDIA Redefines AI Factory Economics Around Three Core Pillars

NVIDIA AI Factories are engineered to maximize return on investment through three interconnected principles: productivity, durability, and fungibility. Each megawatt AI factory costs roughly $60 million, making ROI clarity essential for operators committing capital at this scale

1

. Ian Buck, vice president and general manager of hyperscale and HPC at NVIDIA, emphasized that these assets function as "revenue-generating, fungible, durable, productive parts of an economy" rather than traditional IT costs

2

.

Source: NVIDIA

Source: NVIDIA

The company's approach addresses a fundamental constraint: power capacity limits how much AI infrastructure a data center can deploy. This makes tokens per watt the central measure of AI factory economics and positions efficient AI inference as a competitive advantage. NVIDIA's system-level infrastructure codesign across the full stack—from models and workloads through software to compute, networking, and memory—enables these efficiency gains.

Productivity: 30x Efficiency Gains Through Power-Optimized Design

Power serves as the binding constraint on AI factories, making tokens per second per megawatt the governing metric for earning capacity. More tokens within a fixed power envelope translates directly to higher revenue, while lower cost per million tokens improves margins. According to SemiAnalysis AgentX data, NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than GB300 NVL72, with up to 45x lower cost per million tokens on the DeepSeek V4 Pro model

1

.

Buck confirmed that Blackwell achieved a 30x improvement in tokens per watt compared to previous generations, stating that "with every generation of GPU, we make sure that our tokens per watt is upwards of 10 times more efficient"

2

. These gains stem from extreme codesign across NVIDIA's technology stack, with continuous software optimization keeping installed hardware productive years after deployment.

The economics favor expanding compute rather than contracting it. Cheaper tokens make more use cases economical, and those use cases consume more tokens than the efficiency saved. This dynamic creates sustained demand even as per-token costs decline.

Durability: Extended Hardware Lifecycles Keep Assets Earning

NVIDIA A100 GPUs shipped in 2020 remain in commercial service six years later, demonstrating continued economic value. CoreWeave recently extended bookings for A100 units first introduced in 2020 through 2029

1

. Every major operator has extended depreciation schedules on servers—a financial indicator of when hardware stops earning—and those schedules keep moving outward.

Barkr estimates useful life at five to six years for eight-GPU H100 systems and nine to 10 years for GB300 NVL72, based on resale values. Silicon Data shows six-year-old A100 GPUs still worth a quarter of their original cost, whereas five-year depreciation schedules had valued them at zero more than a year ago. Ornn Data finds the market paying 80% as much to rent an A100 GPU on a five-year contract as on a one-month contract

1

.

CUDA, the software platform for programming NVIDIA GPUs, runs across generations, ensuring no operator's existing hardware becomes stranded when new architectures arrive. Continuous software and kernel optimization keeps improving what existing hardware can accomplish, extending productive life beyond initial projections.

Fungibility: Running Every AI Workload Across Every Phase and Place

NVIDIA AI Factories run every type of AI model—open and proprietary—across language, vision, biology, physics, and robotics. They handle every phase from data processing through pretraining, post-training, and inference, and deploy in every place from hyperscale clouds and AI clouds to sovereign programs, enterprise data centers, and the edge

1

.

Inference increasingly drives AI factory economics as deployed models process requests and produce tokens. Buck clarified that inference doesn't replace training because organizations continually update deployed models as data and market conditions change. "As companies are using these models, they're refining them, they're aligning them, they're adding more data to them," he explained

2

.

Source: SiliconANGLE

Source: SiliconANGLE

Low-latency inference creates another economic tier for workloads where faster reasoning delivers greater value. NVIDIA's Groq 3 LPX inference accelerator works with Vera Rubin to increase per-user token rates for time-sensitive applications. "If there's value in those tokens to have the fastest possible thinking, LPX can be boosted on top of Vera Rubin to make that possible," Buck noted, citing strong interest from fintech and other real-time sectors

2

.

The same platform supports machine learning, deep learning, generative AI, reasoning, agentic AI, and physical AI. Each new AI workload type arrived on hardware already installed, demonstrating fungibility's role in sustaining revenue generation even as work changes. NVIDIA CUDA-X libraries enable factories to run any accelerated workload, while standardized architecture makes everything accessible to any operator through validated reference designs.

System-Level Design Becomes Competitive Differentiator

As agentic systems draw on multiple models, databases, and tools, entire data centers must function as unified computing systems rather than collections of individual chips. Buck emphasized this transition: "Networking, storage, processors and software must operate together at scale while increasing the output generated from every unit of power"

2

.

Partners like CoreWeave allow customers to choose configurations or use higher-level inference services that optimize the balance between throughput and token speed. "CoreWeave can do that for customers. They don't have to feel overwhelmed by all the choices," Buck said, highlighting how NVIDIA's partner ecosystem simplifies deployment complexity

2

.

The shift from individual component performance to system-level infrastructure efficiency represents a fundamental change in how AI factories compete and maximize return on investment. With power as the ultimate constraint, the ability to extract maximum tokens per watt while maintaining flexibility across workloads determines long-term economic viability.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved