2 Sources
[1]
Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
AI factories generating the strongest returns are the ones built to earn more, last longer and serve more kinds of work. AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI factory operators will only commit capital on that scale
[2]
Nvidia links AI factory economics to inference and efficiency
Nvidia ties AI factory economics to tokens and power efficiency AI factory economics increasingly depend on more than access to high-performance graphics processing units. As agentic systems draw on multiple models, databases and tools, the entire data center must work as one computing
Share
Copy Link
NVIDIA outlined its AI factory strategy focused on maximizing return on investment through three pillars: productivity, durability, and fungibility. With each megawatt factory costing $60 million, the company's Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt and up to 45x lower cost per million tokens compared to GB300 NVL72.
NVIDIA AI Factories are engineered to maximize return on investment through three interconnected principles: productivity, durability, and fungibility. Each megawatt AI factory costs roughly $60 million, making ROI clarity essential for operators committing capital at this scale
1
. Ian Buck, vice president and general manager of hyperscale and HPC at NVIDIA, emphasized that these assets function as "revenue-generating, fungible, durable, productive parts of an economy" rather than traditional IT costs2
.
Source: NVIDIA
The company's approach addresses a fundamental constraint: power capacity limits how much AI infrastructure a data center can deploy. This makes tokens per watt the central measure of AI factory economics and positions efficient AI inference as a competitive advantage. NVIDIA's system-level infrastructure codesign across the full stack—from models and workloads through software to compute, networking, and memory—enables these efficiency gains.
Power serves as the binding constraint on AI factories, making tokens per second per megawatt the governing metric for earning capacity. More tokens within a fixed power envelope translates directly to higher revenue, while lower cost per million tokens improves margins. According to SemiAnalysis AgentX data, NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than GB300 NVL72, with up to 45x lower cost per million tokens on the DeepSeek V4 Pro model
1
.Buck confirmed that Blackwell achieved a 30x improvement in tokens per watt compared to previous generations, stating that "with every generation of GPU, we make sure that our tokens per watt is upwards of 10 times more efficient"
2
. These gains stem from extreme codesign across NVIDIA's technology stack, with continuous software optimization keeping installed hardware productive years after deployment.The economics favor expanding compute rather than contracting it. Cheaper tokens make more use cases economical, and those use cases consume more tokens than the efficiency saved. This dynamic creates sustained demand even as per-token costs decline.
NVIDIA A100 GPUs shipped in 2020 remain in commercial service six years later, demonstrating continued economic value. CoreWeave recently extended bookings for A100 units first introduced in 2020 through 2029
1
. Every major operator has extended depreciation schedules on servers—a financial indicator of when hardware stops earning—and those schedules keep moving outward.Barkr estimates useful life at five to six years for eight-GPU H100 systems and nine to 10 years for GB300 NVL72, based on resale values. Silicon Data shows six-year-old A100 GPUs still worth a quarter of their original cost, whereas five-year depreciation schedules had valued them at zero more than a year ago. Ornn Data finds the market paying 80% as much to rent an A100 GPU on a five-year contract as on a one-month contract
1
.CUDA, the software platform for programming NVIDIA GPUs, runs across generations, ensuring no operator's existing hardware becomes stranded when new architectures arrive. Continuous software and kernel optimization keeps improving what existing hardware can accomplish, extending productive life beyond initial projections.
Related Stories
NVIDIA AI Factories run every type of AI model—open and proprietary—across language, vision, biology, physics, and robotics. They handle every phase from data processing through pretraining, post-training, and inference, and deploy in every place from hyperscale clouds and AI clouds to sovereign programs, enterprise data centers, and the edge
1
.Inference increasingly drives AI factory economics as deployed models process requests and produce tokens. Buck clarified that inference doesn't replace training because organizations continually update deployed models as data and market conditions change. "As companies are using these models, they're refining them, they're aligning them, they're adding more data to them," he explained
2
.
Source: SiliconANGLE
Low-latency inference creates another economic tier for workloads where faster reasoning delivers greater value. NVIDIA's Groq 3 LPX inference accelerator works with Vera Rubin to increase per-user token rates for time-sensitive applications. "If there's value in those tokens to have the fastest possible thinking, LPX can be boosted on top of Vera Rubin to make that possible," Buck noted, citing strong interest from fintech and other real-time sectors
2
.The same platform supports machine learning, deep learning, generative AI, reasoning, agentic AI, and physical AI. Each new AI workload type arrived on hardware already installed, demonstrating fungibility's role in sustaining revenue generation even as work changes. NVIDIA CUDA-X libraries enable factories to run any accelerated workload, while standardized architecture makes everything accessible to any operator through validated reference designs.
As agentic systems draw on multiple models, databases, and tools, entire data centers must function as unified computing systems rather than collections of individual chips. Buck emphasized this transition: "Networking, storage, processors and software must operate together at scale while increasing the output generated from every unit of power"
2
.Partners like CoreWeave allow customers to choose configurations or use higher-level inference services that optimize the balance between throughput and token speed. "CoreWeave can do that for customers. They don't have to feel overwhelmed by all the choices," Buck said, highlighting how NVIDIA's partner ecosystem simplifies deployment complexity
2
.The shift from individual component performance to system-level infrastructure efficiency represents a fundamental change in how AI factories compete and maximize return on investment. With power as the ultimate constraint, the ability to extract maximum tokens per watt while maintaining flexibility across workloads determines long-term economic viability.
Summarized by
Navi
[2]
03 Jan 2026•Technology

16 Sept 2026•Technology

25 Aug 2026•Technology

1
Policy and Regulation

2
Technology

3
Technology
