6 Sources
[1]
New Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance and 35x Lower Costs for Agentic AI
Cloud providers including Microsoft, CoreWeave and Oracle Cloud Infrastructure are deploying NVIDIA GB300 NVL72 systems at scale for low-latency and long-context use cases such as agentic coding and coding assistants. The NVIDIA Blackwell platform has been widely adopted by leading inference
[2]
Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell
Baseten, DeepInfra, Fireworks AI and Together AI are reducing cost per token across industries with optimized inference stacks running on the NVIDIA Blackwell platform. A diagnostic insight in healthcare. A character's dialogue in an interactive game. An autonomous resolution from a customer
[3]
AI inference costs dropped up to 10x on Nvidia's Blackwell -- but hardware is only half the equation
Lowering the cost of inference is typically a combination of hardware and software. A new analysis released Thursday by Nvidia details how four leading inference providers are reporting 4x to 10x reductions in cost per token. The dramatic cost reductions were achieved using Nvidia's Blackwell
[4]
NVIDIA Blackwell Ultra delivers 50x higher efficiency for agentic AI
Nvidia's new benchmark data reveals that GB300 NVL72 systems equipped with Blackwell Ultra GPUs achieve up to 50x higher throughput per megawatt and 35x lower cost per token compared to the Hopper platform for low-latency AI workloads. The metrics reflect combined hardware and software advancements
[5]
NVIDIA's Blackwell Ultra Pushes "Agentic AI" Performance to New Heights, Delivering Up to 50× Higher Tokens/Watt & Stronger Long-Context Workloads
NVIDIA's Blackwell Ultra is the modern-day computing option for hyperscalers, and in newer benchmarks, the GB300 NVL72 shows immense performance in low-latency and long context workloads. The AI industry has evolved across multiple layers since its original boom back in 2022, and right now, we are
[6]
NVIDIA Has Managed to Reduce Token Costs by a Whopping 10x With Its Newest Blackwell Platform, Credited to Team Green's "Extreme Codesign" Approach
NVIDIA's Blackwell platform has brought new levels of token optimization to AI inference workloads, as the company reveals a massive milestone in the realm of tokenomics. While NVIDIA has been racing to build new infrastructure in the AI world, one of the company's biggest focuses has been
Share
Copy Link
NVIDIA's GB300 NVL72 systems powered by Blackwell Ultra GPUs achieve up to 50x higher throughput per megawatt and 35x lower cost per token compared to the Hopper platform. Cloud providers including Microsoft, CoreWeave, and Oracle are deploying these systems at scale for agentic AI and coding assistants, while leading inference providers report 4x to 10x cost reductions using open-source models.
NVIDIA has released new performance data showing that its GB300 NVL72 systems equipped with Blackwell Ultra GPUs achieve up to 50x higher throughput per megawatt and 35x lower cost per token compared to the NVIDIA Hopper platform for low-latency workloads
1
. These dramatic efficiency gains target agentic AI applications and AI coding assistants, which have driven explosive growth in software-programming-related AI queries from 11% to approximately 50% last year, according to OpenRouter's State of Inference report1
.
Source: Wccftech
The performance improvements stem from extreme hardware-software codesign that addresses transformer attention layer bottlenecks. Blackwell Ultra Tensor Cores provide 1.5x greater compute performance than standard NVIDIA Blackwell GPUs, while the architecture doubles attention-layer processing through accelerated softmax execution
4
. Cloud providers including Microsoft, CoreWeave, and Oracle Cloud Infrastructure are deploying GB300 NVL72 systems in production for low-latency and long-context workloads such as agentic coding1
.Continuous software optimizations from NVIDIA's TensorRT-LLM, NVIDIA Dynamo, Mooncake, and SGLang teams have significantly boosted Blackwell NVL72 throughput for Mixture-of-Experts (MoE) inference across all latency targets
1
. The TensorRT-LLM library improvements alone have delivered up to 5x better performance on GB200 for low-latency workloads compared with just four months ago1
.Key software optimizations include higher-performance GPU kernels optimized for efficiency and low latency, NVLink Symmetric Memory enabling direct GPU-to-GPU memory access, and programmatic dependent launch that minimizes idle time
1
. SemiAnalysis benchmarks documented that throughput per GPU doubled at certain interactivity levels since October 2025, with NVIDIA stating these developments deliver a 10x increase in tokens per second per user and a 5x improvement in tokens per second per megawatt relative to Hopper4
.Leading inference providers including Baseten, DeepInfra, Fireworks AI, and Together AI are reducing AI inference costs by up to 10x using open-source models on the NVIDIA Blackwell platform
2
. Production deployment data shows significant cost improvements across healthcare, gaming, agentic chat, and customer service as enterprises scale AI from pilot projects to millions of users3
.
Source: VentureBeat
The 4x to 10x cost reductions required combining Blackwell hardware with optimized software stacks and switching from proprietary to open-source models that now match frontier-level intelligence
3
. Hardware improvements alone delivered 2x gains in some deployments, but reaching larger cost reductions required adopting low-precision formats like NVFP4 and moving away from closed-source APIs that charge premium rates3
.Sully.ai cut healthcare AI inference costs by 90%, representing a 10x reduction, while improving response times by 65% for critical workflows like generating medical notes by switching from proprietary models to open-source models running on Baseten's Blackwell-powered platform
2
. The company has returned over 30 million minutes to physicians, time previously lost to data entry and manual tasks2
.Latitude reduced gaming inference costs 4x for its AI Dungeon platform by running large MoE models on DeepInfra's Blackwell deployment
2
. Cost per million tokens dropped from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell, then to 5 cents after adopting Blackwell's native NVFP4 low-precision format2
. Sentient Foundation achieved 25% to 50% better cost efficiency using Fireworks AI's Blackwell-optimized inference stack, processing 5.6 million queries in a single week during its viral launch3
.Related Stories
For long-context workloads with 128,000-token inputs and 8,000-token outputs, such as AI coding assistants reasoning across entire codebases, GB300 NVL72 delivers up to 1.5x lower cost per token compared with GB200 NVL72
1
. Blackwell Ultra's 1.5x higher NVFP4 compute performance and 2x faster attention processing enable agents to efficiently understand entire codebases1
.Chen Goldberg, senior vice president of engineering at CoreWeave, stated: "As inference moves to the center of AI production, long-context performance and token efficiency become critical. Grace Blackwell NVL72 addresses that challenge directly"
1
. CoreWeave was the first AI cloud provider to deploy GB300 NVL72 systems in production4
. Microsoft subsequently deployed what it describes as the world's first large-scale GB300 NVL72 supercomputing cluster, with testing validated by Signal65 recording the cluster achieving over 1.1 million tokens per second on a single rack4
.
Source: NVIDIA
Blackwell Ultra has expanded to a 72-GPU configuration, joining them into a single unified NVLink fabric with 130 TB/s of connectivity
5
. Compared to Hopper, which is confined to an 8-chip NVLink design, NVIDIA has brought superior architecture, rack design, and the NVFP4 precision format, which explains why GB300 dominates in throughput5
."Performance is what drives down the cost of inference," said Dion Harris, senior director of HPC and AI hyperscaler solutions at NVIDIA. "What we're seeing in inference is that throughput literally translates into real dollar value and driving down the cost"
3
. Oracle's OCI platform is deploying GB300 NVL72 systems with plans to scale Superclusters beyond 100,000 Blackwell GPUs to support inference workload demand4
. NVIDIA has previewed its next-generation Rubin platform, projecting a 10x performance improvement over Blackwell4
.Summarized by
Navi
[2]
[3]
30 Jun 2026•Technology

19 Mar 2025•Technology

10 Sept 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
