11 Sources
[1]
Nvidia details Rubin architectural optimizations for inference - improvements target better performance and efficiency from the GPU to the rack
Nvidia's upcoming Vera Rubin platform, set to arrive later this year, will take the stage as the AI world shifts towards an era dominated not by frontier training runs but by the demands of agentic AI inference at massive scale. The hunger for generated tokens in agentic workflows and the demands of delivering them quickly, efficiently, and at low unit cost now dominate the discussion. We've already gone in depth on new performance data around the Vera CPU and how it helps to accelerate agentic AI workloads, but that's not all Nvidia is sharing today. It's also detailing some new features of the Rubin architecture and how those features are meant to increase inference efficiency from the GPU level to rack-scale and data-center-scale implementations of this accelerator platform. The full Vera Rubin NVL72 rack-scale system is built up from 36 Vera CPUs and 72 Rubin GPUs, but our focus today is on the GPU proper. Rubin joins two compute dies onto a single package using the Nvidia High Bandwidth Interface. The resulting chip offers 224 Streaming Multiprocessors (SMs) containing a total of 896 Tensor Cores alongside 288GB of HBM4 memory providing 22 TB/s of memory bandwidth. As an inference-focused accelerator, Nvidia touts Rubin's 50 sparse PFLOPS of NVFP4 inference throughput as its headline performance figure, although that's only one of a dizzying array of data types this chip can handle. Here are some key rates to keep in mind for this chip so far: Let's dive into some of Rubin's refinements for inference workloads to understand how Nvidia aims to keep all of those resources fully utilized. The Rubin Tensor Memory Accelerator efficiently manages growing MoE models First up, Nvidia highlights efficiency improvements in the Tensor Memory Accelerator (TMA) that help feed the Tensor Cores with data. The TMA is a dedicated engine built to handle memory address calculations and perform direct loads of array data into a GPU's shared local memory. Leading AI model architectures have moved from dense models where every parameter is activated per output token to a mixture-of-experts (MoE) architecture where only certain specialized sub-networks are activated per token, based on the guidance of a router that helps judge which experts are best suited to processing a given input. MoE expert weights can be distributed across GPUs in order to efficiently utilize limited per-GPU HBM capacity. Nvidia says that Rubin's TMA has been improved to deal with the challenges of managing the growing numbers of experts in today's leading models. The TMA in Blackwell GPUs needed to maintain separate MoE descriptors in memory for the location of every expert, meaning that the overhead of locating and moving those expert weights requires more compute resources as the number of experts grows. The Rubin TMA now supports GPU kernels that maintain and update a single unified MoE descriptor directly in the TMA instruction at runtime, reducing computation of MoE descriptor metadata and requiring less calculation overhead for data movement. This approach frees up GPU cycles for inference calculations, which is, of course, the place that you want your expensive AI accelerator spending the vast majority of its time. Doubled K-dimension throughput, double the Tensor Core output Rubin also improves the fundamental performance of matrix operations in the Tensor Core by doubling the amount of work those cores can perform on the K dimension, or the shared inner dimension of a pair of matrices to be multiplied. Without going too deep into the math, the size of the K dimension is directly related to the number of times the Tensor Core has to loop over the elements of the two matrices being multiplied. In Nvidia's example, then, the calculation of a result matrix that would require four loop iterations on Blackwell can be completed in only two on Rubin. Nvidia says this improvement has wide-ranging benefits for throughput-, memory-, and latency-bound kernels, and it's helpful for both context processing and decode phases of inference. Softmax on Rubin gets up to a 4X boost versus Blackwell Rubin also focuses on improving the performance of the attention mechanism that's foundational to transformer-based LLMs More advanced models now support context lengths of up to a million tokens, and quickly performing attention calculations on such long input sequences quickly is a key driver for improved inference performance. Softmax is an essential operation in attention calculations, and in order to keep up with the improved Tensor Core throughput in Rubin, Nvidia has once again boosted softmax throughput in the GPU SM's Special Function Unit (SFU). Since it relies on the transcendental math capabilities of the SFU, softmax throughput can become a bottleneck for subsequent inference work, and it's a limitation that Nvidia already sought to address with enhancements to the Blackwell Ultra SFU. Blackwell Ultra doubled FP32 and BF16/FP16 exponential throughput compared to the first-gen Blackwell GB200. Rubin maintains Blackwell Ultra's 2X speedup over Blackwell in FP32 exponential math, and it doubles BF16/FP16 exponential calculations again compared to Blackwell Ultra, leading to a 4X improvement in throughput compared to Blackwell for those lower-precision data types. Finer-grained dependency management, better Tensor Core occupancy Rubin also increases Tensor Core occupancy by providing finer-grained opportunities for coordination between dependent kernels than on Blackwell. One case that Nvidia cites where these dependencies arise is the generation of activations for an LLM, where one kernel produces and stores data that is then used by a subsequent kernel as a prompt proceeds through a neural network. On Blackwell GPUs, a long-running producer kernel on one thread block (perhaps within a CUDA structure like a cluster) might delay the execution of a subsequent consumer kernel on those thread blocks, even as other thread blocks of the producer kernel have finished their work. Rubin offers finer-grained dependency resolution between kernels, such that a consumer kernel can begin executing on individual thread blocks as soon as the producer kernel's output from each thread block becomes available, instead of waiting for the entire batch of producer kernel data to become available. This finer-grained management results in better GPU utilization, lower kernel-to-kernel latency, and ultimately increases tokens per second per user. More efficient inter-GPU communication, lower NVLink overhead All of the improvements we've discussed so far relate to how work happens on one GPU, but the Vera Rubin NVL72 rack-scale accelerator comprises many GPUs connected over an NVLink fabric within the rack. Model weights, key-value cache data, and inter-GPU synchronization messages all move over this fabric, so keeping overhead and latency low is key to realizing maximum performance. GPUs running CUDA kernels can directly initiate communication with other GPUs in the rack using Nvidia Collective Communications Library (NCCL) API, lowering overhead. Nvidia notes that because the GPU performs those operations directly as part of the compute kernel, the efficient execution of those communications becomes critical to performance. On a Blackwell system, an NVLink transfer between GPUs might require data store operations followed by a memory barrier and an atomic flag. The Rubin architecture introduces a feature called counted writes that reduces the amount of coordination and synchronization traffic necessary to share data between GPUs across the fabric. On Rubin, the memory barrier and atomic operations are replaced by a single write counter update on the receiving GPU, reducing network traffic and latency and improving compute utilization by reducing the time spent waiting for coordination overhead. All told, in tandem with the high single-threaded performance of the Vera CPU for agent harnesses, tool calling, code compilation, and more, the improvements in the Rubin GPU for performance on critical inference operations, as well as improved efficiency for data movement on-chip and across the rack, promise to help create a rack-scale and data-center-scale system that will both increase inference performance and lower per-token inference costs in the increasingly agentic future that Nvidia envisions. We're excited to see more of what this GPU can do as deliveries of Vera Rubin systems are set to begin this fall. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[2]
Nvidia shows off Vera Rubin platform for tokenmaxxing
Last week, Nvidia invited a handful of journalists to a briefing about AI infrastructure to help convey the physical reality behind simulated intelligence. "Infrastructure is physical, it's real," said Ian Buck, general manager of Nvidia's hyperscale and HPC computing business. "You can touch it, you can see it, and it's what helps bring AI to life. So that is part of the goal here." The briefing included a tour of a working Nvidia lab - a mini-datacenter - nestled in a residential district of Sunnyvale, a Silicon Valley suburb. Attendees were asked not to reveal the location - which is one of four such facilities - though they're easy enough to discover with a bit of online sleuthing. AI infrastructure is indeed real - or at least realized as revenue when equipment is booked under a bill-and-hold arrangement - and you can touch it if your job involves handling tech gear. It's also controversial as the AI tsunami crests over a society that's uncertain if the technology will bring bounty or harm. Just one day after the press event, Dutch activists peppered a data center serving Microsoft workloads with chemical-filled balloons to protest climate and political concerns. Nvidia presumably would prefer not to have neighbors grousing about power and water consumption at its facilities, though such concerns may be muted in an area so in thrall to the tech sector. In any event, Nvidia's latest data center hardware should require significantly less human intervention than its older kit. During the lab tour, Andrew Bell, a senior VP of hardware engineering, showed off how the company's Vera Rubin NVL72 compute tray is vastly faster to install than prior hardware. "With automation assembly, the compute tray inside the Vera Rubin NVL72 can be assembled in one minute, compared to GB200 compute tray assembly that typically takes 90 minutes. This represents a 90x improvement," he said. The Nvidia event focused on how the company's Vera Rubin platform, announced at Computex 2024 and detailed at CES in January, advances its vision of AI Factories, a type of computing infrastructure dedicated to running AI workloads. The platform consists of six chips: the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 network interface, BlueField-4 data processing unit (DPU), and Spectrum-6 Ethernet switch. Nvidia claims Vera Rubin is said to deliver 10x more tokens per watt than GB200 NVL72 in initial CoreWeave results on DeepSeek-R1. These appear in five rack-mountable systems designed for data centers: the Vera Rubin NVL72 compute tray, Vera Rubin NVL72 NVLink switch tray, the Groq 3 LPX inference accelerator tray, the Vera CPU tray, the BlueField-4 STX storage tray, and the Spectrum-6 SPX switch tray. Nvidia's pitch for all this hardware is that agentic workloads - running AI agents - require a different approach. AI agents work best, the company claims, with sustained inference and low latency across multiple reasoning steps, high decoding throughput, efficient reasoning over long contexts, large key-value cache capacity, and the ability to scale models across linked GPU domains. The company's AI Factory calls for unified compute capability that spans the data center. It's an architectural gambit based on the business of selling tokens. And while GPUs are the star of the show for training and inference, CPUs play a critical role in data processing, networking, scheduling, and storage. Vera represents a bet that making the core faster is better for agentic workloads than increasing the number of cores. Nvidia claims that the Vera CPU Olympus Core accelerates agents 2x, offers 3x core-to-core bandwidth, and delivers 40 percent lower latency via LPDDR5X memory. And to the extent Nvidia's data center customers can move more tokens for their AI inference operations, they should be able to capture more revenue. "Every AI factory is power-constrained," said Buck. "And frankly the most important metric is your delivered performance in a fixed watt data center, your perf-per watt. All that comes together in your AI factory revenue. You may rent infrastructure ... for dollars per hour but you generate revenue with profitable tokens now, in the hundreds of dollars per hour. Your AI Factory revenue is a function of how many tokens you can generate in a fixed watt data center." It's also a function of the market price of tokens. As the supply of tokens increases - through new hardware that makes token generation more efficient and competing AI providers - there's downward pressure on the cost of tokens. The hope at Nvidia, and more acutely at the frontier AI labs, is that demand for tokens will increase as prices decline, at a rate that leaves enough of a profit margin to recoup the capital expenditures funding the current datacenter construction boom. But given reports of companies capping employee token budgets, it's not clear that every organization is seeing the productivity increases touted by Nvidia and its peers. Buck offered a back-of-the-envelope calculation to suggest that AI makes developers more productive, but it was more of a thought experiment than a verifiable claim. Pointing to the surge in GitHub code commits - which haven't been great for GitHub's stability - he posited there are about 30-40 million software developers active in the world, which translates to about $3 trillion in salaries. "They're now producing three times the output or effectively nine trillion dollars of productivity," Buck said. At Nvidia, he added, "the number of [code] check-ins and developer productivity has tripled as a result" of AI coding tools. Buck said that Nvidia uses AI agents for pretty much all of its software development processes. "We use agents also for our chip development," he added. "Our entire chip bug database is all watched and reviewed by agents." The leading AI company uses AI to make its AI software and hardware. Flog AI enough and you can bring it to life. ®
[3]
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
Backed by 300 global partners, Vera Rubin is ramping up worldwide. NVIDIA partners CoreWeave, Google Cloud, Microsoft Azure and Mistral are among many deploying Vera Rubin, which delivers benchmark leadership on performance per watt and lowest token costs. NVIDIA Vera Rubin is here, and it's going gigascale. Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350-plus factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand. The Vera Rubin platform is built from chip to grid to deliver the highest performance per watt and the lowest token cost. CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72 -- landing directly on the metric that matters most for power-constrained AI factories. Advancing Performance With Extreme Co-Design What makes this possible is extreme co-design across seven chips and five rack trays -- Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX and Vera BlueField-4 STX -- all engineered as a single system rather than assembled from separate off-the-shelf products. The NVIDIA Vera CPU is at its center. It redefines what an AI factory CPU can be. Designed and built for the agent era, its custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs, making it the most efficient single-threaded CPU for the agentic workloads that matter most. Accelerating AI Factories With Purpose-built Networking For networking, the platform's sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency and 10x higher packet rates than off-the-shelf Ethernet. For scale-out, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems, 1.6T ConnectX-9 SuperNICs, adaptive routing, advanced congestion control, telemetry, and open operating system support, enabling 1.6x higher RDMA bandwidth than off-the-shelf Ethernet. The world's leading AI infrastructure builders -- including CoreWeave, Microsoft, SpaceXAI and Tesla -- are among the first to bring in Spectrum-6 switches to accelerate their AI factories. NVIDIA Photonics with co-packaged optics for scale-out -- the industry's first such switch in volume manufacturing -- adds 5x lower power and 10x higher MTBI versus pluggable transceivers, with CoreWeave, Lambda and OCI among the first adopters. Spectrum-XGS Ethernet extends performance across sites with 1.9x multi-site throughput because gigascale AI isn't a single building problem. And NVLink Fusion opens the NVIDIA infrastructure platform to third-party XPUs, giving partners a faster path to market on the proven NVLink scale-up stack and ecosystem. Saving Setup Time, Water NVIDIA's three generations of rack-scale co-design produced a Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute. A 45-degree Celsius liquid cooling inlet temperature design enables chiller-free dry-cooler operation. For new AI factories, this higher temperature dry cooling along with the closed-loop liquid cooling system saves millions of gallons of water per megawatt annually. Tuesday, July 21, 8:00 a.m. PT 🔗 NVIDIA Vera Rubin Powers Europe's Open-Model Era Vera Rubin is delivering next-generation performance to Europe's AI infrastructure. It's the foundation for a newly expanded Microsoft and Mistral partnership that brings frontier AI to the region, combining open European models with cloud and customer-controlled environments so governments and regulated industries can adopt it on their own terms. Underpinning the partnership is a new multibillion-dollar agreement focused on expanding AI infrastructure in Europe. Mistral is adding its GPU capacity, drawing on thousands of the latest NVIDIA Vera Rubin GPUs to increase AI compute availability for customers and provide a shared platform for training, inference and large-scale deployment. NVIDIA Vera Rubin is ramping into full production.The rack-scale AI supercomputer unifies seven new chips co-designed as one system, and it will power the next generation of Mistral Compute and Microsoft's European AI infrastructure, drawing on tens of thousands of GPUs. Europe wants the world's most capable AI, running under its own laws, close to home, and fully within its control. And the bar keeps rising: agentic systems can consume up to 15x more tokens than traditional AI applications, making efficient infrastructure a strategic priority. Meeting that expectation takes more than computing capacity: AI that scales efficiently while satisfying regional requirements for data control, governance, resilience and strategic autonomy. Vera Rubin provides the computing foundation for Europe's open-model ecosystem. Its full-stack architecture combines accelerated computing, networking and software so models and agents run efficiently from training through production. Open Models Built for Enterprise AI Together Microsoft, Mistral and NVIDIA are delivering sovereign-ready AI across public cloud, cloud-connected and fully disconnected private cloud environments -- pairing the flexibility of open models with the infrastructure required to operate them at scale. Mistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, and Mistral models are integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local, customers can use the same models, tools and operating patterns across cloud and customer-controlled environments AI on Europe's terms Government agencies can apply AI to sensitive workflows. Healthcare and financial services organizations can deploy within regional requirements. Manufacturers can process data on site for faster decisions. Compared with NVIDIA GB200 NVL72, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens, providing more intelligence within the same power footprint. Sovereign AI should not force organizations to choose among innovation, economics and control. With Vera Rubin as the computing foundation -- and Microsoft and Mistral delivering sovereign cloud infrastructure and open European models -- Europe can pursue all three. Read the Microsoft and Mistral press release for more details. Tuesday, July 21, 8:00 a.m. PT 🔗 NVIDIA Vera Rubin NVL72 on CoreWeave Demonstrates 10x More Tokens Per Megawatt Than Blackwell in Benchmark Embracing extreme co-design with NVIDIA, Coreweave is delivering an order of magnitude performance leap on Vera Rubin NVL72. Close collaboration with partners like CoreWeave is mission-critical to bringing up a new generation of NVIDIA accelerated computing into AI factories. After months of co-engineering work, CoreWeave became the first AI cloud to bring up and validate Vera Rubin NVL72 -- and it is now sharing the first measured performance numbers from live hardware. CoreWeave ran a DeepSeek-R1 benchmark on Vera Rubin NVL72 and saw 10x improvement in tokens per second per megawatt compared with Grace Blackwell NVL72. Tokens per megawatt is the metric that determines whether AI infrastructure can profitably scale. More tokens per megawatt means more intelligence from the same power budget, or the same workload on significantly less power. AI labs and enterprises such as Jane Street will use the Vera Rubin platform to scale its AI factories on the CoreWeave cloud. Beating Bottlenecks With NVIDIA Spectrum-X DeepSeek R1's MoE architecture makes all-to-all GPU communication a critical requirement at scale: Each token must be routed across distributed expert sub-networks. Vera Rubin NVL72's 260 TB/s all-to-all NVLink 6 fabric removes that constraint, enabling the rack to behave as a single unified accelerator. CoreWeave is among the first to deploy the NVIDIA Spectrum-X Ethernet SN6600-LD as the switching fabric for Vera Rubin NVL72. Built on the 102.4 Tb/s Spectrum-6 switch chip and featuring a liquid-cooled design, CoreWeave deploys dense switching racks, delivering 1.64 Pb/s per rack with 100% more capacity than previous generation air-cooled switches. CoreWeave also provides a fully non-blocking, multi-plane, multi-rail spine and leaf fabric connecting Vera Rubin NVL72 GPUs without oversubscription. Tuesday, July 21, 8:00 a.m. PT 🔗 NVIDIA Vera Rubin NVL72 Drives New Google Cloud A5X Instance for Ineffable Intelligence NVIDIA Vera Rubin NVL72 is powering Google Cloud's first A5X instance, now up and running for London startup Ineffable Intelligence. Ineffable Intelligence develops a new generation of intelligent "superlearner" systems that continuously learn through experience to discover new breakthroughs across all fields. Ineffable Intelligence's agents learn directly from interaction with their environments, rather than from static datasets. Instead of using large language models, Ineffable is developing "superlearner" systems through reinforcement learning, generating experience across continuously simulated, massively parallel environments and rapidly translating that experience into policy updates and evaluation. These tightly coupled learning loops place exceptional demands on compute, memory bandwidth and interconnect, requiring infrastructure that can operate at enormous scale with extremely low latency. "The next era of research requires the next era of hardware," said Lasse Espeholt, co-founder of Ineffable Intelligence. "We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners." NVIDIA Vera Rubin NVL72 is designed for this kind of agentic training, delivering predictable latency, high utilization and significantly more intelligence per dollar than previous generation systems, making it a natural platform choice for large-scale reinforcement learning. Google Cloud A5X instances, announced at Google Cloud Next, are bare-metal instances built on NVIDIA Vera Rubin NVL72 rack-scale systems, delivering up to 10x lower inference cost per token and 10x higher token throughput per megawatt than the prior generation. A5X uses NVIDIA ConnectX‑9 SuperNICs combined with next-generation Google Virgo networking, enabling clusters that can scale to tens of thousands of NVIDIA Rubin GPUs within a single site and up to nearly a million GPUs across multisite configurations, giving customers a unified, AI‑optimized stack for training, tuning and serving frontier, open, agentic and physical AI models while optimizing for performance, cost and sustainability. This infrastructure is designed to help unlock the next generation of reinforcement learning systems for breakthroughs in superlearning and superintelligence. Tuesday, July 21, 8:00 a.m. PT 🔗 NVIDIA Vera CPU Doubles Orchestration Speed for DeepInfra AI Cloud Benchmark results from DeepInfra show that the NVIDIA Vera CPU is more than twice as fast and can support more concurrent AI agents compared with other CPUs. Cloud platform DeepInfra, an early access participant in the NVIDIA open AI ecosystem, independently designed and ran benchmarks using its production AI agent infrastructure. DeepInfra processes nearly five trillion tokens a week, with about 30% driven by agentic systems. Its cloud platform is built for high-throughput AI inference. The benchmarks demonstrate support for up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs, while improving infrastructure utilization and cost efficiency. These results show that the NVIDIA Vera CPU delivers the cost efficiency, low latency and throughput that production agentic AI demands. As AI agents take on more complex reasoning, planning, tool use and data movement, CPU performance has become increasingly important for orchestrating work around each model call. Part of NVIDIA's extreme co-design approach to AI factories, the Vera CPU is built for agentic workloads. DeepInfra's benchmark highlights how the NVIDIA Vera CPU helps cloud providers improve infrastructure utilization, increase cost efficiency and support more concurrent AI agents at the same quality of service.
[4]
Nvidia says Vera Rubin is in full production, with OpenAI set to deploy at scale in Q3
Nvidia confirmed Vera Rubin is in full production, with OpenAI deploying at scale in Q3 and CoreWeave reporting ten times the token output. Nvidia confirmed on Monday that its Vera Rubin platform has reached full production, with Ian Buck, the company's vice president of accelerated computing, telling reporters at Nvidia headquarters that systems are now shipping to customers including OpenAI, CoreWeave, Google Cloud, Microsoft Azure, Meta, and Dell. OpenAI plans to adopt Vera Rubin at scale during the third quarter, according to Bloomberg, which first reported the briefing. CoreWeave, one of the first cloud providers to receive the hardware, told Bloomberg that its NVL72 racks are delivering ten times the token output of the previous generation. The NVL72 is a full-rack system that pairs 72 Rubin GPUs with Vera CPUs and uses liquid cooling to eliminate internal cabling, a design change Nvidia demonstrated at the headquarters event. Buck said the cooling approach allows the company to remove copper connections that previously limited how tightly components could be packed together. The Vera CPU is the piece Nvidia built to replace the processors it previously bought from others, and the company used the briefing to draw direct comparisons with AMD. Nvidia claimed the Vera CPU is nearly twice as fast as AMD's Turin chip on Python workloads, a benchmark chosen because Python dominates the software stack that runs most AI inference. Anthropic and OpenAI are among the first labs to receive the processor, alongside Perplexity, SpaceX, and Oracle. The production milestone comes at an awkward moment for Nvidia's stock. The company's shares have risen nine percent this year, while the broader chip index has climbed 66 percent over the same period, according to Bloomberg. Intel, ARM, and AMD have all more than doubled. Analysts project Nvidia's revenue will increase 82 percent to roughly $393 billion for the fiscal year, but that growth rate has not translated into the kind of share price momentum investors saw during the Blackwell cycle. Nvidia has steadily built out the Vera Rubin story over the past two months. Jensen Huang declared the platform in full production at Computex in early June and named Anthropic, OpenAI, SpaceX, and Oracle as early recipients. Monday's briefing added the performance numbers and customer testimonials that the keynote stage did not provide, turning a claim into a set of measurable benchmarks. The numbers Nvidia chose to highlight are its own and its customers' rather than independent tests, a distinction worth noting. CoreWeave's tenfold token figure and the Python benchmark against AMD came from the company's own briefing, not from a third party. Volume shipments across all named customers, and independent verification of the performance claims, are still ahead.
[5]
Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains
Artificial intelligence chip king Nvidia Corp. today revealed a fresh trove of performance benchmarks and architectural milestones for its next-generation Vera Rubin platform as it edges closer to global availability. The new numbers are impressive, but on a higher level they also underscore the potency of Nvidia's approach that combines custom silicon with its broader hardware ecosystem and aggressive software optimization to squeeze even greater efficiency gains for AI workloads. The latest milestones showcase the effort the chipmaker has made toward fine-tuning its ecosystem to tackle the increased performance demands and exploding infrastructure costs of next-generation "agentic" AI workloads. Impressive gains One of the key revelations came from Nvidia's partner CoreWeave Inc., the neocloud company that provides rented, cloud-based access to its graphics processing unit platforms. In early production runs, CoreWeave showed that it was able to squeeze out an incredible 10 times more tokens per watt when running DeepSeek Ltd.'s R1 model on the new Vera Rubin platform, compared with the previous generation GB200 NVL72 system based on Blackwell. It's an impressive leap, but CoreWeave showed that the optimizations made to Vera Rubin can also be applied to the GB200 NVL72 system as well, improving throughput per megawatt by more than four times over a three-month span. Nvidia said these gains were validated across 250,000-plus distinct configurations and more than 1.4 million GPU hours of testing. Nvidia also wheeled out some new numbers for its Vera central processing units for running autonomous AI agents that can perform work on behalf of humans with minimal supervision. The Vera CPUs were built on a custom microarchitecture specification code-named Olympus core, and are highly optimized to process the irregular control flows of AI agents. In the latest benchmarks, Nvidia demonstrated 1.9 times faster agentic performance and a six-fold improvement in latency over x86-based alternatives. The company also showed that Vera surpassed Advanced Micro Devices Inc.'s flagship EPYC Turin CPU by almost 100% on selected industry benchmarks, underscoring the importance of custom silicon in eliminating CPU-side bottlenecks. Infra optimizations There's no doubt that Nvidia's processors are among the best in the business, but the company made it clear that these gains could only be achieved thanks to its strategy that's focused on hardware and software co-design. The company also develops the required software and networking systems needed to optimize the performance of its chips so as to maximize power and cooling efficiency. Nvidia revealed that by dynamically optimizing the full infrastructure and energy stack, it could deploy 40% more GPUs within the same power envelope. At the same time, it showed how it can also minimize the environmental impact of its new processors by cooling them in a 45°C closed-loop liquid-cooled system that saves about 4 million gallons of water per megawatt annually over standard cooling methods. The company also talked about its latest generation of networking fabrics, showing how they continue to outperform the best generic Ethernet architectures. The sixth-generation NVLink 6 interconnect delivered 2.3 times higher simulated decode throughput for massive large language models than Ethernet-based networks. Meanwhile, the company's Spectrum-X platform enabled 1.6 times faster remote direct memory access bandwidth with 1.7 times fewer switches. This integrated network results in five times greater optical power efficiency and 10 times more reliability, Nvidia said. In particular, the company revealed, its latest Spectrum-X platform, Spectrum-6, is arriving at AI factories from the likes of CoreWeave, Microsoft Corp., Nebius B.V., SpaceXAI Corp. and Tesla Inc. Nvidia is currently racing to ship the new Vera Rubin systems to customers and partners, including the likes of Google Cloud, Microsoft Azure, Meta Platforms Inc., Oracle Cloud Infrastructure, Dell Technologies Inc., OpenAI Group PBC and CoreWeave. As those platforms inch closer to general availability, the release of the latest benchmarks is clearly a strategic move by Nvidia, aiming to show that it now has all of the pieces in the puzzle for its customers to build the "AI factories" of the future. With reporting from Robert Hof
[6]
First NVIDIA Vera Rubin NVL72 benchmarks show 10X improvement over Grace Blackwell
"Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure," NVIDIA confirms as its rack-scale Vera Rubin platform rolls out to multiple sites in 30 countries. And when it comes to performance, we've got CoreWeave's first benchmark on DeepSeek-R1. And yes, a 10X improvement over Grace Blackwell NVL72 definitely falls into the "order-of-magnitude performance leap" that NVIDIA attributes to it, with this number referring to tokens per second per megawatt. As an efficiency and power-budget-based metric, it directly correlates to how effectively AI infrastructure built with Vera Rubin scales when dealing with the same power budget and workload as Grace Blackwell systems. This 10X figure comes from CoreWeave running the same DeepSeek-R1 benchmark on both Vera Rubin NVL72 and Grace Blackwell NVL72. A key part of the 10X improvement and impressive evolution is how Vera Rubin leverages NVIDIA Spectrum-X to deal with bandwidth bottlenecks. We're talking about a jaw-dropping 1.64 Pb/s per rack. "CoreWeave is among the first to deploy the NVIDIA Spectrum-X Ethernet SN6600-LD as the switching fabric for Vera Rubin NVL72," NVIDIA confirms. "Built on the 102.4 Tb/s Spectrum-6 switch chip and featuring a liquid-cooled design, CoreWeave deploys dense switching racks, delivering 1.64 Pb/s per rack with 100% more capacity than previous-generation air-cooled switches." With AI and data center power use becoming a point of contention in surrounding communities, improved efficiency is quickly becoming one of the main focuses when talking about large rack-scale systems built on platforms like Vera Rubin. In addition to its power-saving capabilities, the Vera Rubin NVL72 rack is also designed to save "millions of gallons of water per megawatt," thanks to its cable-, fan-, and hose-free trays and liquid-cooling inlet that enables dry-cooler operation.
[7]
NVIDIA Vera Rubin NVL72 Enters The Stage With A Monstrous 10x Uplift In Token Throughput Versus Blackwell, Achieves 800,000 Tokens/s Vs GB200's 80,000 at The Same 150MW
NVIDIA's Vera Rubin is finally here, delivering unimaginable amounts of AI performance versus Blackwell as the first systems power on. NVIDIA Vera Rubin Now Deployed & Pumps Out 800,000 Tokens/s In DeepSeek-R1, Crushing Blackwell At The Same Wattage Well, it's here: the most powerful AI platform that the world has seen so far, Vera Rubin. Part of NVIDIA's Extreme Co-Design initiative, which aims to deliver a full-fledged stack of AI-ready hardware and software, enabling unprecedented levels of performance per watt & delivering the lowest token cost, Vera Rubin ushers in a new era of AI. The NVIDIA Vera Rubin platform is a diverse set of hardware and software solutions, which are the realization of the five-layer cake framework forming a comprehensive AI ecosystem. The stack houses a total of six trays, with some of the world's most advanced chips ever made: * NVIDIA Vera Rubin NVL72 Compute Tray (Rubin + Vera + Grace Bluefield + ConnectX-9) * NVIDIA Vera Rubin NVL72 NVLink Switch Tray (NVLink6 Switch) * NVIDIA Vera CPU Tray (Vera) * NVIDIA Spectrum-6 SPX Switch Tray (Spectrum-6 Switch) * NVIDIA Groq 3 LPX Tray (Groq 3) * NVIDIA Vera Bluefield-4 STX Storage Tray (Vera Bluefield) NVIDIA's rack-scale solution process has also seen continuous refinements. Just like how the Raptor engines got refined by SpaceX, resulting in the Raptor 3, which is a sleeker and neater version with more advanced capabilities, NVIDIA refined its Vera Rubin trays, eliminating the use of cables, fans, and housings, in favor of a more streamlined and liquid-cooled system that fits into the rack seamlessly with a modular approach. This shows how the engineering efforts went beyond just chips and code. Now, NVIDIA's first Vera Rubin NVL72s are powering on, and the first resultsa re coming in. Last month, we reported how Coreweave and other major cloud providers were installing & validating NVIDIA's Vera Rubin platforms; well today, we get to see how they perform too. In the first DeepSeek R1 tests shared by Coreweave, the company has managed to sustain a 10x higher token throughput per watt versus Grace Blackwell. An NVIDIA GB200 NVL72 Grace Blackwell system managed just 80,000 Tokens/s at 150MW. The NVIDIA VR200 NVL72 Vera Rubin platform offers 800,000 tokens/s at the same wattage, showcasing a massive increase in MoE workloads. And this is only the beginning; NVIDIA's Vera is all set to disrupt the market when more systems come online and showcase their potential. Backed by a strong suite of networking solutions that will enable companies to achieve Gigascale levels of AI, NVIDIA is once again showcasing why it is the leader in AI. Follow Wccftech on Google to get more of our news coverage in your feeds.
[8]
NVIDIA Rubin GPUs Bring 10x Increase in Agentic AI Performance Versus Blackwell as Its Architecture Gets Fully Unpacked, Featuring 336 billion Transistors
NVIDIA has fully disclosed its Rubin AI GPU architecture, bringing a monumental increase over Blackwell for the Agentic era. NVIDIA Reimagines Data Center With Its Rubin GPUs, Bringing The Fastest AI Performance In The World The Rubin GPU is at the heart of NVIDIA's next-generation data center platforms; it is a chip that took several years to make and is now finally entering general availability. With the first Rubin AI platforms already landing across the globe at major tech firms, it was high time that NVIDIA disclosed the full details of the architecture that went into making this supercharged AI chip that will shape the Agentic landscape. According to NVIDIA, the Rubin chip delivers 10x higher agentic throughput per unit of energy versus Blackwell, and it does so with three major architectural changes: * Enhanced Tensor Cores with Expanded Precision Capabilities * New HBM4 Memory Subsystem * 3rd Generation Transformer Engine But first, let's start with the architectural breakdown by looking inside the chip. Rubin Architecture Dissected - A Multi-Billion Transistor Monster NVIDIA's Rubin chip is made on TSMC's 3nm (N3P) process node and features two reticle-limited dies to achieve high-density and efficiency characteristics that define the chip. These two dies are connected in a single unified package through NVIDIA's High-Bandwidth Interface (NV-HBI). The chip packs a total of 336 billion transistors, a 62% increase over the Blackwell GB300 chip. Under the hood of each Rubin GPU are 8 GPCs or Graphics Processing Clusters with a total of 224 SMs or Streaming Multiprocessors and 896 Tensor Cores. That should round up to 28 SMs & 112 Tensor Cores per GPC, but based on the block diagram, there are two types of GPCs, one with 30 SMs, and the other with 26 SMs. There are four GPCs per die, which are connected to two blocks of large decentralized L2 cache, a Gigathread Engine, four HBM channels (4096-bit per die), and an NVLink interface. The chip also houses a PCIe Gen6 controller to interface with the host CPU, and the NVLink v6 IO provides 3600 GB/s linked with the NVLink switch. The MIG controller partitions the GPU for multiple workloads & there's also NV-DEC for accelerated decoding. The chip also features Confidential Computing through TEE-I/O, which secures data at rest, in transit, and in use across AI workflows. The Fastest HBM4 on The Market For Faster Data Movement On the memory front, NVIDIA Rubin makes use of HBM4, which is the latest HBM standard. The need for faster data movement has become critical in today's AI landscape, and Rubin goes all out by bringing 288 GB of capacity across eight 12-Hi stacks (that's 36 GB per stack &3 GB per DRAM). These stacks deliver a peak bandwidth of 22 TB/s, making them the fastest HBM4 solution on the market. NVIDIA also utilizes its new and enhanced Tensor Memory Accelerator (TMA) for moving highly complex data efficiently through the chip, while NVLink 6 provides 3600 GB/s of scale-up speeds for all-to-all GPU-GPU communication. The NVLink-C2C interface provides 1800 GB/s of CPU-to-GPU communication, while the PCIe Gen6 x16 lanes provide 256 GB/s bandwidth for host CPU connectivity. HBM4 Boosts Token Generation The HBM4 memory for Rubin has been designed for maximum power and compute efficiency, leading to faster decoding and token generation capabilities. Most modern workloads spend time doing long contexts, large KV caches, and interactive token generation, which make bandwidth even more critical. With a 2.8x increase over Blackwell's memory solution (22 TB/s vs 8 TB/s), the HBM4 solution provides: * Capacity: Supports model residency, larger context windows, larger KV caches, and higher concurrency without unnecessary KV-cache offload. * Bandwidth: Supports the token-by-token generation phase, during which model weights and KV state must be moved rapidly enough to keep compute engines productive. * Memory subsystem: TMA and memory-locality strategies help software use the memory subsystem efficiently, supporting high achieved bandwidth for complex data layouts. HBM4 amplifies Rubin, which is essential for supporting multi-trillion-parameter models, extending context length without the hassle of KV cache offloading, and supporting much higher concurrency rates as well as long-horizon inference workloads. Accelerating MoE & Tokens Through Enhanced Tensor Memory Accelerator Rubin also comes with an enhanced Tensor Memory Accelerator, or TMA, which efficiently locates and moves expert weights as the number of experts grows within MoE models. TMA reduces the data movement overhead through improved descriptor handling that allows software to route and work more efficiently. The Inline Descriptor also sees updated support for TMA and allows the kernel to keep one unified descriptor for tensors that share the same layout & override fields such as memory pointer and stride directly in the TMA instruction at runtime. The result is more efficient MoE model scaling even with increased expert counts. This enables the chip to allocate more time towards inference computation for higher throughput. Faster Tensor Cores For Faster Throughput The enhanced Tensor Cores have been pumped up to offer double the throughput per clock cycle. This is achieved by doubling the amount of data a tensor core can process along the K Dimension. This helps throughput-bound kernels and memory/latency bound kernels. NVIDIA explains that the benefit of a larger K dimension is fewer K loop iterations. For example, a GEMM that required four K iterations on Blackwell can be completed on just two iterations on Rubin chips. There are also improvements to attention made by combining activation sparsity with adaptive compression, and improved softmax throughput. The attention pipeline begins with a dense QK^T computation to generate the intermediate attention scores. Rubin can then load that intermediate data from Tensor Memory into a structured 2:4 sparse compressed form, generating both nonzero values and the metadata needed to use them efficiently while reducing the scores' write cost and storage requirements. This allows the later attention stages to operate on less data while preserving the dense output format expected by the rest of the model. There's also an increase in Rubin's exponential throughput capabilities over baseline Blackwell. Rubin offers 2x the FP32 throughput (per clock per SM) versus GB200 and the same as GB300, while the BF16/FP16 throughput is quadrupled over baseline Blackwell (GB200) and double vs Blackwell Ultra (GB300). Rubin also enables more fine-grained coordination between dependent kernels. This allows consumer work to begin earlier as required input data becomes available, rather than waiting for a larger set of producer work to complete. The result is a more tightly packed GPU timeline, reduced idle gaps, and improved overlap between dependent kernels. This is particularly valuable for agentic inference, where activations move sequentially through the model and kernel-to-kernel latency directly affects tokens per second per user. Incremental Gains To Power Efficiency With DSX MaxLPS Another factor that is crucial for today's AI data centers is power efficiency. The Vera Rubin NVL72 are state of the art systems and features some new tricks to expand on this, such as power steering and Intelligent Power Smoothing with energy storage into a single execution domain designed to maximize useful token output within a fixed power envelope. Transient peaks can hamper the overall throughput by stranding the available capacity at an AI factory. To address this, Vera Rubin systems are equipped with SoC (state-of-charge) intelligent power-smoothing power supplies that absorb these transient swings to reduce average power consumption by approximately 10% versus the prior generation. They also reduce 50ms peak power by 20%, delivering a streamlined power profile that can sustain max power requirements. For AI factories, NVIDIA DSX OS, which includes DSX MaxLPS, extends across GPUs, racks, liquid cooling, and workloads, offering an operational layer for scheduling, lifecycle management, and health automation that enables the provision of 40% more GPUs within the same power budget, increasing the throughput of each deployment. All of this combined is what makes Rubin a state-of-the-art chip for AI data centers, and a chip that will embark on a new journey for future AI capabilities, extending to Agentic AI workflows & to what comes beyond. Follow Wccftech on Google to get more of our news coverage in your feeds.
[9]
Nvidia's latest Rubin deal points to a bigger growth market
Nvidia's (NVDA) newest Japan project looks like an enormous chip sale. It is a lot more important to U.S. investors. A Japanese group aims to construct anAI factory made up of about 27,500 of Nvidia's Rubin GPUs and 13,750 Nvidia Vera central processing units. The 140-megawatt system will be used to train models for robotics, manufacturing, and other practical AI applications. That would be a big deployment even for Nvidia, which already sells to the world's biggest cloud providers. But the identity of the buyer matters as much as the number of processors. This is not just another data center that a U.S. hyperscaler is developing. It is a government-backed national infrastructure project supported by Japanese industries, telecom businesses, and technological groupings. That might expose the next big category of Nvidia customers. Countries increasingly desire computer power at home, models that comprehend local languages, and technologies that can safeguard sensitive industrial data within national borders. If other governments follow Japan's lead, sovereign AI factories might be yet another regular source of demand for Nvidia, on top of cloud providers, AI labs, and big corporates. The project will not be an instant cash bonanza. Construction is anticipated to begin in April 2027, with operations scheduled for June 2028. No purchase price or revenue recognition timeline has been disclosed. Its greater worth is strategic. Japan isn't just buying chips from Nvidia. It is creating a national AI ecosystem around Nvidia's chips, networking, software, and development tools. "Japan invented modern manufacturing. Now, it is building the AI factories that will power the next industrial revolution," Nvidia CEO Jensen Huang said. Nvidia needs AI demand beyond the biggest cloud companies Nvidia's current growth remains extraordinary. Revenue for the first quarter surged 85% from a year earlier to $81.6 billion. Revenue from data centers increased 92% to $75.2 billion, with a gross margin of around 75%. The numbers suggest the biggest cloud and AI companies are still spending big. They are also dangerous for long-term investors. Nvidia is making a lot of money off a tiny number of customers. Three direct customers accounted for 21%, 17%, and 16% of total sales, the company stated in its last quarterly filing. Hyperscalers made up about half of the data center revenue in the quarter. The other half was from AI clouds, industrial businesses, enterprises, and sovereign customers. Japan's proposal bolsters that second category. The more Nvidia can develop with governments and industrial customers, the less its growth will be dependent solely on another surge in expenditure increases by a few American computer titans. That could reduce customer concentration, though it would not eliminate the risk. Even a project of country scale can depend on a single government budget, one consortium, and a small pool of infrastructure providers. Construction delays, power restrictions, or political changes can delay revenue. Sovereign AI also possesses attributes that could render it resilient. Governments see domestic computer capacity as economic infrastructure. Among its goals are data sovereignty, national security, industrial competitiveness, and access to artificial intelligence systems made for the local environment. Those priorities can encourage investments, even if it is difficult to quantify the short-term financial return. Nvidia has already begun to prepare investors for this expanded customer mix. Its new reporting framework separates hyperscaler revenue from what it calls AI clouds, as well as industrial and enterprise demand. Nvidia says the latter category includes purpose-built AI factories across industries and countries. That's about what the Japan deployment looks like. Noetra leads the project, with support from 44 enterprises and organizations. Core members include Sony Group, SoftBank, NEC, and Honda, with engineers from Preferred Networks and Japan's National Institute of Advanced Industrial Science and Technology helping to develop the models. That structure provides Nvidia access to more than one data center customer. The models and development tools developed in the project will be used in manufacturing, logistics, health care, telecommunications, mobility, and robotics. As those businesses construct applications around Nvidia technology, the first sale of infrastructure might create more demand for edge processors, simulation software, and robot-development systems. Japan is buying Nvidia's full AI factory, not only Rubin chips The headline figure is 27,500 Rubin GPUs. The bigger sale is a lot bigger. Noetra's AI factory will also use 13,750 Vera CPUs, Nvidia Spectrum-X Ethernet networking, BlueField data-processing units, and the company's DSX framework for creating and operating massive artificial intelligence systems. That's important because Nvidia is increasingly looking to position itself as an infrastructure firm, not just a chip seller, for investors. Once revenue is from a stand-alone processor sale. A full stack system can include Nvidia across compute, networking, storage, software development, and ongoing operations. It also makes it more difficult for a consumer to replace one component with a competitive product later. Nvidia's DSX reference design encompasses the entire AI factory stack, from computation and networking to storage. The company says it's a repeatable strategy for developing big clusters with predictable performance and efficiency. Japan's effort will test that method on a national scale. Noetra wants to start creating a Japanese reasoning model in the fiscal year ending March 2027. It then plans to develop a model that can interpret text, photos, video, and audio by fiscal 2028, followed by systems that can understand real settings by fiscal 2030. The last step is the most crucial to Nvidia's long-term growth story. Physical AI is about systems that can understand the real environment and control machinery. Uses include robots that carry items in warehouses, examine industries, assist in hospitals, or drive automobiles. Japan is a natural market. Its big industrial enterprises already have manufacturing knowledge, robotics capabilities, and lots of data from the physical world. "Nvidia is providing the computational platform to turn those assets into trainable AI models. The Japanese government picked Noetra and its research collaborators for the national program on multimodal models for AI robots and physical AI from 15 applications. The published models will be introduced in stages and provided to Japanese developers and businesses. That might magnify Nvidia's ecosystem advantage. A factory taught on Rubin hardware may build models later served via Nvidia's Cosmos, Isaac, Jetson, and comparable software platforms. Separately, Nvidia claimed Japanese manufacturers and robotics businesses are embracing those technologies for use in industries, mobility, construction, agriculture, and health care. That's similar to how Nvidia has succeeded in cloud computing. First, Nvidia supplies the computing infrastructure used to train the model. Developers build around Nvidia software, and customers may later deploy those models on machines containing Nvidia edge processors. Finally, customers use the intelligence obtained via machines and edge devices that can also have Nvidia technology. Japan's Rubin project might thus stimulate demand at various levels from a single infrastructure commitment. Bloomberg / Getty Images What Nvidia investors should watch before calling it a windfall The first question is whether the project will be constructed and operational on schedule. Companies participating say construction will commence in April 2027, with operations slated to begin in June 2028. That implies the project could demonstrate future demand but won't add much to Nvidia's next few quarterly earnings reports. Investors shouldn't count all 27,500 GPUs as immediate revenue. The firms have not specified price, delivery date, payment terms, or whether the systems will be implemented in various phases. Rubin availability is the second difficulty. The Rubin platform is ramping into full production, with server makers and supply chain partners preparing systems at scale, according to Nvidia. Nvidia and its partners will have to build thousands of processors, racks, and networking components in addition to supplying cloud providers and other national initiatives; Japan's purchase adds to proof of demand. The third question is whether Japan makes usable models. Financing can develop a data center for government. It doesn't promise that the models educated there will be appealing to developers, increase manufacturing productivity, or lead to commercially successful robots. Noetra aims to disseminate its models broadly, which may help adoption. Yet open access can also create uncertainty about where the ultimate economic value will reside. As long as training and deployment are tied to the Nvidia platform, Nvidia wins. The fourth problem is if the model repeats elsewhere. Nvidia has other sovereign AI opportunities than Japan. The company has worked with SoftBank on domestic Japanese AI infrastructure in the past and is working with SK Telecom on a gigawatt-scale AI cloud in South Korea. Several comparable programs would be significantly more important than a single nationwide deployment. Key takeaways for Nvidia investors * Japan plans an AI factory containing about 27,500 Rubin GPUs and 13,750 Vera CPUs. * The system will also use Nvidia networking, data-processing units, software and its DSX infrastructure design. * Construction is expected to begin in April 2027, with operations targeted for June 2028. * The project expands Nvidia's potential customer base beyond U.S. hyperscalers. * Government-backed AI factories could reduce customer concentration while extending Nvidia's software ecosystem. * Revenue timing, infrastructure execution and eventual model adoption remain uncertain. The last problem is valuation. The company has already become one of the largest firms in the world, and investors are looking for its next architectures to drive huge sales. Deploying 27,500 chips sounds dramatic, yet in its last quarter, Nvidia generated $75.2 billion in data-center revenue. This is the financial framework that the Japan initiative must fit into. One order changing the financial statement at Nvidia is irrelevant. The project demonstrates how the company can continue to grow even after the top cloud providers have developed many generations of AI technology. Governments invest in AI infrastructure for different reasons than technology companies do. They may create domestic capabilities to maintain data sovereignty, support local language models, automate vital industries, or reduce reliance on foreign AI services. It can profit from each purpose but still provide the underlying American technology. That makes for a powerful, if slightly ironic, business model. Japan is looking for a local foundation model to help reduce reliance on foreign artificial intelligence systems. The country is looking to Nvidia's U.S. hardware and software stack for building it. That is the hidden implication for U.S. shareholders. Sovereign AI doesn't necessarily mean a weaker Nvidia as countries build their own local models. It might also extend Nvidia's market by giving countries a reason to build their own hyperscale infrastructure. The Japan project won't dramatically affect Nvidia's next quarter. But it may point to where Nvidia finds its next generation of clients. The Arena Media Brands, LLC THESTREET is a registered trademark of TheStreet, Inc. This story was originally published July 19, 2026 at 12:13 PM.
[10]
Nvidia's Rubin reassurance protects a much bigger AI bet
Nvidia (NVDA) CEO Jensen Huang has pushed back against reports that manufacturing problems could delay the company's next artificial intelligence platform, Bloomberg reported. If those reports were wrong, it would protect investors far more than just a single product launch. Nvidia's rapid growth is partly a function of its ability to roll out ever more powerful devices before customers finish adopting the previous generation. That quick cadence pushes cloud providers to spend, keeps competitors from catching up, and provides consumers a reason to stay within Nvidia's hardware and software ecosystem. Next, we have Vera Rubin. Nvidia claims Rubin-based products will go to partners in the second half of 2026, Bloomberg notes. Any major delay may disrupt customer plans, just as the corporation is trying to turn the excitement around AI agents and robotics into another big source of demand. Huang said in Tokyo on July 15 that Rubin hardware was already in production and headed toward "giant" volumes, according to Bloomberg, rejecting reports of manufacturing difficulties involving a specialized circuit board. His assurance is significant because Rubin is not merely a more rapid successor to Blackwell. It is the technology Nvidia hopes will power the next generation of AI factories and serve as a bridge into physical AI when artificial intelligence moves beyond chatbots and begins managing robots, factories, and autonomous machinery. "Vera Rubin is already in production. Giant amounts of production incoming," Huang said, as Tom's Hardware confirmed. Nvidia's growth depends on keeping Rubin on schedule Nvidia enters the Rubin transition in a position of tremendous financial strength. The corporation reported record first-quarter revenue of $81.6 billion, up 85% from a year ago. Data-center revenue surged 92% to $75.2 billion, while Nvidia forecast revenue of around $91 billion for the next quarter, using its fiscal 2027 first-quarter statistics. Those data indicate that customers continue to consume Blackwell systems at massive volume. They also create expectations. Once a firm reaches the size of Nvidia, it takes more and bigger additions of income to keep growing fast. A delayed architecture might postpone data center construction, upset orders with suppliers, and offer customers more time to explore alternatives. Nvidia first unveiled Rubin in January as a six-chip architecture built on graphics processing units, central processing units, networking, and storage. Rubin has started full production, with partner availability expected in the second half of 2026, the company stated. By March, Nvidia had extended the platform to seven chips and pitched it as infrastructure for agentic artificial intelligence that can handle multistep tasks with little human input. The whole Vera Rubin platform is now in production. That message was bolstered by Nvidia in May, when it said server makers and supply-chain partners were ramping up Rubin systems. Huang's current comments are obviously more than a typical denial. They are a defense of Nvidia's core pledge to investors: that it can migrate from one major platform to another without a long product gap stifling its growth. It's not just individual chips contributing to the company's recent edge. Now it creates full systems that incorporate CPUs, networking, software, and racks. That method can boost performance but can also raise execution risk. More components need to function together, and manufacturers need to construct ever denser and more sophisticated systems. This integration is on the magnitude of Vera, the platform's central processing unit. Nvidia says the Vera CPU is in full production and can do specific AI-agent tasks 1.8x quicker than standard x86 CPUs. Rubin's success will depend on Nvidia and its partners turning those individual technologies into full systems customers can reliably install. That makes manufacturing timing a direct investment issue, not just an engineering detail. PHILIP FONG / Getty Images Japan shows why Rubin is bigger than another data-center chip Huang's choice to reach out to Rubin in Japan also alludes to the greater possibility for the platform. Japan has world-class manufacturers, factory automation businesses, and robotics experts. It also has a dwindling workforce, which provides companies a strong economic incentive to automate more physical tasks. Japan's preliminary census estimates showed a population of 123.05 million in October 2025, down 3.1 million from 2020. More than 90% of Japan's municipalities suffered population reduction, according to the Statistics Bureau. This demographic pressure makes robotics more than a speculative technical trend. To sustain output, Japanese firms may need to use machines that can learn, adapt, and execute a greater variety of activities if the labor pool declines. Japan's administration is on the right track. In June, the Ministry of Economy, Trade, and Industry updated its AI Robotics Strategy, keeping the target of deploying about 10 million robots by 2040 in 18 key sectors, according to NHK World Japan. The plan comprises labor-intensive businesses such as manufacturing, health care, and food services. And that's where Nvidia wants to be the computational layer behind that transformation. In its review of the Japanese AI and robotics ecosystem on July 15, it mentioned work with cloud providers, manufacturers, universities, and robotics developers using Nvidia technology. That might help diversify Nvidia's AI narrative. The company is currently focusing its data-center growth on a relatively small number of big cloud providers and technology enterprises. Robotics may boost demand from manufacturers, logistics companies, hospitals, and industrial firms. They also relate to workload. Developers are able to train robot models in data centers, test them in simulations, and then run them on processors within physical machines. Nvidia can potentially sell technology at each step. Its Isaac robotics platform offers models, simulation tools, data pipelines, and computer systems for building and deploying AI-powered robots. This full-stack strategy resembles the approach that made Nvidia dominant in data centers. The corporation doesn't want to sell the processor inside a robot. It wants developers to train the model using software from Nvidia, test it using simulation tools from Nvidia, refine it using servers from Nvidia, and manage it using edge computers from Nvidia. Rubin may shore up the data center side of the chain by backing the big AI factories required to train ever-more-sophisticated physical AI models. That's the greater gamble Huang's production comments are defending. What Nvidia investors should watch next The first question is whether Rubin systems will start to reach customers in the back half of 2026 as predicted. Producing is not the same as mass-deploying to customers. Nvidia and its manufacturing partners need to build, test, and ship full racks in sufficient numbers. Then, customers require power, cooling, and networking infrastructure to install them. Investors should be listening for signs of Rubin revenue, fixed delivery timelines, and client deployments during upcoming earnings calls from Nvidia. The second question is whether Blackwell demand holds during the transition. Demand for AI computing is outstripping supply, so customers may continue buying Blackwell. But some purchasers may choose not to order if Rubin adds enough extra performance to make waiting worthwhile. Nvidia has to deal with that shift without generating a revenue lull or developing a bunch of old gear consumers don't want. The third development to watch is whether physical AI begins to produce measurable business. Japan is a really interesting demonstration market with both modern manufacturing and very strong demographic pressure. Successful deployments there could drive uptake in other aging economies and labor-constrained industries. Key takeaways for Nvidia investors * Huang says Vera Rubin is already in production, despite reported manufacturing concerns. * Nvidia expects partners to offer Rubin-based products during the second half of 2026. * Rubin's timing matters because Nvidia must sustain rapid growth from an increasingly large revenue base. * The platform is designed for AI agents and the data centers that train physical-AI systems. * Japan's shrinking population creates a strong economic incentive for factory and service-sector automation. * Robotics could broaden Nvidia's customer base beyond large cloud providers. Another variable is China. Huang said Nvidia only started selling H200 chips to the U.S. while the government was starting to assess permits on a case-by-case basis. The Commerce Department's H200 export policy provides for case-by-case licenses if exporters and buyers meet security conditions, Reuters reported. But those sales would still be subject to decisions made in Washington and Beijing, and they may boost the bottom line. Rubin is something that Nvidia can influence more directly: execution. The corporation has to show that more complex technologies can move from announcement to production to client data centers without a harmful delay. Huang's denial eases some worries, but investors still need to see shipments and revenue. And that is why the Rubin argument is important. Nvidia is no longer being evaluated just as the top supplier of AI chips. Investors expect it to continue an aggressive product cadence while moving its platform into AI agents, autonomous machines, and robotics. Keeping Rubin on track maintains that bigger thesis. A successful launch would demonstrate that Nvidia can continue to feed the data center expansion and provide the computing infrastructure for a new industrial market. A delay, however, would threaten both assumptions at once. The Arena Media Brands, LLC THESTREET is a registered trademark of TheStreet, Inc. This story was originally published July 16, 2026 at 2:47 PM.
[11]
NVIDIA CEO Rips Apart The "Chip Delay" Narrative, Says "Giant Amounts" of Vera Rubin Coming & Unveils Japan's First AI Factory With 27,500 Rubin GPUs
Jensen Huang's Japan visit comes with the unveiling of the country's first Vera Rubin AI factory, as the CEO of NVIDIA confirms chips are on track. A Flood of Vera Rubin Is Coming, & NVIDIA's Roadmap Is On Track, Says CEO Jensen Huang, Pointing Out That AI Is Here To Stay For The Long Haul During his ongoing visit to Japan, NVIDIA's CEO Jensen Huang made several statements regarding the future outlook of AI and the company's roadmap for AI. Essentially, Jensen is reaffirming the statements from a week ago, which quashed rumors surrounding chip design issues and delays affecting Rubin platforms. When questioned on talks regarding chip delays, Jensen said that those rumors are not true at all, and once again confirmed that the Vera Rubin platform is in production, and "Giant Amounts of production" are incoming. We know that the NVIDIA Vera Rubin platform is already in full production, following the volume production of its Vera CPUs, which are aiming to be a major success in the DC CPU segment. Rumors regarding NVIDIA's chip and rack delay were at their peak a few weeks ago, which not only cited production callbacks for the upcoming Oberon NVL144 solutions, but also the next-generation Kyber NVL576 solutions, which will harness Rubin Ultra chips. There are still ongoing talks on whether NVIDIA will be able to ship the original 4-reticle die solution during the Rubin Ultra ramp or if they will tone it down to a 2-reticle solution. What we know is that most of the initial design changes and drawbacks are rectified almost immediately as NVIDIA works with its robust supply chain and ecosystem partners. This was proved with the Blackwell and Blackwell Ultra launches, as we have disclosed previously. Vera Rubin AI Factories To Advance Physical AI NVIDIA's Vera Rubin and Vera CPUs are already landing at major AI firms, and in Japan, the company announced the country's first AI Factory. Part of Japan's FRONTia Project, which is responsible for the Development of Multimodal Foundation Models with a View to AI Robotics and Physical AI, the AI factory will be made in partnership with Noetra Corp. "Japan invented modern manufacturing. Now, it is building the AI factories that will power the next industrial revolution," said Jensen Huang, founder and CEO of NVIDIA. "NVIDIA is honored to partner with Japan and its industrial leaders to build the AI infrastructure that will power the country's industries, its economy and a new generation of innovation." The Japanese Vera Rubin AI factory will consist of 13,750 Vera CPUs and 27,500 Rubin GPUs. It will be able to output 140MW of Data Center capacity based on the NVIDIA DSX platform. The Factory will house NVIDIA Vera Rubin NVL72 racks and will scale through Spectrum-X networking. The AI factory will play a major role in advancing the global AI market by 2040, representing a $133 billion TAM. Japan, the birthplace of modern-day humanoids and robots, is an essential region for NVIDIA to invest in. Various Japanese companies have already partnered with NVIDIA and will leverage its latest Jetson Thor robotics platform to advance their capabilities and the Physical AI era. NVIDIA Will Spend The Next Decade Building Infrastructure For AI Now, on the topic of AI itself, Jensen made some interesting comments. According to him, most technology cycles last anywhere from 10-15 years before they plateau (essentially reaching the peak). But for AI, we are still in the early rounds; it's only been a few months since the tech really started to gain traction, and as Jensen puts it, "We're at the beginning of this one", referring to the AI cycle. So he says that in another 10 years, it will be an interesting question whether AI goes up or if it flattens out, but Jensen believes that AI "will never go down". In another interview with the press, Jensen said that "we are a long way from an AI bubble", referring to the fact that AI is here to stay for a long time and that infrastructure needs to be built for a long time, "at least a decade," to support AI in the long-term. AI is still in its "useful" era with the rise of Agentic AI, and we haven't even seen its full potential, let alone the Physical and Future AI eras. "Every nation and every company should own and control its intelligence infrastructure. Open models make that possible," said Jensen Huang, founder and CEO of NVIDIA. "They give countries, enterprises and researchers the freedom to inspect, improve, adapt, secure and deploy AI for their own needs. Together with Japan's AI leaders, we are advancing an open AI ecosystem that accelerates discovery, strengthens national capability and ensures every society can participate in -- and benefit from -- the AI revolution." Jensen was also asked whether he sees Rapidus as a potential partner of NVIDIA in the future, and he said that the demand for AI will be beneficial for all semiconductor manufacturers, and they look forward to seeing Rapidus's progress in this field. So not much into a potential collab, but the door remains there. With that said, Jensen's Japan visit has been just as tremendous as his Taiwan and Korea visits. Follow Wccftech on Google to get more of our news coverage in your feeds.
Share
Copy Link
Nvidia's Vera Rubin platform has reached full production, with CoreWeave reporting 10 times more tokens per watt compared to the previous GB200 NVL72 system. Major customers including OpenAI, Microsoft Azure, Google Cloud, and Meta are deploying the next-generation AI infrastructure, which combines 72 Rubin GPUs with custom Vera CPUs to handle agentic AI workloads more efficiently.
Nvidia Vera Rubin has officially entered full production, marking a significant milestone in the company's push to dominate next-generation AI infrastructure. Ian Buck, vice president of accelerated computing at Nvidia, confirmed that systems are now shipping to major customers including OpenAI, CoreWeave, Microsoft Azure, Google Cloud, Meta, Oracle Cloud Infrastructure, and Dell
4
. OpenAI plans to deploy the AI platform at scale during the third quarter, according to reports from the company's headquarters briefing4
.
Source: Wccftech
The production ramp represents the culmination of extreme hardware-software co-design across seven chips and five rack trays, all engineered as a single system rather than assembled from separate components
3
. This approach targets the specific demands of agentic AI workloads, which can consume up to 15 times more tokens than traditional AI applications3
.The most striking performance benchmark comes from CoreWeave, which reported achieving 10 times more tokens per watt when running DeepSeek-R1 on the Vera Rubin NVL72 system compared to the previous generation GB200 NVL72 based on Blackwell
3
. This dramatic improvement in performance per watt directly addresses the power constraints that define modern AI factories, where delivered performance within a fixed watt data center determines profitability2
.CoreWeave also demonstrated that optimizations developed for Vera Rubin can be applied retroactively to GB200 NVL72 systems, improving throughput per megawatt by more than four times over a three-month period
5
. These gains were validated across 250,000-plus distinct configurations and more than 1.4 million GPU hours of testing5
.The Rubin GPU sits at the heart of the platform's inference capabilities. Each chip joins two compute dies using the Nvidia High Bandwidth Interface, delivering 224 Streaming Multiprocessors containing 896 Tensor Cores alongside 288GB of HBM4 memory with 22 TB/s of memory bandwidth
1
. As an inference-focused accelerator, Nvidia highlights Rubin's 50 sparse PFLOPS of NVFP4 inference throughput as its headline performance figure1
.
Source: The Register
Key architectural improvements target the specific challenges of modern transformer-based LLMs and mixture-of-experts models. The Tensor Memory Accelerator now supports GPU kernels that maintain a single unified MoE descriptor directly in the instruction at runtime, reducing computation overhead as expert counts grow
1
. Rubin also doubles the K-dimension throughput in Tensor Cores, cutting the number of loop iterations required for matrix multiplication in half compared to Blackwell .Softmax performance, essential for attention calculations in long-context models, receives up to a 4x boost versus Blackwell through enhancements to the Special Function Unit
1
. These improvements matter as advanced models now support context lengths of up to one million tokens.The Vera CPU represents Nvidia's bet that faster cores beat more cores for agentic AI workloads. Built on the custom Olympus core microarchitecture, the processor delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs
3
. In recent benchmarks, Nvidia demonstrated 1.9 times faster agentic performance and a six-fold improvement in latency over x86-based alternatives5
.The company also claimed the Vera CPU surpasses AMD's EPYC Turin chip by nearly 100% on selected industry benchmarks, particularly Python workloads that dominate the software stack for AI inference
4
. Anthropic and OpenAI are among the first labs to receive the processor, alongside Perplexity, SpaceX, and Oracle4
.Networking infrastructure plays a critical role in the platform's efficiency. The sixth-generation NVLink 6 scale-up fabric delivers more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet
3
. For scale-out networking, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs, enabling 1.6x higher RDMA bandwidth than standard Ethernet3
.
Source: Tom's Hardware
Leading AI infrastructure builders including CoreWeave, Microsoft, SpaceXAI, and Tesla are among the first to deploy Spectrum-6 switches in their data centers
3
. The sixth-generation NVLink delivered 2.3 times higher simulated decode throughput for massive language models compared to Ethernet-based networks5
.Related Stories
Three generations of rack-scale co-design produced a Vera Rubin NVL72 system with no cables, fans, or hoses in the compute tray, cutting assembly time from 90 minutes to just one minute—a 90x improvement
2
. Andrew Bell, senior vice president of hardware engineering, demonstrated the streamlined installation process during a lab tour at Nvidia's Sunnyvale facility2
.The platform's 45-degree Celsius liquid cooling inlet temperature design enables chiller-free dry-cooler operation, saving approximately 4 million gallons of water per megawatt annually compared to standard cooling methods
3
. By dynamically optimizing the full infrastructure and energy stack, Nvidia can deploy 40% more GPUs within the same power envelope5
.Vera Rubin production spans 350-plus factory sites across 30 countries, representing what Nvidia calls the largest and most mature rack-scale supply chain ever assembled
3
. The platform underpins a newly expanded partnership between Microsoft and Mistral that brings frontier AI to Europe, with Mistral adding GPU capacity drawing on thousands of Rubin GPUs3
.Jensen Huang first declared the platform in full production at Computex in early June, naming Anthropic, OpenAI, SpaceX, and Oracle as early recipients
4
. The recent headquarters briefing added performance numbers and customer testimonials that transform the Computex announcement into measurable benchmarks4
.The timing proves awkward for Nvidia's stock performance. While the company's shares have risen 9% this year, the broader chip index has climbed 66% over the same period, with Intel, ARM, and AMD all more than doubling
4
. Analysts project Nvidia's revenue will increase 82% to roughly $393 billion for the fiscal year, though this growth hasn't translated into the share price momentum seen during the Blackwell cycle4
.The business model centers on token economics. Buck emphasized that AI factory revenue depends on how many tokens can be generated in a fixed watt data center, with infrastructure rented for dollars per hour generating revenue through profitable tokens in the hundreds of dollars per hour
2
. As token supply increases through more efficient hardware and competing providers, downward pressure on token cost creates uncertainty about whether demand will grow fast enough to justify current capital expenditures2
.Summarized by
Navi
[1]
[2]
[4]
25 Feb 2026•Technology

06 Jan 2026•Technology

29 Oct 2025•Technology

1
Technology

2
Science and Research

3
Technology
