7 Sources
[1]
Nvidia Wants to Own Every Chip Inside AI Data Centers
Nvidia's Vera Rubin platform combines CPUs and GPUs into a single system, reflecting the company's growing ambition to power every layer of AI infrastructure. Nvidia is hyping up its new Vera Rubin chip system this week, revealing new performance benchmarks for the GPU and CPU combo ahead of rival AMD's annual product event in San Francisco on Thursday. During a lengthy technical workshop last week at the company's headquarters in Santa Clara, California, Nvidia executives boasted to a small group of journalists about the chip system's increased power and efficiency capabilities. The biggest takeaway: Nvidia, which has long specialized in making GPUs, is increasingly trying to position itself as a supplier of CPUs that can power AI agents. While GPUs are still the main hardware that companies use to train and run their AI models, the industry's shift toward more complex, agentic systems has increased demand for CPUs, which can orchestrate data flows, networking, and other software tasks. That's likely one reason Nvidia has been eager to position itself as a supplier of complete AI systems, rather than just AI chips. Vera Rubin is Nvidia's successor to its hybrid superchip system Grace Blackwell, and represents the linchpin of its near-term future powering the AI industry. It's designed to offer one CPU for every two GPUs. In a single Vera Rubin NVL 72 super chip system, there are 36 Vera CPUs for every 72 Rubin GPUs. Nvidia is also selling the Vera CPU as a standalone product, and it has reportedly told Chinese customers these could be ready as soon as August. Nvidia executives emphasized that its new Vera Rubin NVL72 racks -- a stack of chips packed into a single liquid-cooled platform -- are much more "plug-and-play" than some of its earlier products. During a brief tour of a Nvidia data center lab in Silicon Valley, Nvidia executives shared that OpenAI already has one Vera Rubin rack in use. Nvidia CEO Jensen Huang didn't make an appearance at the workshop in Santa Clara last week; he was in Japan announcing the chipmaker's new partnerships with a number of Japanese firms to develop AI for robotics. The briefings were instead led by Ian Buck, Nvidia's longtime vice president of accelerated computing and the architect behind the company's CUDA software. "We're on a roadmap to crank out new architectures, not just GPUs but CPUs," Buck told reporters. "We're going to keep innovating, because it's do this or die in Silicon Valley." The meetings were held in Huang's executive briefing center, where multiple desks nearby were piled with bags of Taiwanese snacks that the CEO brought back from his recent trip to Computex, a massive annual semiconductor trade show in Taipei, an Nvidia spokesperson told WIRED. Nvidia claims that the Vera Rubin NVL72 system will process ten times as many tokens per watt as the company's Grace Blackwell super chip. The company says that its Vera CPU is also faster at processing agentic AI tasks compared to rival CPUs from AMD and Intel (though the tests it ran to support those benchmarks appears to have used slightly older generations of its competitors' CPUs). Localized memory subsystems on the new chips will also offer nearly three times as much memory bandwidth as Blackwell, which will likely be an appealing feature to many companies amid an ongoing shortage of high band-width memory. Nvidia says it has also significantly reduced the number of cables needed to connect its chips to racks in multi-rack server systems, to the point where the company is touting Vera Rubin as "cable-free compute" and "hot-swappable." This means customers can theoretically reduce the amount of time it takes to install each rack from a couple of hours to a few minutes, a point that was brought up by both Buck and Andrew Bell, Nvidia's senior vice president of hardware engineering. And the new chip system is 100 percent liquid-cooled, which can reduce the amount of energy needed to cool the chips, since air-cooling is more energy intensive. Ever since Nvidia unveiled Vera Rubin in the spring of 2025, the company has been slowly dribbling out more details about the chip system while insisting it will be released on schedule. Huang has repeatedly said Vera Rubin is ramping to "full production" and will ship in the second half of this year, with early customers including Microsoft, OpenAI, and Oracle. Nvidia is particularly sensitive to any suggestion of delays after its previous-generation Blackwell chips reportedly overheated when connected together in the company's customized server racks, forcing it to make design changes and push back shipments. Nvidia's marketing push for Vera Rubin is happening just ahead of rival AMD's annual conference, where executives are expected to tout its next-generation AI and data center chips. On Sunday, AMD revealed more details about its Helios AI chip rack, which is designed to compete with Nvidia's new wares. Both AMD and Nvidia have been vying for large-scale, multi-year contracts to supply AI hyperscalers like Meta and Amazon and AI labs like OpenAI, Anthropic, and SpaceXAI with chips. Over the past two years, AMD has significantly grown its share of the market for CPUs used in data centers. The company has long been recognized as a pioneer of the modern chiplet architecture used in x86 processors, which still account for the vast majority of data center CPU revenue. Nvidia, by contrast, builds its data center CPUs on ARM, an alternative chip architecture known for its power efficiency. Nvidia executives Buck and Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, both emphasized that Vera Rubin abandons the chiplet architecture used by many modern processors in favor of a single, monolithic chip. Coutand argued that stitching together multiple chiplets imposes "a heavy tax on memory bandwidth and data movement," whereas the monolithic design of Vera Rubin allows data to move more quickly across a single integrated circuit.
[2]
Nvidia deep dives Vera CPU for AI data centers -- SPEC CPU 2026 benchmarks revealed, Olympus architecture specifics, and more
Nvidia's Vera CPU is its first bid to become a key player in the data center CPU market. Although Grace has seen some success (most notably with Grace standalone deployments at Meta), Vera is Nvidia's first CPU with a custom core design. It's arriving at an ideal time, as well, with the server CPU market exploding in the last few months on the back of agentic AI demand. Vera isn't a chip built to chip away at the market share of AMD and Intel in the cloud. It's built to grab market share in an expanding market, as hyperscalers look to widen AI infrastructure beyond legacy clouds. As such, it's designed in a much different way than Nvidia's x86 competitors, and it even holds some unique architectural design points compared to the swath of Arm-based designs. Nvidia has slowly revealed more details about Vera as it ramps into general availability, which is on track for the back half of this year. Now, we have a full picture of the chip. Nvidia shared its Vera white paper, along with unofficial SPEC CPU 2026 results comparing Vera to AMD's Turin-based Epyc 9755. We're going to break down the white paper here, including all of the details about the Olympus core and a look at the benchmarks Nvidia ran. At the end of this piece, we'll also take a brief look at the larger context of Vera and how it fits into Nvidia's wider AI ecosystem compared to standalone deployments. But plenty of ink has been spilled about Vera's technical capabilities and Nvidia's next-gen AI infrastructure vision. Let's start with the important thing: the benchmarks. Nvidia Vera CPU benchmarks We've seen Vera in action before, though only through a series of selected benchmarks ran at Nvidia HQ by Phoronix. In the Vera white paper, Nvidia shared benchmarks for SPEC CPU 2026, specifically the integer suite from SPECrate, against AMD's Epyc 9755, with both chips running in a dual-socket configuration. Before getting into the results, there are some important notes about how SPEC runs work, and the reporting criteria for them. Nvidia's run here isn't official, as Vera was tested in a reference system due to the fact that it's not broadly available yet. It's ramping for general availability in the second half of the year. Due to that, Nvidia is unable to report its results. That's why you see "estimated" in some of the charts below. Nvidia ran SPEC CPU 2026; it's not extrapolating expected performance like we've seen from AMD so far with its upcoming Venice chips. SPEC CPU 2026 is split into four suites, but Nvidia tested the SPECrate integer suite, which is focused on system throughput with integer-based workloads. The "rate" result is looking at how much work is completed within a certain amount of time. Here, each thread in the system has a copy of the workload. The score is how much time it takes for those workloads to complete, regardless of thread count, naturally giving chips with more cores an advantage. If you want more detail on the benchmarks included in the suite, make sure to read our original coverage of SPEC CPU 2026. Here are the overall results: Nvidia didn't share the exact results for the 9755 it tested, short of the overall score of 898. Taking that overall score into account, Vera is 3% ahead of the 9755. It's worth noting that Vera is ahead here despite a large thread disadvantage. An overall score of 898 for a dual-socket Epyc 9755 system isn't unreasonable compared to publicly-submitted SPEC CPU 2026 runs, though higher results have been published. SPEC CPU ships as source code, which the tester must compile with their compiler of choice, and that can heavily influence results (particularly with vendor-specific compilers). Nvidia used GNU 15.2 with both systems. Above, you can see Vera's results stacked up against the 9755, but these aren't comparing the numbers directly. Nvidia has normalized the per-core performance, which isn't how SPECrate results are normally shared. According to the overall numbers, Vera is still completing more work within the same amount of time, despite a thread disadvantage, but the margins aren't in the range of a 70% or 80% advantage as the above chart suggests. We asked Nvidia about the results given that they're obfuscated by comparison; we could not reverse-engineer the Epyc 9755's scores with the information Nvidia has provided. Here's the response it gave: "Per-core performance under a fully loaded socket is important because agentic AI and RL run many sandboxes concurrently, while each agent step remains sequential and latency-sensitive. It measures how much performance each core sustains amid contention for shared power, memory, cache, and fabric. We therefore normalize by physical core, with SMT enabled on both systems." The "agentic" workloads Nvidia has highlighted here are code compilation and interpretation workloads, which is something an agent is often doing, querying repos for dependencies and building source code. Below are data science workloads (or Exploratory Data Analysis), and below that are data processing workloads like SQLite database management. The results here align with Nvidia's overall messaging of Vera, that it's highly competent at data-rich, backend operations. Although Nvidia is sharing per-thread results, it argues that SPECrate is still the correct benchmark to run. The per-thread results here are in the context of a fully-loaded socket. Here's the justification from the white paper: "This metric is non-trivial for agentic AI and RL systems, where many sandboxes, tools, and environments run concurrently rather than as isolated single-thread tests. Fully loaded per-core performance captures how well each core sustains throughput while sharing socket-level power, memory bandwidth, cache, and fabric resources." In addition to running the workloads, Nvidia analyzed the code execution for architectural benchmarks, which you can see in the gallery above. Nvidia claims an overall IPC gain of up to 1.9x compared to Turin, up to 2.3x more branch predictions and 3.5x taken branches per cycle, and up to 2.4x higher instruction fetch operations per cycle. Outside of SPEC, Nvidia shared a few benchmarks highlighting the capabilities of the Olympus core. First up is PageRank, an algorithm developed by Google to originally rank web pages, which highlights Olympus' prefetch engine. Nvidia scaled this workload to higher core counts, showing Vera maintaining much of its single-core performance up to 32 cores, while the Turin chip hits a wall around 20 cores. In addition to the above results, Nvidia shared some tests of the Vera memory system compared to Turin. These microbenchmarks are good for validating Nvidia's specifications, but they're looking at architectural performance, not application performance. An architectural advantage translates into a performance advantage, but not always in a linear, expected fashion. Nvidia used internally-developed tools for the memory tests, though they're available on GitHub for anyone to run. First is loaded memory latency, stressing the memory subsystem as bandwidth usage increases. Vera has much higher bandwidth overall, but you can see the Turin chip hit a latency wall below its maximum, which Nvidia attributes to Non-Uniform Memory Access (NUMA) domain traversal and CCD-to-CCD latency. Looking at per-core bandwidth, Nvidia claims Vera provides more than four times the bandwidth of AMD's 9755. The suggestion here is that "real-world" per-core bandwidth is even better than Nvidia's specs lead on (or perhaps worse than AMD's). Maybe the most consequential of these tests is the one you can see above, looking at core-to-core latency. It's no secret that crossing the CCD on AMD's chiplet-based architecture incurs a big latency penalty. You can see that in action even in our Ryzen 9 9950X3D2 review, and the penalties compound as you scale up the number of CCDs. In fairness to AMD here, chiplet-based designs aren't built for this type of cross-CCD traversal, preferring to keep workloads localized and optimizing for core density. Vera's design goal is clearly to keep latencies consistent across the entire die and sacrificing core density in the process. Nvidia's Ian Buck told us that this design trade-off "will come at the cost of the legacy workload," when we recently visited Nvidia HQ. That's important context. Nvidia isn't gunning to steal existing market share from AMD and Intel as much as it's trying to grab market share in an expanding market before AMD and Intel can. Some financial institutions (including Morgan Stanley and Bank of America) suggest the server CPU market could double in size (or grow even larger) by 2030. That context is important because there will be a continuing demand for CPUs that can handle workloads Vera is not optimized for, and it'll be interesting to see how AMD and Intel tackle that dynamic with future products, trying to keep a legacy base of customers while pushing ahead into the expanded market. Nvidia clearly has a vision of how that expanded market looks, and to that end, hasn't shared SPEC CPU floating point results. Presumably, this is due to the fact that SPEC's vectorized suite is focused primarily on HPC workloads, whereas Nvidia focused on what it believes are critical agentic workloads that are integer-based. Vera has a vector engine complete with SVE, but that doesn't seem like Nvidia's focus. In an end-to-end Nvidia system, those vectorized workloads would be offloaded to a Rubin GPU. Still, we don't have any vector results for Vera yet. Up to this point, we've only seen integer results, which is strange given the memory system at play in Vera.
[3]
Nvidia details its next-generation Vera CPU for AI, setting up challenge to AMD and Intel
Nvidia became the most valuable company because of insatiable demand for its graphics processing unit, or GPU, the primary chip used for creating and deploying artificial intelligence. But the chip giant is now shipping its own central processing units, or CPUs, which cloud providers could decide to deploy in place of those from Advanced Micro Devices and Intel, opening up a new battleground in AI servers. On Tuesday, Nvidia released new information about its data center CPU, called Vera, including specifications and the kind of benchmarks and architectural information that prospective customers need to fully evaluate the chip. Nvidia representatives said Vera chips were delivered to clients, including OpenAI, Anthropic, and SpaceX, in June. Nvidia is seen as the company that sets the direction for the information technology industry. But in CPUs, Nvidia is the challenger once again, competing against two well-established players in Intel and AMD, which have deep ties to hyperscalers and cloud giants. The company's investment in CPUs is another example of Nvidia's strategy to vertically integrate its systems and produce more of the chips and technology inside them every year. It aims to sell the entire system as a full rack of computing power, instead of simply selling chips by themselves. It's a strategy that Nvidia says will help engineers squeeze more performance out of their GPUs, helping the company's systems remain the computers of choice for leading AI labs as competition from AMD and custom chips heats up. Before the AI boom, the CPU was the most important and most expensive part in a server. The first generation of AI servers available when ChatGPT was released in 2022 paired as many as eight GPUs to one CPU, signaling a shift towards Nvidia's GPUs. But the rise of agentic AI, which can run independently in the background with minimal human input, has returned attention to the CPU, which is needed to babysit and feed data to an agent. Financial markets have noticed, and CPU incumbents AMD and Intel are among the two best-performing chip stocks so far in 2026, up 128% and 149% respectively, besting Nvidia's rise of 8%.
[4]
Nvidia has 'shipped hundreds of thousands of Grace standalone servers' -- GPU firm pivots messaging as CPUs take center stage in agentic data centers
Nvidia's Ian Buck, vice president of hyperscale and high-performance computing and the inventor of CUDA, says the company has "shipped... let's put it in the hundreds of thousands of Grace standalone servers." In May, Nvidia disclosed that it had shipped over 2.5 million Grace CPUs in total, and the company announced a partnership with Meta to deploy standalone Grace servers in February. Buck's comments suggest the scale of deployment may be even larger, however, as Nvidia tries to compete in a market dominated by other players. It's an interesting comment, though not a surprising one. Nvidia has become the dominating force of Silicon Valley as demand for its GPUs skyrocketed during an unprecedented data center buildout for AI inference. Since peaking earlier this year, however, around $1 trillion in Nvidia's market cap has been wiped away as investors rally behind CPU makers like Intel. Evolving agentic AI workloads have changed the hardware balance, shifting away from as many as eight GPUs per CPU, toward a one-to-one ratio in some cases. Nvidia wants to ride that train with its new Vera CPU, which was architected specifically for those types of workloads. Even before the recent rise of agents, however, Nvidia says it has seen demand for its CPUs for data-hungry workloads. "They weren't running a web server [with Grace]... or they aren't being used for, what the cloud uses, of cheap, dollar-per-core," Buck said. "They were being deployed for the backend, data-rich operations, like the data processing." Grace represents an on-ramp for Nvidia into data center CPUs. It uses 72 stock Arm Neoverse V2 cores, but it's differentiated by Nvidia's Scalable Coherency Fabric (SCF). Vera uses an updated SCF, but it also features Nvidia's first custom core design, called Olympus. Grace cracked the door, and Vera represents Nvidia's big entrance into the market against AMD and Intel. Regardless of where Vera ends up in the battle of next-gen data center CPUs -- which is heating up now, as AMD is expected to launch its Zen 6 Venice CPUs this week -- the design is vastly different from what we've seen out of Intel and AMD. Most notably, Vera is monolithic, placing all of its 88 cores on a single piece of silicon. AMD and Intel, years ago at this point, pivoted away from monolithic dies in favor of chiplets, allowing an extremely high density of cores at the cost of latency and coherency issues. Vera is radically different in that regard, not only being built on a single die, but also dedicating significant die space to the fabric. "One of the reasons we don't have 128 cores is because we've dedicated so much of the die area toward the fabric," Buck said. "It's 3.4 TB/s of bandwidth inside of that CPU that is dedicated toward allowing every core to talk to every cache, every memory [controller] at full speed without any collisions." For clarification's sake, Buck is referencing 3.4 TB/s of core-to-core bandwidth in Vera. There's up to 1.2 TB/s of aggregate memory bandwidth (14 GB/s per core) through the LPDDR5X interface. But just as chiplet-based designs made trade-offs in per-thread performance, Vera will likely make trade-offs for its unique architecture. The majority of data center workloads are still "legacy" tasks that hyperscalers have built for, and even with seemingly insatiable demand for AI infrastructure, that is unlikely to change for several years. Buck recognizes this trade-off, asking: "Can Intel and others build rich fabrics? Do they have the IP and the ecosystem to do it and connect it all the way through to LP memory? They need to tell you when they're going to do it... but that trade-off will come at the cost of the legacy workload." Earlier this year, at GTC in March, Buck was even more clear. "The world is not going to be served by one SKU of CPU, and that is not our intention," the executive said in a news conference at the time. Still, it's clear Nvidia has ambitions with data center CPUs beyond what headlines are floating around on the New York Stock Exchange. Nvidia says CPUs represent a $200 billion TAM (Total Addressable Market) opportunity for the company, a rather rosy forecast compared to the rest of the industry, which sees a TAM of around $120 billion by 2030 (though recent estimates have climbed as high as $170 billion). And agentic AI is expanding that market, with Morgan Stanley in April estimating that agents could add as much as $60 billion to the data center CPU market. Vera is in full production alongside Nvidia's next-gen AI infrastructure, including Rubin GPUs, ConnectX-9 NICs, SpectrumX Ethernet switches, and the various components that go into building a Vera Rubin NVL72 rack. The company says there are around 1.3 million components that go into a rack, and it has a list of over 300 partners globally to build them. As part of our visit to Nvidia HQ last week, we saw a Vera Rubin NVL72 rack in action, running workloads for OpenAI. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[5]
NVIDIA Vera CPU Rolls Over Traditional x86 Chips With An Agentic-AI Focused Design With 6x Faster Performance & 40% Lower Latency
NVIDIA's Vera CPU trounces x86 chips, such as AMD's EPYC, in Agentic AI workloads, delivering a 6x performance gain with 40% lower latency. NVIDIA Vera's Monolithic CPU Design & Architectural Changes Achieve Over 2x Performance Gain Over x86 CPUs Such As AMD EPYC With Just 88 Cores, Beating 128 Cores Today, NVIDIA is going all out by sharing its full performance disclosure of the Vera CPU against x86 offerings. We just talked about the architecture in detail over here, and now, we will look at the performance capabilities that this chip has on offer. NVIDIA is evaluating Vera across four key categories that matter the most in the Agentic AI era: * Memory Bandwidth and Latency - Shows how Vera keeps cores supplied with data through high-bandwidth memory, low-latency access, and a high-throughput coherent fabric. * Application Workload Performance - Demonstrates performance across agentic benchmarks, data analytics, graph processing, and other CPU-intensive workloads. * Core IPC - Measures the impact of the Olympus microarchitecture, including the wide Front End, advanced branch prediction, deep out-of-order execution, and expanded execution resources. * RL and Agentic AI Performance - Highlights Vera's value for latency-sensitive, branch-heavy, and highly concurrent agent environments. But before starting with the benchmarks, we want to state that NVIDIA has heavily used the AMD EPYC Turin "Zen 5" chip in its comparison, highlighting the monolithic & NUMA nature of its chip & how it significantly elevates its position over x86 chips. Architectural Benchmarks NVIDIA has divided the benchmarks into two categories, one that focuses on the architectural side of things and the others at how the architectures improve AI workflows. As such, the first architectural benchmark showcases Vera's IPC uplifts across various workloads that stress large instruction footprints, dense control flow, compiler and runtime behavior, and long dependency chains. Here, Vera offers up to a 1.9x uplift over AMD's Zen 5 (EPYC Turin), showcasing its strong single-threaded gains, a key driver for Agentic AI and RL apps. The Branch Prediction on Vera's Olympus cores is up to 2.3x faster and 2x faster on average versus AMD's Zen 5. The Taken-branches per cycle see a 3.5x uplift. Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle. The Olympus core features a 64 KB instruction cache with a high-bandwidth fetch of 128 bytes per cycle out and a 10-wide decode without relying on an x86-style uOP cache, which offers up to a 2.4x gain over Zen 5 in Instruction Fetch Ops per cycle. And lastly, we have the Backend Ops per cycle, which are up to 4.3x faster and over 3x faster on average versus Zen 5. Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle. AI Workload Benchmarks The first benchmark is centered around memory latency. Here, NVIDIA states that traditional x86 CPUs these days rely on chiplet architecture. The chiplet design leads to shortcomings in the memory subsystems when many requests are made across cores. These cause bottlenecks to appear in the memory subsystem and the coherency fabric, and as a result, additional hops increase latency and reduce effective bandwidth, leading to performance drawbacks. In the comparison, the AMD EPYC Turin 9755 CPU can be seen jumping from similar latency in the beginning (~120ns) up to 350ns as soon as memory bandwidth demand increases. Vera, on the other hand, can deliver 3x the bandwidth with 40x lower peak latency at a 90% memory utilization rate. NVIDIA's Vera CPUs also offer much higher memory bandwidth per core. The Turin chip has 128 cores, with each core being fed roughly 3.1 GB/s of bandwidth. Vera has 88 cores, and each core is fed with around 12.7 GB/s of bandwidth, 4x higher than EPYC, and the per-core memory bandwidth is achieved up to 14 GB/s in certain use cases. Such per-core memory bandwidth characteristics are important for Agentic AI workloads where many threads operate on large, irregular data sets and can quickly become limited by memory throughput rather than compute resources. Further impacting the performance of chiplet-based solutions is the addition of dozens of separate I/O and compute chiplets. When data moves from one core to another, it may need to leave the source core die, traverse the package fabric through an I/O die, reach a different core die, and then follow a similar path back for coherency responses. Each die hop adds latency, consumes fabric bandwidth, and creates variability depending on where the communicating cores are physically located. The result is higher core-to-core latency & impacts performance. In the core-to-core latency heat map, you can see that NVIDIA Vera delivers 50% lower latency than the AMD EPYC 9755 chip. What's even more impressive is that Vera shows consistent latency across all of its core pairs, further emphasizing how its monolithic architecture is made from the ground up for AI workflows. Next up, we move to Agentic benchmarks, which are the core reason why Vera was made in the first place. In Spec CPU 2006, NVIDIA measures the per-core performance for a fully loaded socket system. The software stack here measures performance across Python execution, code compilation, static analysis, and tool-driven software workflows. Fully loaded per-core performance captures how well each core sustains throughput while sharing socket-level power, memory bandwidth, cache, and fabric resources. The NVIDIA Vera CPU delivered up to 1.8x higher performance in Python versus AMD's EPYC. The rest of the benchmarks show up to 1.7x performance; that's almost 2x the performance bump vs x86 competitors. In Graph Traversal algorithms, NVIDIA's Vera CPU offers a 2.6x performance uplift over AMD Turin. Graph workflows are bandwidth and cache-sensitive, and here, NVIDIA's memory and cache subsystems inside Vera are showcasing displaying their full potential. NVIDIA is also showing off its Graph Traversal Core scaling in Page Rank workloads, which also stresses the core pipeline, fabric, and memory bandwidth. Here, Vera achieves a monumental 29.3x gain at a core count of 32. Lastly, we have ClickHouse, which is a high-performance columnar database used for online analytical processing. Vera achieves a 20% uplift over the AMD EPYC Turin chip and shows it to be a great fit for scan-heavy, aggregation-heavy, and concurrency-sensitive analytics and ETL (Extract-Transform-Load) workloads. NVIDIA's Vera CPU represents a significant architectural leap forward for high-performance computing, particularly in the demanding realm of Agentic AI workloads. By embracing a monolithic design rather than the chiplet-based approach common in modern x86 CPUs, Vera overcomes critical bottlenecks in memory subsystems, coherency fabric, and inter-core communication. This results in dramatically lower and more consistent latency, while delivering substantially higher memory bandwidth per core. Gunning A $200b TAM Oppurtunity The NVIDIA Vera CPU is aiming to become a highlight of the Agentic AI CPU race. The company has already said that Vera is in a position to become the leading CPU supplied in 2026, and there are plenty of early adopters that have started using Vera, such as OpenAI, Anthropic, SpaceX, & Perplexity. The CPU is being deployed across various supercomputer centers and cloud data centers, and will see full-fledged support from countless ecosystem partners. Optimized from the ground up for irregular, bandwidth-hungry, and concurrency-intensive AI workflows, Vera not only sustains higher per-core throughput under full load but also sets a new standard for efficient, scalable processing in AI datacenters. Follow Wccftech on Google to get more of our news coverage in your feeds.
[6]
NVIDIA Vera CPU Is Architected For The Agentic AI Era, as It Delivers Max Single-Core & Single-Thread Performance Versus x86; Full Architectural Breakdown Shows
The NVIDIA Vera CPU is designed for Reinforcement Learning and Agentic AI systems, offering 88 Olympus cores for max single-core performance. NVIDIA's x86-Killer, Vera CPU, Is Now Available, Built Purely With Agentic AI In Mind Surge in agentic AI workloads and reinforcement learning emerging as a key scaling mechanism for improving model capability have made CPUs a vital component for AI systems. Without a fast CPU (execution, evaluation, orchestration), a GPU won't be able to be supplied with context, and that leads to major drawbacks and inefficiencies. This is why NVIDIA designed Rubin, a CPU that is purpose-built for agents, and a chip that eliminates these bottlenecks, delivering the throughput and efficiency that the industry demands. Today, we will give you a full rundown of Vera, NVIDIA's latest and greatest CPU for AI, based on custom Arm cores. The Insides of The Vera CPU On a high-level, NVIDIA's Vera CPU is based on the Olympus Core architecture, which utilizes the custom Armv9.2 IP. It is based on a monolithic compute die and comes with technologies such as 2nd Gen NVIDIA Scalable Coherency Fabric (SCF), Confidential Computing (TEE-I/O Capable), and x16 PCIe Gen6 CXL 3.1. Key benefits of the NVIDIA Scalable Coherency Fabric include: * Low-latency communication between CPU clusters * Efficient access to shared system-level cache resources * High-bandwidth connectivity to memory controllers * Native integration with NVLink-C2C and accelerator infrastructure The Brand Prediction features Neural & Special-Purpose Predictors with 2 taken branchers per cycle. It has a 10-Wide Instruction Decode, a Novel Graph and Advanced Prefetcher, 96 KB of L1D, 64 KB of L1I, 2 MB of L2, and 164 MB of L3 cache. Each core has access to 14 GB/s of bandwidth. The NVLink C2C interconnect offers up to 1.8 TB/s of coherent CPU-GPU bandwidth. In total, there are 88 Olympus "high-performance" cores and 176 NVIDIA Spatial Multi-threading threads on Vera. The chip integrates a wide, high-IPC microarchitecture that is designed to offer maximum single-threaded performance, achieving efficient concurrency across thousands of software environments. Olympus - The AI Peak Olympus, the core powering Vera, is referred to as the compute engine of the CPU. It is built from the ground up with one goal: to maximize IPC (Instructions Per Cycle) for irregular, branch-heavy, latency-sensitive, and highly concurrent software. Versus Grace, Vera achieves a 50% gain in IPC. There are four key components of the Vera CPU architecture: * Front End: Optimized Branch Predictors * Mid Core: Critical Path Acceleration * Execution Engine: Complex Vector Optimization * Cache Subsystem: Latency-Optimized Instruction and Data Paths The following is the block diagram of the Olympus core: Olympus Neural Branch Predictor Starting with the dissection of the Olympus core, we first have the Neural Branch Predictor, which is designed to improve prediction accuracy on complex, branch-heavy software. The NBP learns longer-range control-flow behavior and statistically biased branch patterns that are common in large runtime environments, interpreters, databases, compilers, and agent workloads. The branch prediction improvements allow Olympus's front end to be supplied with useful instructions, which in turn enable higher instruction throughput and more consistent performance. Olympus Mid-Core The Mid Core on Olympus turns decoded instructions into executable operations while maxing instruction-level and memory-level parallelism. The Olympus Mid Core is designed to identify and execute independent work hidden behind these bottlenecks, enabling higher utilization of execution resources and improved overall IPC. Mid Core comes with a wide rename and allocation engine that is supported by a large reorder buffer and extensive physical register resources. The Mid Core Register renaming eliminates false dependencies between instructions, allowing the processor to expose additional parallelism and execute instructions as soon as their true operands become available. A large instruction window enables Olympus to look significantly further ahead in the instruction stream, finding useful work while other instructions wait on memory or branch resolution. Olympus also houses several technologies to break dependencies, such as: * Memory Renaming: Accelerates store-to-load dependency chains by allowing dependent instructions to execute before a load completes when the data relationship can be predicted or determined. This is particularly beneficial for pointer-heavy software, graph traversal, runtime frameworks, and complex object-oriented workloads. * Value Prediction: Identifies stable dependency chains and predicts future values before they are produced, allowing dependent instructions to execute speculatively while correctness is verified later. This capability is especially effective for repetitive software patterns and sequential data structures. * Move elimination and value optimization: Techniques to achieve higher utilization. Olympus Execution Engine The execution engine within the Olympus core is designed to sustain high instruction throughput across various workloads, including Agentic AI runtimes, RL environments, databases, and more. After the instructions pass through the rename and scheduling stages, they are dynamically dispatched to specialized execution resources optimized for integer arithmetic, branch processing, vector computation, and memory operations. The combination of a wide execution backend and deep out-of-order engine allows Olympus cores to continue making progress even when some parts of the workloads are slowed down by memory access, branch resolution, or long dependency chains. So for the Execution Engine, NVIDIA packs 8 simple Integer ALUs, 2 Complex ALUs, and 4 dedicated branch execution units. The simple ALUs handle common arithmetic and logical operations, while the complex ALUs accelerate higher-latency operations such as multiplication, division, CRC, and shift-intensive workloads. The dedicated branch units work in conjunction with the advanced branch prediction subsystem to rapidly resolve control-flow decisions and minimize pipeline disruption. This combination is particularly beneficial for modern agentic workloads, which often exhibit irregular control flow and frequent branching as agents reason, evaluate state, and execute actions. For vector and AI-oriented computation, Olympus integrates six 128-bit SVE vector execution units, including two crypto-capable vector pipelines, and supports FP8 precision. These units accelerate vectorized workloads commonly found in data processing, machine learning preprocessing, compression, encryption, scientific computing, and analytics frameworks. Support for Arm Scalable Vector Extension (SVE) enables efficient execution of both traditional HPC workloads and emerging AI data-processing pipelines. The Execution Engine is tightly integrated with a high-bandwidth memory subsystem through four load pipelines and 2 store pipelines. Together, these resources enable Olympus to maintain high utilization across both compute-intensive and memory-intensive workloads. Vera's Spatial Multi-Threading Offers Higher Single-Threaded Throughput Than Traditional SMT Approaches NVIDIA is making use of a new multi-threading solution in Vera CPUs called Spatial Multi-Threading. Compared to traditional SMT (Simultaneous Multi-Threading), the NVIDIA approach makes it so that resources can be partitioned across two hardware threads inside the Olympus core instead of relying primarily on a narrower core. Traditional SMT achieves this by allowing both hardware threads to opportunistically share the same core resources. In densely threaded sandbox environments, that sharing can create contention across branch prediction, decode, execution, load/store, cache, and memory resources. This creates a "noisy neighbor" effect, where activity on one thread can reduce the performance of the other thread sharing the same core within a sandbox. For agentic AI, reinforcement learning, and cloud services -- where many short-running tasks execute concurrently and tail latency matters -- this can increase variability and make performance less predictable. With this, Vera offers a high-throughput single-threaded core when needed, while the thread duo handles management. With 88 cores and 176 threads, NVIDIA's Vera is a flexible CPU that offers both high single-threaded performance and dense vGPU-style deployments. Spatial Multithreading improves determinism, isolation, and quality of service compared to traditional SMT approaches. The result is a CPU architecture that can run large numbers of concurrent agent tasks while maintaining more consistent latency and throughput. The Vera Memory Subsystem - LP5X Lowers Power Input Versus DDR DIMMs NVIDIA went with LPDDR5X memory for its Vera CPUs by fusing the chip with eight controllers, offering up to 1.2 TB/s of bandwidth. The LPDDR5X memory comes with data center-class ECC (Error-Correcting Code) and enables the platform to significantly reduce power consumption vs traditional DDR4-based servers. The chip enables up to 1.5 TB of LPDDR5X capacity. The use of LPDDR5X memory in the SOCAMM2 form factor also plays a vital role in driving up power efficiency, while the three primary objectives for this memory type include: * High Bandwidth - Deliver the memory bandwidth required to sustain modern AI, analytics, and HPC workloads. * Power Efficiency - Maximize bandwidth per watt using LPDDR5X technology to reduce overall platform power. * Enterprise Serviceability - Provide a modular, field-replaceable memory architecture suitable for hyperscale and enterprise deployments. The LPDDR5X memory on Vera operates at 9600 MT/s, offering up to 14 GB/s of memory bandwidth to each Olympus core. The memory capacities are fully scalable from 256 GB up to 1.5TB. Now, the use of LP5X also has a major reason, and that is to drive down the power envelope. Traditional DRAM standards such as DDR5 (DIMM, RDIMM, MRDIMM) do offer higher bandwidth, but they also use much higher power. RDIMM platforms can consume over 100W, while MRDIMM solutions can use over 200W of power in bandwidth-heavy workloads. SOCAMM2, on the other hand, achieves up to 1.2 TB/s speeds while consuming just 30-40W, and that's for a fully-loaded system. The IOs - PCIe 6.4, 2nd Gen NVLink, Dual-Socket Scale, and Confidential Computing On the IO, Security and Scale-Up front, NVIDIA Vera CPUs offer lots of technological innovations such as the following: * Second-Generation NVLink-C2C - High-bandwidth coherent interconnect for CPU-to-GPU and CPU-to-CPU communication. * Two-Socket Scale-Up - High-speed coherent processor-to-processor connectivity for scalable dual-socket systems. * PCI Express 6.4 - High-performance I/O connectivity for storage, networking, and accelerator devices. * Compute Express Link (CXL) 3.1 - Extends the PCIe infrastructure to support memory expansion, pooling, and composable system architectures. * Confidential Computing - Secure scale-up environment across 2-socket and NVL rack. Starting with PCIe, Vera leverages the latest PCIe 6.4 standard with each chip offering up to 88 PCIe lanes, while dual-socket configurations offer 176 lanes. These operate at speeds of up to 64 GT/s, doubling the throughput of the PCIe 5 standard, and ensure enough bandwidth for AICs and storage solutions. Vera also comes with a wide array of PCIe lane bifurcation, as detailed in the table below: Further extension to CXL 3.1 enables next-gen memory expansion and composable infrastructure with the support of CXL Type-3 memory devices beyond the standard LPDDR5X solution. Vera supports Arm CCA, RME-DA, and RME-CDA, providing the foundation for confidential VM isolation and secure device assignment. Per-VM encryption keys enable cryptographic isolation between tenants running on the Vera CPU, while support for TDISP-capable devices extends the trusted execution environment beyond the CPU. As the industry's first Confidential Computing CPU to enable TDISP for coherent devices such as GPUs, Vera provides high-performance, coherent access without bounce buffers or unnecessary data copies while maintaining strong workload isolation. On the security, reliability, and availability front, Vera offers the following: NVIDIA's Vera CPU represents a groundbreaking advancement in processor design, purpose-built to power the next era of agentic AI and reinforcement learning workloads. Leveraging the innovative Olympus core architecture with custom Armv9.2 technology, Vera delivers exceptional single-threaded performance through a high-IPC microarchitecture featuring advanced neural branch prediction, sophisticated mid-core parallelism techniques like memory and value prediction, a powerful execution engine with vector capabilities, and an optimized cache subsystem. Its Spatial Multi-Threading approach ensures superior determinism and isolation compared to traditional SMT, while the integration of high-bandwidth LPDDR5X memory, second-generation NVLink-C2C, PCIe 6.4, CXL 3.1, and comprehensive confidential computing features enables seamless CPU-GPU orchestration, massive scalability, and enhanced security. By eliminating traditional bottlenecks in execution, evaluation, and data movement, Vera achieves up to 50% better IPC over its predecessors, offering unprecedented efficiency, throughput, and low-latency performance tailored for complex, concurrent AI environments -- positioning it as a true x86 alternative and a cornerstone for upcoming AI systems. Follow Wccftech on Google to get more of our news coverage in your feeds.
[7]
Nvidia Unveils Vera to Speed Expansion in Server Processor Market
With Vera, Nvidia is betting on an architecture optimized for artificial intelligence applications, particularly autonomous agents that require rapid communication between CPUs and GPUs. The company says its processor delivers up to 50% higher performance than competing x86 chips on these workloads, thanks to an architecture that prioritizes per-core speed, memory bandwidth and lower latency. Vera will be sold as a standalone chip or integrated into different compute racks, notably alongside GPUs from the Vera Rubin platform. The push comes as the server CPU market regains importance with the rise of artificial intelligence. Nvidia is seeking to broaden its footprint beyond graphics processors, even as AMD and Intel retain strong legacy positions in the segment. Several analysts believe Vera could form a new category of processors geared to the most demanding AI uses, but its success will hinge on adoption by major cloud service providers, a market Nvidia is now targeting with a fully integrated offering. Nvidia shares are up 1.6% today.
Share
Copy Link
Nvidia has unveiled comprehensive details about its Vera CPU, marking a bold entry into the data center processor market dominated by AMD and Intel. Built with a custom Olympus architecture and monolithic design, Vera delivers up to 6x faster performance in agentic AI workloads. The chip is already shipping to OpenAI, Anthropic, and SpaceX as Nvidia pushes to control every layer of AI infrastructure.
Nvidia has released detailed specifications and benchmarks for its next-generation Vera CPU for AI, establishing a direct challenge to AMD and Intel in the lucrative data center chip market
3
. The GPU giant, which became the world's most valuable company on the strength of its graphics processors, is now positioning itself to capture a larger share of AI infrastructure by supplying both CPUs and GPUs in integrated systems1
.
Source: Wccftech
The Vera CPU features Nvidia's first custom core design, called Olympus architecture, with 88 cores built on a monolithic die rather than the chiplet designs favored by competitors
2
. This architectural choice reflects Nvidia's focus on agentic AI workloads, which require low latency and high memory bandwidth to manage autonomous agents running in the background with minimal human input. Ian Buck, Nvidia's vice president of accelerated computing and the inventor of CUDA, emphasized the company's commitment: "We're on a roadmap to crank out new architectures, not just GPUs but CPUs"1
.Nvidia shared unofficial SPEC CPU 2026 results comparing Vera against AMD's EPYC 9755 Turin processor in dual-socket configurations
2
. Despite having fewer cores than AMD's 128-core design, Vera achieved a 3% advantage in overall throughput. More striking are the gains in agentic AI workloads, where Vera delivers up to 6x faster performance with 40% lower latency compared to traditional x86 chips5
.The Olympus architecture achieves up to 1.9x IPC uplift over AMD's Zen 5, with branch prediction up to 2.3x faster on average
5
. Nvidia dedicated significant die space to its Scalable Coherency Fabric, which provides 3.4 TB/s of core-to-core bandwidth and up to 1.2 TB/s of aggregate memory bandwidth through LPDDR5X interfaces4
. Each core receives roughly 12.7 GB/s of bandwidth, 4x higher than EPYC's per-core allocation5
.Vera is the CPU component of Nvidia's Vera Rubin platform, which combines CPUs and GPUs into integrated systems designed for plug-and-play deployment in AI data centers
1
. The Vera Rubin NVL72 system pairs 36 Vera CPUs with 72 Rubin GPUs in a single liquid-cooled rack, offering one CPU for every two GPUs. Nvidia representatives confirmed that Vera chips were delivered to clients including OpenAI, Anthropic, and SpaceX in June, with OpenAI already operating one Vera Rubin rack3
.
Source: Wired
The system processes ten times as many tokens per watt as the previous Grace Blackwell superchip and offers nearly three times as much memory bandwidth
1
. Nvidia has also significantly reduced cabling requirements, touting Vera Rubin as "cable-free compute" that can be installed in minutes rather than hours. The 100 percent liquid cooling approach reduces energy consumption compared to air-cooling systems. Microsoft and Oracle are among the early customers expecting shipments in the second half of this year .Related Stories
The rise of agentic AI has fundamentally altered the hardware balance in the AI server market. Early AI servers paired as many as eight GPUs to one CPU when ChatGPT launched in 2022, but agentic systems now require ratios closer to one-to-one
3
. This shift has propelled AMD and Intel stock prices up 128% and 149% respectively in 2026, outpacing Nvidia's 8% gain as investors recognize the expanding role of CPUs in latency-sensitive AI tasks3
.
Source: Tom's Hardware
Nvidia estimates CPUs represent a $200 billion total addressable market opportunity, higher than industry forecasts of $120 billion to $170 billion by 2030
4
. Morgan Stanley estimated in April that agents could add as much as $60 billion to the data center CPU market. Buck revealed that Nvidia has "shipped hundreds of thousands of Grace standalone servers," with over 2.5 million Grace CPUs delivered in total as of May4
. Meta deployed standalone Grace servers in February for data-rich backend operations.Nvidia's detailed disclosure arrives just before AMD's annual conference, where the company is expected to unveil next-generation AI and data center chips including Zen 6 Venice CPUs
1
. AMD revealed its Helios AI chip rack on Sunday, designed to compete directly with Vera Rubin. Nvidia remains sensitive to any suggestion of delays after Blackwell chips reportedly overheated in customized server racks, forcing design changes and shipment delays1
.The monolithic die approach distinguishes Vera from competitors who adopted chiplet designs years ago for higher core density. Buck acknowledged the trade-offs: "The world is not going to be served by one SKU of CPU, and that is not our intention"
4
. While Vera excels at agentic AI workloads, legacy cloud workloads still dominate data center operations. Nvidia is also selling Vera as a standalone product, with reports indicating Chinese customers could receive shipments as soon as August1
. The company's vertical integration strategy aims to maintain its position as AI labs face increasing competition from AMD and custom chip designs.Summarized by
Navi
[2]
[4]
26 May 2026•Technology
26 Jan 2026•Technology
08 Jul 2026•Technology

1
Technology

2
Science and Research

3
Policy and Regulation
