14 Sources
[1]
Nvidia Wants to Own Every Chip Inside AI Data Centers
Nvidia's Vera Rubin platform combines CPUs and GPUs into a single system, reflecting the company's growing ambition to power every layer of AI infrastructure. Nvidia is hyping up its new Vera Rubin chip system this week, revealing new performance benchmarks for the GPU and CPU combo ahead of rival AMD's annual product event in San Francisco on Thursday. During a lengthy technical workshop last week at the company's headquarters in Santa Clara, California, Nvidia executives boasted to a small group of journalists about the chip system's increased power and efficiency capabilities. The biggest takeaway: Nvidia, which has long specialized in making GPUs, is increasingly trying to position itself as a supplier of CPUs that can power AI agents. While GPUs are still the main hardware that companies use to train and run their AI models, the industry's shift toward more complex, agentic systems has increased demand for CPUs, which can orchestrate data flows, networking, and other software tasks. That's likely one reason Nvidia has been eager to position itself as a supplier of complete AI systems, rather than just AI chips. Vera Rubin is Nvidia's successor to its hybrid superchip system Grace Blackwell, and represents the linchpin of its near-term future powering the AI industry. It's designed to offer one CPU for every two GPUs. In a single Vera Rubin NVL 72 super chip system, there are 36 Vera CPUs for every 72 Rubin GPUs. Nvidia is also selling the Vera CPU as a standalone product, and it has reportedly told Chinese customers these could be ready as soon as August. Nvidia executives emphasized that its new Vera Rubin NVL72 racks -- a stack of chips packed into a single liquid-cooled platform -- are much more "plug-and-play" than some of its earlier products. During a brief tour of a Nvidia data center lab in Silicon Valley, Nvidia executives shared that OpenAI already has one Vera Rubin rack in use. Nvidia CEO Jensen Huang didn't make an appearance at the workshop in Santa Clara last week; he was in Japan announcing the chipmaker's new partnerships with a number of Japanese firms to develop AI for robotics. The briefings were instead led by Ian Buck, Nvidia's longtime vice president of accelerated computing and the architect behind the company's CUDA software. "We're on a roadmap to crank out new architectures, not just GPUs but CPUs," Buck told reporters. "We're going to keep innovating, because it's do this or die in Silicon Valley." The meetings were held in Huang's executive briefing center, where multiple desks nearby were piled with bags of Taiwanese snacks that the CEO brought back from his recent trip to Computex, a massive annual semiconductor trade show in Taipei, an Nvidia spokesperson told WIRED. Nvidia claims that the Vera Rubin NVL72 system will process ten times as many tokens per watt as the company's Grace Blackwell super chip. The company says that its Vera CPU is also faster at processing agentic AI tasks compared to rival CPUs from AMD and Intel (though the tests it ran to support those benchmarks appears to have used slightly older generations of its competitors' CPUs). Localized memory subsystems on the new chips will also offer nearly three times as much memory bandwidth as Blackwell, which will likely be an appealing feature to many companies amid an ongoing shortage of high band-width memory. Nvidia says it has also significantly reduced the number of cables needed to connect its chips to racks in multi-rack server systems, to the point where the company is touting Vera Rubin as "cable-free compute" and "hot-swappable." This means customers can theoretically reduce the amount of time it takes to install each rack from a couple of hours to a few minutes, a point that was brought up by both Buck and Andrew Bell, Nvidia's senior vice president of hardware engineering. And the new chip system is 100 percent liquid-cooled, which can reduce the amount of energy needed to cool the chips, since air-cooling is more energy intensive. Ever since Nvidia unveiled Vera Rubin in the spring of 2025, the company has been slowly dribbling out more details about the chip system while insisting it will be released on schedule. Huang has repeatedly said Vera Rubin is ramping to "full production" and will ship in the second half of this year, with early customers including Microsoft, OpenAI, and Oracle. Nvidia is particularly sensitive to any suggestion of delays after its previous-generation Blackwell chips reportedly overheated when connected together in the company's customized server racks, forcing it to make design changes and push back shipments. Nvidia's marketing push for Vera Rubin is happening just ahead of rival AMD's annual conference, where executives are expected to tout its next-generation AI and data center chips. On Sunday, AMD revealed more details about its Helios AI chip rack, which is designed to compete with Nvidia's new wares. Both AMD and Nvidia have been vying for large-scale, multi-year contracts to supply AI hyperscalers like Meta and Amazon and AI labs like OpenAI, Anthropic, and SpaceXAI with chips. Over the past two years, AMD has significantly grown its share of the market for CPUs used in data centers. The company has long been recognized as a pioneer of the modern chiplet architecture used in x86 processors, which still account for the vast majority of data center CPU revenue. Nvidia, by contrast, builds its data center CPUs on ARM, an alternative chip architecture known for its power efficiency. Nvidia executives Buck and Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, both emphasized that Vera Rubin abandons the chiplet architecture used by many modern processors in favor of a single, monolithic chip. Coutand argued that stitching together multiple chiplets imposes "a heavy tax on memory bandwidth and data movement," whereas the monolithic design of Vera Rubin allows data to move more quickly across a single integrated circuit.
[2]
Nvidia deep dives Vera CPU for AI data centers -- SPEC CPU 2026 benchmarks revealed, Olympus architecture specifics, and more
Nvidia's Vera CPU is its first bid to become a key player in the data center CPU market. Although Grace has seen some success (most notably with Grace standalone deployments at Meta), Vera is Nvidia's first CPU with a custom core design. It's arriving at an ideal time, as well, with the server CPU market exploding in the last few months on the back of agentic AI demand. Vera isn't a chip built to chip away at the market share of AMD and Intel in the cloud. It's built to grab market share in an expanding market, as hyperscalers look to widen AI infrastructure beyond legacy clouds. As such, it's designed in a much different way than Nvidia's x86 competitors, and it even holds some unique architectural design points compared to the swath of Arm-based designs. Nvidia has slowly revealed more details about Vera as it ramps into general availability, which is on track for the back half of this year. Now, we have a full picture of the chip. Nvidia shared its Vera white paper, along with unofficial SPEC CPU 2026 results comparing Vera to AMD's Turin-based Epyc 9755. We're going to break down the white paper here, including all of the details about the Olympus core and a look at the benchmarks Nvidia ran. At the end of this piece, we'll also take a brief look at the larger context of Vera and how it fits into Nvidia's wider AI ecosystem compared to standalone deployments. But plenty of ink has been spilled about Vera's technical capabilities and Nvidia's next-gen AI infrastructure vision. Let's start with the important thing: the benchmarks. Nvidia Vera CPU benchmarks We've seen Vera in action before, though only through a series of selected benchmarks ran at Nvidia HQ by Phoronix. In the Vera white paper, Nvidia shared benchmarks for SPEC CPU 2026, specifically the integer suite from SPECrate, against AMD's Epyc 9755, with both chips running in a dual-socket configuration. Before getting into the results, there are some important notes about how SPEC runs work, and the reporting criteria for them. Nvidia's run here isn't official, as Vera was tested in a reference system due to the fact that it's not broadly available yet. It's ramping for general availability in the second half of the year. Due to that, Nvidia is unable to report its results. That's why you see "estimated" in some of the charts below. Nvidia ran SPEC CPU 2026; it's not extrapolating expected performance like we've seen from AMD so far with its upcoming Venice chips. SPEC CPU 2026 is split into four suites, but Nvidia tested the SPECrate integer suite, which is focused on system throughput with integer-based workloads. The "rate" result is looking at how much work is completed within a certain amount of time. Here, each thread in the system has a copy of the workload. The score is how much time it takes for those workloads to complete, regardless of thread count, naturally giving chips with more cores an advantage. If you want more detail on the benchmarks included in the suite, make sure to read our original coverage of SPEC CPU 2026. Here are the overall results: Nvidia didn't share the exact results for the 9755 it tested, short of the overall score of 898. Taking that overall score into account, Vera is 3% ahead of the 9755. It's worth noting that Vera is ahead here despite a large thread disadvantage. An overall score of 898 for a dual-socket Epyc 9755 system isn't unreasonable compared to publicly-submitted SPEC CPU 2026 runs, though higher results have been published. SPEC CPU ships as source code, which the tester must compile with their compiler of choice, and that can heavily influence results (particularly with vendor-specific compilers). Nvidia used GNU 15.2 with both systems. Above, you can see Vera's results stacked up against the 9755, but these aren't comparing the numbers directly. Nvidia has normalized the per-core performance, which isn't how SPECrate results are normally shared. According to the overall numbers, Vera is still completing more work within the same amount of time, despite a thread disadvantage, but the margins aren't in the range of a 70% or 80% advantage as the above chart suggests. We asked Nvidia about the results given that they're obfuscated by comparison; we could not reverse-engineer the Epyc 9755's scores with the information Nvidia has provided. Here's the response it gave: "Per-core performance under a fully loaded socket is important because agentic AI and RL run many sandboxes concurrently, while each agent step remains sequential and latency-sensitive. It measures how much performance each core sustains amid contention for shared power, memory, cache, and fabric. We therefore normalize by physical core, with SMT enabled on both systems." The "agentic" workloads Nvidia has highlighted here are code compilation and interpretation workloads, which is something an agent is often doing, querying repos for dependencies and building source code. Below are data science workloads (or Exploratory Data Analysis), and below that are data processing workloads like SQLite database management. The results here align with Nvidia's overall messaging of Vera, that it's highly competent at data-rich, backend operations. Although Nvidia is sharing per-thread results, it argues that SPECrate is still the correct benchmark to run. The per-thread results here are in the context of a fully-loaded socket. Here's the justification from the white paper: "This metric is non-trivial for agentic AI and RL systems, where many sandboxes, tools, and environments run concurrently rather than as isolated single-thread tests. Fully loaded per-core performance captures how well each core sustains throughput while sharing socket-level power, memory bandwidth, cache, and fabric resources." In addition to running the workloads, Nvidia analyzed the code execution for architectural benchmarks, which you can see in the gallery above. Nvidia claims an overall IPC gain of up to 1.9x compared to Turin, up to 2.3x more branch predictions and 3.5x taken branches per cycle, and up to 2.4x higher instruction fetch operations per cycle. Outside of SPEC, Nvidia shared a few benchmarks highlighting the capabilities of the Olympus core. First up is PageRank, an algorithm developed by Google to originally rank web pages, which highlights Olympus' prefetch engine. Nvidia scaled this workload to higher core counts, showing Vera maintaining much of its single-core performance up to 32 cores, while the Turin chip hits a wall around 20 cores. In addition to the above results, Nvidia shared some tests of the Vera memory system compared to Turin. These microbenchmarks are good for validating Nvidia's specifications, but they're looking at architectural performance, not application performance. An architectural advantage translates into a performance advantage, but not always in a linear, expected fashion. Nvidia used internally-developed tools for the memory tests, though they're available on GitHub for anyone to run. First is loaded memory latency, stressing the memory subsystem as bandwidth usage increases. Vera has much higher bandwidth overall, but you can see the Turin chip hit a latency wall below its maximum, which Nvidia attributes to Non-Uniform Memory Access (NUMA) domain traversal and CCD-to-CCD latency. Looking at per-core bandwidth, Nvidia claims Vera provides more than four times the bandwidth of AMD's 9755. The suggestion here is that "real-world" per-core bandwidth is even better than Nvidia's specs lead on (or perhaps worse than AMD's). Maybe the most consequential of these tests is the one you can see above, looking at core-to-core latency. It's no secret that crossing the CCD on AMD's chiplet-based architecture incurs a big latency penalty. You can see that in action even in our Ryzen 9 9950X3D2 review, and the penalties compound as you scale up the number of CCDs. In fairness to AMD here, chiplet-based designs aren't built for this type of cross-CCD traversal, preferring to keep workloads localized and optimizing for core density. Vera's design goal is clearly to keep latencies consistent across the entire die and sacrificing core density in the process. Nvidia's Ian Buck told us that this design trade-off "will come at the cost of the legacy workload," when we recently visited Nvidia HQ. That's important context. Nvidia isn't gunning to steal existing market share from AMD and Intel as much as it's trying to grab market share in an expanding market before AMD and Intel can. Some financial institutions (including Morgan Stanley and Bank of America) suggest the server CPU market could double in size (or grow even larger) by 2030. That context is important because there will be a continuing demand for CPUs that can handle workloads Vera is not optimized for, and it'll be interesting to see how AMD and Intel tackle that dynamic with future products, trying to keep a legacy base of customers while pushing ahead into the expanded market. Nvidia clearly has a vision of how that expanded market looks, and to that end, hasn't shared SPEC CPU floating point results. Presumably, this is due to the fact that SPEC's vectorized suite is focused primarily on HPC workloads, whereas Nvidia focused on what it believes are critical agentic workloads that are integer-based. Vera has a vector engine complete with SVE, but that doesn't seem like Nvidia's focus. In an end-to-end Nvidia system, those vectorized workloads would be offloaded to a Rubin GPU. Still, we don't have any vector results for Vera yet. Up to this point, we've only seen integer results, which is strange given the memory system at play in Vera.
[3]
Nvidia has 'shipped hundreds of thousands of Grace standalone servers' -- GPU firm pivots messaging as CPUs take center stage in agentic data centers
Nvidia's Ian Buck, vice president of hyperscale and high-performance computing and the inventor of CUDA, says the company has "shipped... let's put it in the hundreds of thousands of Grace standalone servers." In May, Nvidia disclosed that it had shipped over 2.5 million Grace CPUs in total, and the company announced a partnership with Meta to deploy standalone Grace servers in February. Buck's comments suggest the scale of deployment may be even larger, however, as Nvidia tries to compete in a market dominated by other players. It's an interesting comment, though not a surprising one. Nvidia has become the dominating force of Silicon Valley as demand for its GPUs skyrocketed during an unprecedented data center buildout for AI inference. Since peaking earlier this year, however, around $1 trillion in Nvidia's market cap has been wiped away as investors rally behind CPU makers like Intel. Evolving agentic AI workloads have changed the hardware balance, shifting away from as many as eight GPUs per CPU, toward a one-to-one ratio in some cases. Nvidia wants to ride that train with its new Vera CPU, which was architected specifically for those types of workloads. Even before the recent rise of agents, however, Nvidia says it has seen demand for its CPUs for data-hungry workloads. "They weren't running a web server [with Grace]... or they aren't being used for, what the cloud uses, of cheap, dollar-per-core," Buck said. "They were being deployed for the backend, data-rich operations, like the data processing." Grace represents an on-ramp for Nvidia into data center CPUs. It uses 72 stock Arm Neoverse V2 cores, but it's differentiated by Nvidia's Scalable Coherency Fabric (SCF). Vera uses an updated SCF, but it also features Nvidia's first custom core design, called Olympus. Grace cracked the door, and Vera represents Nvidia's big entrance into the market against AMD and Intel. Regardless of where Vera ends up in the battle of next-gen data center CPUs -- which is heating up now, as AMD is expected to launch its Zen 6 Venice CPUs this week -- the design is vastly different from what we've seen out of Intel and AMD. Most notably, Vera is monolithic, placing all of its 88 cores on a single piece of silicon. AMD and Intel, years ago at this point, pivoted away from monolithic dies in favor of chiplets, allowing an extremely high density of cores at the cost of latency and coherency issues. Vera is radically different in that regard, not only being built on a single die, but also dedicating significant die space to the fabric. "One of the reasons we don't have 128 cores is because we've dedicated so much of the die area toward the fabric," Buck said. "It's 3.4 TB/s of bandwidth inside of that CPU that is dedicated toward allowing every core to talk to every cache, every memory [controller] at full speed without any collisions." For clarification's sake, Buck is referencing 3.4 TB/s of core-to-core bandwidth in Vera. There's up to 1.2 TB/s of aggregate memory bandwidth (14 GB/s per core) through the LPDDR5X interface. But just as chiplet-based designs made trade-offs in per-thread performance, Vera will likely make trade-offs for its unique architecture. The majority of data center workloads are still "legacy" tasks that hyperscalers have built for, and even with seemingly insatiable demand for AI infrastructure, that is unlikely to change for several years. Buck recognizes this trade-off, asking: "Can Intel and others build rich fabrics? Do they have the IP and the ecosystem to do it and connect it all the way through to LP memory? They need to tell you when they're going to do it... but that trade-off will come at the cost of the legacy workload." Earlier this year, at GTC in March, Buck was even more clear. "The world is not going to be served by one SKU of CPU, and that is not our intention," the executive said in a news conference at the time. Still, it's clear Nvidia has ambitions with data center CPUs beyond what headlines are floating around on the New York Stock Exchange. Nvidia says CPUs represent a $200 billion TAM (Total Addressable Market) opportunity for the company, a rather rosy forecast compared to the rest of the industry, which sees a TAM of around $120 billion by 2030 (though recent estimates have climbed as high as $170 billion). And agentic AI is expanding that market, with Morgan Stanley in April estimating that agents could add as much as $60 billion to the data center CPU market. Vera is in full production alongside Nvidia's next-gen AI infrastructure, including Rubin GPUs, ConnectX-9 NICs, SpectrumX Ethernet switches, and the various components that go into building a Vera Rubin NVL72 rack. The company says there are around 1.3 million components that go into a rack, and it has a list of over 300 partners globally to build them. As part of our visit to Nvidia HQ last week, we saw a Vera Rubin NVL72 rack in action, running workloads for OpenAI. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[4]
Nvidia details its next-generation Vera CPU for AI, setting up challenge to AMD and Intel
Nvidia became the most valuable company because of insatiable demand for its graphics processing unit, or GPU, the primary chip used for creating and deploying artificial intelligence. But the chip giant is now shipping its own central processing units, or CPUs, which cloud providers could decide to deploy in place of those from Advanced Micro Devices and Intel, opening up a new battleground in AI servers. On Tuesday, Nvidia released new information about its data center CPU, called Vera, including specifications and the kind of benchmarks and architectural information that prospective customers need to fully evaluate the chip. Nvidia representatives said Vera chips were delivered to clients, including OpenAI, Anthropic, and SpaceX, in June. Nvidia is seen as the company that sets the direction for the information technology industry. But in CPUs, Nvidia is the challenger once again, competing against two well-established players in Intel and AMD, which have deep ties to hyperscalers and cloud giants. The company's investment in CPUs is another example of Nvidia's strategy to vertically integrate its systems and produce more of the chips and technology inside them every year. It aims to sell the entire system as a full rack of computing power, instead of simply selling chips by themselves. It's a strategy that Nvidia says will help engineers squeeze more performance out of their GPUs, helping the company's systems remain the computers of choice for leading AI labs as competition from AMD and custom chips heats up. Before the AI boom, the CPU was the most important and most expensive part in a server. The first generation of AI servers available when ChatGPT was released in 2022 paired as many as eight GPUs to one CPU, signaling a shift towards Nvidia's GPUs. But the rise of agentic AI, which can run independently in the background with minimal human input, has returned attention to the CPU, which is needed to babysit and feed data to an agent. Financial markets have noticed, and CPU incumbents AMD and Intel are among the two best-performing chip stocks so far in 2026, up 128% and 149% respectively, besting Nvidia's rise of 8%.
[5]
Nvidia reveals Vera CPU details, claims its custom Olympus cores outperform AMD Epyc Turin
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. What just happened? Nvidia has released more information about its upcoming Vera data center CPU this week, including key details about the custom Olympus cores that are purpose-built for agentic AI workloads. The company also claimed that the new core architecture offers faster single-core performance than competing products. Part of Nvidia's Vera Rubin platform, the Arm-based Vera CPU is designed to perform complex reasoning, facilitate tool execution, generate code, and plan tasks for AI agents. Developed as a multi-purpose chip, it can also handle scientific data processing and reinforcement learning, which demand more CPU power than typical AI workloads. Each Vera CPU incorporates 88 Olympus cores, 176 physical threads, a 64KB four-way L1 instruction cache, and 164MB of unified L3 cache on a monolithic compute die. It also features a 48-instruction decode queue. The mid-core incorporates a 96KB six-way L1 data cache and a 2MB eight-way L2 cache, alongside a hardware graph prefetcher for data-intensive graph and analytics workloads. Each custom core comprise a wide front-end with optimized branch predictors, a mid-core with critical path acceleration, and an execution engine with complex vector optimization. They offer high-bandwidth instruction fetch supporting up to 16 instructions per cycle and include a 10-wide decode engine processing up to ten fused instructions per cycle. Olympus' branch prediction subsystem includes a neural branch predictor capable of reducing wrong-path execution on specific branch patterns. Combined with the wide decode path, it also helps improve front-end throughput for latency-sensitive workloads, such as agentic AI and reinforcement learning. The Olympus core architecture integrates Nvidia's Spatial Multithreading technology, enabling flexible resource partitioning. It also supports the Scalable Coherency Fabric for 3.4 TB/s of on-die core-to-core bandwidth, 1.2 TB/s of aggregate SOCAMM2 LPDDR5X memory bandwidth, and up to 1.5 TB of memory capacity. Vera supports 1.8 TB/s connectivity on NVLink-C2C, PCIe 6.4, and CXL 3.1. It supports dual-socket scaling, too, with each socket providing 176 PCIe Gen 6 lanes. To improve latency in dual-socket systems, Vera uses a software-based solution where each socket presents itself as a single NUMA domain, avoiding fragmentation and creating a clean two-NUMA-node topology. To back up its claims about Vera's performance, Nvidia published its internal SPEC CPU 2026 test results from earlier this month. The data seems to show that Vera delivers up to 1.8 times the performance of AMD's Epyc Turin 9755 systems in select agentic AI workloads, such as 714.cpython_r. However, the results have yet to be independently verified.
[6]
NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs
The Vera CPU is speeding up next-generation NVIDIA chip designs, and early testing shows up to 1.5x higher performance on key verification and simulation workloads. The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems. To help meet that challenge, NVIDIA is collaborating with industry leaders Cadence and Synopsys to optimize critical electronic design automation (EDA) applications for the NVIDIA Vera CPU. NVIDIA is now deploying Vera across EDA workflows used to develop its next generation of CPUs and GPUs, demonstrating how high-performance CPU architecture can help accelerate some of the industry's most demanding engineering workloads. Accelerating Critical EDA Workloads What's at stake is the pace of chip production. In turn, the tempo of industry technologies that stand to benefit from boosted EDA workloads, driving development momentum. Simulation, verification and implementation technologies play a central role in semiconductor development. Long before a chip reaches manufacturing, engineers spend years validating behavior, identifying corner cases and refining designs through thousands of iterations. While GPUs and AI have accelerated many aspects of chip design, several critical EDA workloads remain heavily dependent on CPU performance. Logic simulation, formal verification and portions of digital implementation often depend on fast individual cores, efficient memory systems and strong overall throughput. That makes CPU architecture an important factor in determining how quickly engineering teams can validate designs, explore alternatives and move products toward tapeout. Highlighting Early Results With Cadence and Synopsys NVIDIA's initial testing includes several leading EDA applications. The results highlight Vera's ability to accelerate two of the most compute-intensive stages of modern chip design. Early testing on selected production-class workflows shows promising results. Cadence Jasper, a formal verification platform, uses smart proof technology and machine learning to find and fix bugs and improve verification productivity early in the design cycle. Synopsys VCS, a high-performance functional verification solution used to simulate and validate complex chip designs before fabrication, used the same number of cores in the test. Both applications showed up to 1.5x higher performance on selected workloads. Beyond benchmark results, NVIDIA is working closely with both companies on application profiling, software optimization and system-level tuning designed to improve engineering productivity across a broader range of workflows over time. Bringing Vera to the Design Process NVIDIA is deploying Vera throughout the EDA workflows used to create future NVIDIA processors. Vera combines 88 custom NVIDIA Olympus CPU cores with a high-efficiency LPDDR5X memory subsystem and second generation NVIDIA Scalable Coherent Fabric designed to deliver strong per-core performance, high memory bandwidth and consistent low latency for demanding engineering applications. These capabilities are particularly important for workloads that mix latency-sensitive jobs with large-scale regression testing across compute farms. Faster execution can shorten individual verification runs, while greater throughput enables engineers to evaluate more design alternatives and complete more validation within the same development window. Going From RTL to Silicon After defining a processor's architecture and microarchitecture, engineers describe much of its behavior at the register-transfer level (RTL). Multiple verification and implementation technologies then work together to transform that design into manufacturable silicon. These workflows span logic simulation, formal verification, regression testing and digital implementation, helping engineers validate functionality, identify corner cases and transform designs into manufacturable silicon. Because these stages are interconnected, improvements in verification throughput can help organizations identify issues earlier and reduce costly downstream design iterations. Building Future NVIDIA Chips With NVIDIA CPUs The deployment of Vera across NVIDIA's own engineering workflows reflects a broader strategy: accelerate each workload with the compute architecture best suited to the task. In EDA, GPUs and AI continue to speed many algorithms, while high-performance CPUs remain essential for critical simulation, verification and implementation workloads. Together, they help improve the performance of the overall design cycle. Looking ahead, NVIDIA plans to build on Vera with the next-generation Rosa CPU, powered by the NVIDIA Rigel core, while continuing to optimize leading EDA applications across its CPU roadmap. By using NVIDIA CPUs to help design future NVIDIA CPUs and GPUs, the company is creating a continuous feedback loop between silicon design, software optimization and systems engineering, with each generation helping build the next. Learn more about NVIDIA at DAC 2026.
[7]
Nvidia is putting its Vera CPUs to work alongside AI agents to speed up chip design
Nvidia Corp. says it's now running the chip design software its engineers are using to design the next generation of its graphics processing units, on its own silicon. It's partnering with Cadence Systems Inc. and Synopsys Inc., the two biggest providers of electronic design automation or EDS software, which are now optimizing their platforms to run on Nvidia's Vera central processing units to accelerate chip design workloads. The chipmaker said that Cadence's Jasper, a formal verification platform, and Synopsys VCS, a logical simulation tool that's used to validate chip designs before fabrication, improved their performance by 1.5-times when running on Vera central processing units. The announcement came as Nvidia revealed that it's also integrating its PhysicsNeMo physics-AI libraries and a set of GPU math libraries into the Nvidia Agent Toolkit, as part of a push to increase the role of autonomous artificial intelligence agents in the chipmaking process. AI agents can now call accelerated solvers in the same way as they use any other third-party tool, Nvidia said at the 2026 Design Automation Conference in Long Beach, California, today. Accelerating EDA workloads Nvidia is trying to accelerate the pace of chip development by speeding up EDA workloads while increasing automation in various parts of the design process. The chipmaker explained that simulation, verification and implementation are all crucial steps in the semiconductor design process. Traditionally, these steps have always been performed by human engineers, and they are painstaking processes. It's not uncommon for engineers to spend years validating behavior, identifying problems and refining designs through thousands of iterations before they settle on a final blueprint for a new generation semiconductor. Nvidia has been looking to accelerate these processes for years, but while GPUs and AI have helped in some areas, many aspects of EDA are heavily dependent on CPU performance. For instance, things like logic simulation, formal verification and some parts of digital implementation are reliant on fast individual cores, efficient memory systems and rapid overall throughput. These qualities are best provided by CPU architectures, which means they continue to play a vital role in validating new chip designs and exploring alternatives. By optimizing the Vera CPUs for EDA workloads, Nvidia says it has shown it can accelerate two of the most compute-intensive aspects of the early chip design lifecycle. Cadence Jasper is a verification platform that uses smart proof technology and machine learning algorithms to identify and fix bugs and accelerate productivity, while Synopsys VCS is used to simulate and validate complex chip designs before they're built. In Nvidia's early tests, it showed it was able to increase the performance of both applications by 1.5-times. The chipmaker said it's going to work with Cadence and Synopsys to optimize other EDA workflows on Vera, and ultimately hopes to accelerate the design of its successor, codenamed the Rosa CPU, which will be powered by the next-generation Nvidia Rigel core. Automating chip design Nvidia is creating a kind of continuous feedback loop, where its current generation Vera CPUs speed up the development of future generations, but that's not all it's doing. By adding the PhysicsNeMo and CUDA-X libraries to the Nvidia Agent Toolkit, it's also stepping up agentic involvement in the chip design process. Creating more sophisticated chips requires engineers to connect physics, simulations and performance analysis across increasingly complex design cycles. This kind of work is perfect for autonomous AI engineers, which can run simulations to generate high-fidelity data at much greater speeds than humans can. With today's update, PhysicsNeMo and CUDA-X become agent-ready tools that can facilitate these simulations. PhysicsNeMo gives agents the physics skills needed to train and deploy AI models, while the CUDA-X libraries bring accelerated solvers and quantum chemistry capabilities into agentic engineering workflows. In addition, Nvidia said it's updating the CUDA-X libraries to support "iterative sparse solvers" on its GPUs for the first time. This refers to the sparse linear algebra arithmetic that underpins physical simulations ranging from fluid flow to structural stress to electromagnetics. The new libraries announced today include cuISS for iterative solvers, cuDSS for direct sparse solvers that are used in circuit and device simulation, and cuEST for the quantum-chemistry simulations that predict how materials behave at atomic scales. Nvidia said that its partners have already seen extremely encouraging results with the new libraries. For instance, the design, test and emulation software firm Keysight Technologies Inc. said it has been able to speed up electromagnetic simulations by up to 10-times using the new cuDSS libraries, while the EDA software firm Silvaco Group Inc. was able to run a 3.2-billion-mesh-node photonic edge coupler simulation in under four hours on a cluster of 32 GPUs. According to Nvidia, no CPU-based simulation would be able to match this performance. The chipmaker also talked about how Cadence's AuraStack AI Super Agent, designed to automate printed circuit boards and advanced packaging workloads, now runs on cuDSS on its Millenium M2000 supercomputer. Cadence said it has seen design verification workflows increase by 15-times as a result of this change - an especially encouraging figure, considering that verification workloads across the chip design industry consume billions of compute hours annually. All of this software is being made freely available to chip designers. PhysicsNeMo is available under the Apache 2.0 license, while the new CUDA-X libraries are free, drop-in replacements for the handcrafted code that engineers would normally have to write for themselves. For Nvidia this makes perfect sense, for PhysicsNeMo and CUDA-X can only run on its own silicon. "Engineering has reached an inflection point. AI can now work with tools of physics, simulation and design," said Nvidia Vice President and General Manager of Computational Engineering. "With Nvidia Agent Toolkit, developers can build agentic engineers that reason using physics, run simulations and generate high-fidelity data to become a new engine for innovation in chip and system design."
[8]
NVIDIA Uses 88-Core Vera CPU to Accelerate Next-Generation CPU and GPU Design
NVIDIA has started using clusters of its own Vera processors to develop future CPUs, GPUs, and AI accelerators. The company says initial electronic design automation testing shows performance improvements of up to 1.5 times in selected verification workloads compared with its existing development infrastructure. Vera is NVIDIA's first internally designed high-performance data-center CPU following the Grace generation. It contains 88 custom Olympus CPU cores and is paired with an LPDDR5X memory subsystem. The platform also implements the second generation of NVIDIA's Scalable Coherency Fabric, which is designed to maintain high-bandwidth communication between processors, accelerators, and memory. The processors are already being deployed in NVIDIA's Portland data center. Rather than using Vera only as a companion processor for Rubin GPUs, NVIDIA is applying it to the design of the company's next generation of silicon. This "chips designing chips" approach is not new in the semiconductor industry, but it provides a practical demonstration of the CPU in a demanding internal production workload. NVIDIA worked with Cadence and Synopsys to optimize two important EDA applications. Cadence Jasper is a formal-verification platform used to identify logical errors before a chip enters production. Synopsys VCS performs functional simulation and verification of complex designs. Both workloads remain heavily dependent on CPU performance, even as GPUs and machine learning increasingly assist other stages of semiconductor development. The reported maximum 1.5× improvement was recorded in selected Jasper and VCS workloads using an equal number of CPU cores. NVIDIA has not published complete benchmark configurations, power consumption, absolute completion times, or results across the entire EDA workflow. The figures should therefore be treated as vendor results rather than independent CPU benchmarks. The development is relevant because it gives Vera a credible purpose beyond NVIDIA's rack-scale AI systems. EDA applications benefit from strong single-threaded performance, memory bandwidth, and aggregate throughput, making them useful indicators of a server CPU's real capabilities. Internal deployment also suggests that the hardware and software stack is sufficiently mature for production work. NVIDIA's CPU roadmap will continue with Rosa, which is expected to replace Olympus with newer Rigel cores. The company says it will continue collaborating with EDA developers to optimize these applications for future processors. SpecificationNVIDIA Vera CPU cores88 custom Olympus cores MemoryLPDDR5X subsystem InterconnectSecond-generation NVIDIA Scalable Coherency Fabric Tested applicationsCadence Jasper and Synopsys VCS Claimed improvementUp to 1.5× in selected workloads Current deploymentNVIDIA Portland data center SuccessorRosa CPU with Rigel cores Sources: NVIDIA Vera EDA report, NVIDIA Newsroom
[9]
Nvidia Vera Rubin: Inside the agentic AI factory that rewrites the CPU playbook
On the surface, this week's Vera Rubin launch is another major platform moment for Nvidia Corp., as the company maintains a steady drumbeat of artificial intelligence infrastructure innovation. Nvidia is positioning Vera Rubin as a full-stack system designed to improve performance per watt and reduce token costs, with production ramping across a broad global partner base, including cloud and AI infrastructure providers. Nvidia says the platform spans seven co-designed chips, integrates new networking and is already being deployed by partners such as CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. However, the bigger story isn't that Nvidia launched another AI system, since that's nothing new for the king of AI. Rather, it's that the company is making an aggressive case that the AI era, especially the rise of agentic AI, requires rethinking the central processing unit, the network and the system architecture as a single, interdependent design problem rather than a pile of best-of-breed parts. That is why Vera (pictured) matters. From cloud economics to agentic bottlenecks For years, the CPU roadmap in the data center was largely shaped by cloud economics. Hyperscalers wanted more cores, higher throughput and lower costs, which rewarded chiplet-heavy designs optimized for scale-out efficiency. Nvidia's analyst briefing, and another for reporters last week, framed that era as one in which core counts grew roughly nine time and single-thread performance gains doubled, leaving the market with processors well-suited for classic cloud workloads but less suited to the latency-sensitive, branch-heavy behavior of AI agents. As Nvidia's Hannah Coutand explained in the briefing, "AI is asking for a new CPU," because "agentic AI is putting CPU back on the critical path." That framing is important. In an agentic system, the graphics processing unit performs the reasoning, while the CPU is constantly in the loop, handling tool calls, code execution, queries, orchestration and data handling between reasoning steps. These workloads create continuous loops of reasoning, acting, observing and evaluating, which put simultaneous pressure on single-thread performance, memory bandwidth, latency and scaling efficiency. In other words, throwing more generic cores at the problem is not enough if the real bottleneck is the sequential work between GPU inference passes. Inside Vera: A CPU designed for the AI loop This is the central rationale behind Vera's new CPU architecture. Rather than extending the conventional server CPU playbook, Nvidia built a custom Arm-based processor around its new Olympus core. Vera features 88 custom cores, 176 hardware threads, up to 1.2 terabits per second of LPDDR5X memory bandwidth, 164 megabytes of unified L3 cache, and up to 1.8 TB/s of coherent CPU-GPU bandwidth via NVLink-C2C. Nvidia says the design target is what it calls the "max single-threaded CPU at scale," meaning a processor that preserves strong per-core responsiveness while still scaling across highly concurrent agent workloads. The analyst briefing made the design intent unusually clear. Coutand stated that Nvidia "didn't set out to go win CPUs," but instead recognized that "the CPU was becoming a bottleneck in the AI factory" and designed a better processor to improve "AI factory economics." That is classic Nvidia strategy: Start with the system bottleneck, then build the silicon needed to remove it. Nvidia's Ian Finder provided the technical details behind that claim. He described Olympus as a ground-up custom core featuring a 10-wide decode engine, aggressive reordering logic and a graph prefetcher tuned for pointer-chasing patterns common in compilers, graph structures and agent runtimes. His point was that CPU performance in this new era is less about chasing clock speeds and more about increasing the amount of useful work each cycle can do. That emphasis on instructions per cycle, branch handling and memory behavior is exactly what you would expect if the target is agent orchestration rather than old-school enterprise middleware. Fabric and memory: Making data movement part of compute Just as important, Vera is a reminder that in AI infrastructure, the network is no longer a peripheral technology but part of the compute architecture. Inside the CPU, Nvidia's second-generation Scalable Coherency Fabric serves as the data-movement backbone, linking cores, caches, LPDDR5X controllers, I/O and NVLink-C2C interfaces with multiterabyte-per-second bandwidth. Nvidia contrasts this monolithic fabric with chiplet-based designs that incur a "chiplet tax" in the form of higher latency and lower effective bandwidth as traffic crosses die boundaries. For agentic workloads, which are sensitive to loaded latency and cross-core data sharing, those differences translate directly into GPU utilization and end-to-end responsiveness. The memory subsystem follows the same philosophy. By pairing LPDDR5X with an enterprise-ready module form factor, Vera aims to deliver high bandwidth per core and better bandwidth-per-watt than conventional DDR-based servers. In an AI factory with thousands of deployed servers, shaving tens of watts from the CPU-plus-memory envelope while increasing bandwidth frees more of the power budget for GPUs and high-speed networking. Beyond the rack: The role of Spectrum-X Outside the CPU, the platform extends to the rack and the cluster. Within AI infrastructure there is a significant distinction between scale-up and scale-out networking. NVLink connects GPUs within a rack, enabling them to act as a unified accelerator with all-to-all bandwidth and in-network compute. Spectrum-X Ethernet provides the scale-out fabric that ties those racks together across the AI factory. Spectrum-X is more strategically important than many realize. In traditional enterprise infrastructure, Ethernet can be treated as a largely modular layer. In AI factories, the network directly affects token throughput, latency, utilization and ultimately economics. If mixture-of-experts models and agentic systems create much heavier east-west traffic and more distributed coordination, generic Ethernet becomes a tax on the entire system. Nvidia's answer is a purpose-built Ethernet stack: 102.4T Spectrum-6 switches, 1.6T ConnectX-9 SuperNICs, adaptive routing, congestion control, telemetry and open software, all tuned for RDMA and AI traffic patterns. The goal is to make Ethernet behave more like an AI-specific fabric while preserving operational familiarity, so that scale-out networking enhances AI factory performance rather than undermining it. Extreme co-design as a competitive weapon This brings us to Nvidia's "extreme co-design." The company says that Vera Rubin NVL72, the Vera CPU rack, BlueField-4 infrastructure processors, Spectrum-6 switching, and the rest of the platform were engineered as a single system rather than assembled from separate off-the-shelf products. With agentic AI, infrastructure services such as networking, storage, telemetry, security and context handling are now part of the inference pipeline itself. That means CPUs, GPUs, DPUs and switches need to be tuned together to keep expensive accelerators fed and productive without burning host CPU cycles on infrastructure work. Most semiconductor vendors can compete credibly in one layer of the stack; a few can reach two. NVIDIA now has meaningful assets across GPUs, CPUs, scale-up networking, scale-out networking, DPUs, interconnect software and system design. That breadth lets it optimize for delivered AI output -- tokens per watt, cost per token and usable throughput -- not just component specs. Nvidia's next share gain story: CPUs The most interesting industry implication of Vera is that CPUs may become Nvidia's next share-gain story. Nvidia is not trying to displace x86 across every general-purpose data center workload. It doesn't need to. The company's own sizing suggests a large incremental CPU opportunity tied specifically to AI-driven workloads and infrastructure patterns. If the CPU's role in the AI factory is increasingly to orchestrate agents, feed GPUs, manage memory movement and support low-latency tool execution, then Nvidia can leverage its GPU dominance to pull its own CPU into the design. This is the same playbook the company has used elsewhere: win the control point, then expand adjacencies. Because Nvidia already owns the strategic budget line in AI infrastructure through GPUs, it is uniquely positioned to define what the surrounding CPU, network and data processing unit should look like. Vera's tight integration with NVLink-C2C, BlueField and Spectrum-X means buyers considering Rubin-class systems are not evaluating the CPU in isolation. They are evaluating a full AI factory architecture. In the near term, that will matter most for AI clouds, hyperscalers and large model builders with agentic or reinforcement-learning-heavy workloads. Over time, though, the definition of a "good" data center CPU may shift more broadly. If Nvidia is right, the next important CPU category will not be the cheapest cloud workhorse or the highest-core-count generalist. It will be the processor that best removes friction from the AI loop. Final thoughts The fundamental tenet of my research has always been that rapid market share shifts occur when markets transition, and the CPU industry hasn't seen a significant transition in a long time. But this is the AI era, and it's seemingly redefining all industries. Nvidia is making the case that the future CPU is no longer a standalone component decision. It is a systems decision, tightly coupled to GPUs, memory, networking and infrastructure processors -- and Nvidia intends to own as much of that system as possible. This week, AMD is holding its own "Advancing AI" summit, and we should get a good look at how it plans to address the challenges Nvidia laid out above. Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.
[10]
NVIDIA's Vera CPU Slashes Chip Verification Times at Cadence and Synopsys by 1.5x, Speeding Next-Gen Silicon
NVIDIA accelerates next-gen CPU and GPU chip designs with its Vera CPUs, delivering a 50% boost to improve system-level processes. Cadence & Synopsys Are Leveraging NVIDIA's Vera CPUs To Accelerate Their Chip-Making Processes NVIDIA's CUDA-X libraries & cuLitho software are already enabling faster chip design while significantly reducing lithography costs. Now, NVIDIA is working with chip designers & industry partners to optimize Electronic Design Automation (EDA) applications using its Vera CPUs. The key EDA processes include Simulation, Verification, and Implementation. These three are the most crucial steps before a chip is sent for manufacturing. During these processes, engineers are focused on validating the various behaviors of the chip & continue to refine/optimize the design. Most of these processes are heavily reliant on CPU performance, hence requiring faster cores and efficient memory systems, and Vera gives these processes an impeccable upgrade in all regards. As per the initial tests done on NVIDIA's Vera CPUs, the chip was able to provide up to a 1.5x boost in performance: * Cadence Jasper, a formal verification platform, uses smart proof technology and machine learning to find and fix bugs and improve verification productivity early in the design cycle. * Synopsys VCS, a high-performance functional verification solution used to simulate and validate complex chip designs before fabrication, used the same number of cores in the test. Vera's faster execution reduces the time taken during individual verification runs, delivering higher throughput that allows engineers and firms to evaluate more design alternatives and complete more validation within the same development window. After the initial processes, the behavior of the chip is described in the register-transfer level (RTL) stage. The RTL processes include logic simulation, formal verification, regression testing, & digital implementation. NVIDIA Agent Toolkit capabilities include: * AI physics skills: NVIDIA PhysicsNeMo libraries help agents train and deploy customizable AI physics models for complex design and simulation tasks, turning model architectures into callable tools for engineering workflows. * Iterative sparse solvers: New NVIDIA cuISS (CUDA Iterative Sparse Solvers) library accelerates large sparse linear systems in physics-based and engineering simulations. Designed for flexibility and performance on GPUs, its modern, composable solvers and preconditioners help developers build scalable, production simulation engines for agentic engineering workflows. * Direct sparse solvers: NVIDIA cuDSS (CUDA Direct Sparse Solvers) accelerates large, complex sparse linear systems central to electronic design automation (EDA) and scientific simulation. It delivers high performance and numerical robustness for critical workloads like device, circuit, and system simulations with scalability to multi-GPU and multi-node deployments in production environments. * Quantum chemistry: NVIDIA cuEST (CUDA Electronic Structure Theory) brings high-accuracy quantum chemistry simulations to device-relevant scales, enabling density functional theory (DFT) and post-DFT methods to be integrated into production workflows at scale. cuEST brings production value to customers by supporting a wide range of modern functionals and making increasingly large ground-state and excited-state simulations manageable on NVIDIA GPUs. The Nemotron 3 Ultra Open AI models also bring agentic coding advanced to chip design with higher accuracy and deep domain expertise for RTL coding. All in all, NVIDIA's Vera CPUs bring the capabilities to EDA and RTL processes with accelerated capabilities, faster throughput across several algorithms, and at higher efficiency. NVIDIA is also working on its next-generation Rosa CPUs with the Rigel core architecture, which will continue to optimize leading EDA applications, and these chips will be a key enabler of NVIDIA's own CPUs & GPUs coming in the future. Follow Wccftech on Google to get more of our news coverage in your feeds.
[11]
NVIDIA Aiming To Produce Up To 1000 Racks With Vera Per Day After Shipping "Hundreds of Thousands" of Grace Standalone Servers As It Guns For Dominance In The AI CPU Market
NVIDIA says that it will be able to make 1,000 Vera Rubin Racks per day following the success of its Grace CPUs in AI markets. Vera Will Be A Monumental Success For NVIDIA After Grace As The Firm Races To Build 1,000 Racks Per Day Since its inception, NVIDIA has been known as a GPU maker, but the company has slowly started to move away from that, and now recognizes itself as a full-stack system provider. That is why the company is making extra efforts to accelerate its CPU roadmap, and while Grace was its first full-on take on a CPU, the Vera chip is looking to take things to a grander scale. Yesterday, we covered the Vera CPU architecture and its performance across various workloads. During the technical pre-brief, NVIDIA also shared some interesting insights into its CPU roadmap and how it is addressing the growing requirement for these chips across AI deployments that are coming online. As per NVIDIA's Vice President of HPC and Hyperscale systems, Ian Buck, NVIDIA has shipped "hundreds of thousands" of Grace standalone servers. These are just racks with Grace CPUs & previously, the company has said that they have shipped over 2.5 million Grace chips to date. These numbers could be even higher for the total number of CPUs shipped. Based on the current estimates and customer responses, Vera is looking to exceed these figures significantly. And NVIDIA is preparing ahead of this, working with over 300 global partners to make sure that Vera is in ample supply to meet the growing demands of Agentic AI. NVIDIA's SVP of hardware engineering, Andrew Bell (via The Information), states that they have around a dozen manufacturing partners that will be able to produce up to 1,000 Vera racks per day. Based on calculations, it would amount to at least $630 billion in revenue in a single quarter. NVIDIA generated $82 billion in revenue in its previous financial quarter. Now this revenue will be shared between NVIDIA and its partners, but it is a hefty jump and shows that the company is building a strong ecosystem with its partners to ensure Vera Rubin racks are assembled and delivered to AI firms and cloud data centers in a timely fashion. Last month, NVIDIA commenced the volume ramp of its Vera Rubin platform and has already positioned itself to become the leading CPU supplier in the world this year as it aims to hit $20 billion in revenue from CPUs alone. The platform is all set, and the competition is fierce, with the likes of AMD's EPYC Venice ramping up at the same time on an advanced process node. The next few months will be interesting to see the outcome in the AI segment as all chipmakers aim at one singular goal in mind: to dominate the Agentic AI race. News Source: @GavinSBaker Follow Wccftech on Google to get more of our news coverage in your feeds.
[12]
NVIDIA Vera CPU Rolls Over Traditional x86 Chips With An Agentic-AI Focused Design With 6x Faster Performance & 40% Lower Latency
NVIDIA's Vera CPU trounces x86 chips, such as AMD's EPYC, in Agentic AI workloads, delivering a 6x performance gain with 40% lower latency. NVIDIA Vera's Monolithic CPU Design & Architectural Changes Achieve Over 2x Performance Gain Over x86 CPUs Such As AMD EPYC With Just 88 Cores, Beating 128 Cores Today, NVIDIA is going all out by sharing its full performance disclosure of the Vera CPU against x86 offerings. We just talked about the architecture in detail over here, and now, we will look at the performance capabilities that this chip has on offer. NVIDIA is evaluating Vera across four key categories that matter the most in the Agentic AI era: * Memory Bandwidth and Latency - Shows how Vera keeps cores supplied with data through high-bandwidth memory, low-latency access, and a high-throughput coherent fabric. * Application Workload Performance - Demonstrates performance across agentic benchmarks, data analytics, graph processing, and other CPU-intensive workloads. * Core IPC - Measures the impact of the Olympus microarchitecture, including the wide Front End, advanced branch prediction, deep out-of-order execution, and expanded execution resources. * RL and Agentic AI Performance - Highlights Vera's value for latency-sensitive, branch-heavy, and highly concurrent agent environments. But before starting with the benchmarks, we want to state that NVIDIA has heavily used the AMD EPYC Turin "Zen 5" chip in its comparison, highlighting the monolithic & NUMA nature of its chip & how it significantly elevates its position over x86 chips. Architectural Benchmarks NVIDIA has divided the benchmarks into two categories, one that focuses on the architectural side of things and the others at how the architectures improve AI workflows. As such, the first architectural benchmark showcases Vera's IPC uplifts across various workloads that stress large instruction footprints, dense control flow, compiler and runtime behavior, and long dependency chains. Here, Vera offers up to a 1.9x uplift over AMD's Zen 5 (EPYC Turin), showcasing its strong single-threaded gains, a key driver for Agentic AI and RL apps. The Branch Prediction on Vera's Olympus cores is up to 2.3x faster and 2x faster on average versus AMD's Zen 5. The Taken-branches per cycle see a 3.5x uplift. Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle. The Olympus core features a 64 KB instruction cache with a high-bandwidth fetch of 128 bytes per cycle out and a 10-wide decode without relying on an x86-style uOP cache, which offers up to a 2.4x gain over Zen 5 in Instruction Fetch Ops per cycle. And lastly, we have the Backend Ops per cycle, which are up to 4.3x faster and over 3x faster on average versus Zen 5. Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle. AI Workload Benchmarks The first benchmark is centered around memory latency. Here, NVIDIA states that traditional x86 CPUs these days rely on chiplet architecture. The chiplet design leads to shortcomings in the memory subsystems when many requests are made across cores. These cause bottlenecks to appear in the memory subsystem and the coherency fabric, and as a result, additional hops increase latency and reduce effective bandwidth, leading to performance drawbacks. In the comparison, the AMD EPYC Turin 9755 CPU can be seen jumping from similar latency in the beginning (~120ns) up to 350ns as soon as memory bandwidth demand increases. Vera, on the other hand, can deliver 3x the bandwidth with 40x lower peak latency at a 90% memory utilization rate. NVIDIA's Vera CPUs also offer much higher memory bandwidth per core. The Turin chip has 128 cores, with each core being fed roughly 3.1 GB/s of bandwidth. Vera has 88 cores, and each core is fed with around 12.7 GB/s of bandwidth, 4x higher than EPYC, and the per-core memory bandwidth is achieved up to 14 GB/s in certain use cases. Such per-core memory bandwidth characteristics are important for Agentic AI workloads where many threads operate on large, irregular data sets and can quickly become limited by memory throughput rather than compute resources. Further impacting the performance of chiplet-based solutions is the addition of dozens of separate I/O and compute chiplets. When data moves from one core to another, it may need to leave the source core die, traverse the package fabric through an I/O die, reach a different core die, and then follow a similar path back for coherency responses. Each die hop adds latency, consumes fabric bandwidth, and creates variability depending on where the communicating cores are physically located. The result is higher core-to-core latency & impacts performance. In the core-to-core latency heat map, you can see that NVIDIA Vera delivers 50% lower latency than the AMD EPYC 9755 chip. What's even more impressive is that Vera shows consistent latency across all of its core pairs, further emphasizing how its monolithic architecture is made from the ground up for AI workflows. Next up, we move to Agentic benchmarks, which are the core reason why Vera was made in the first place. In Spec CPU 2006, NVIDIA measures the per-core performance for a fully loaded socket system. The software stack here measures performance across Python execution, code compilation, static analysis, and tool-driven software workflows. Fully loaded per-core performance captures how well each core sustains throughput while sharing socket-level power, memory bandwidth, cache, and fabric resources. The NVIDIA Vera CPU delivered up to 1.8x higher performance in Python versus AMD's EPYC. The rest of the benchmarks show up to 1.7x performance; that's almost 2x the performance bump vs x86 competitors. In Graph Traversal algorithms, NVIDIA's Vera CPU offers a 2.6x performance uplift over AMD Turin. Graph workflows are bandwidth and cache-sensitive, and here, NVIDIA's memory and cache subsystems inside Vera are showcasing displaying their full potential. NVIDIA is also showing off its Graph Traversal Core scaling in Page Rank workloads, which also stresses the core pipeline, fabric, and memory bandwidth. Here, Vera achieves a monumental 29.3x gain at a core count of 32. Lastly, we have ClickHouse, which is a high-performance columnar database used for online analytical processing. Vera achieves a 20% uplift over the AMD EPYC Turin chip and shows it to be a great fit for scan-heavy, aggregation-heavy, and concurrency-sensitive analytics and ETL (Extract-Transform-Load) workloads. NVIDIA's Vera CPU represents a significant architectural leap forward for high-performance computing, particularly in the demanding realm of Agentic AI workloads. By embracing a monolithic design rather than the chiplet-based approach common in modern x86 CPUs, Vera overcomes critical bottlenecks in memory subsystems, coherency fabric, and inter-core communication. This results in dramatically lower and more consistent latency, while delivering substantially higher memory bandwidth per core. Gunning A $200b TAM Oppurtunity The NVIDIA Vera CPU is aiming to become a highlight of the Agentic AI CPU race. The company has already said that Vera is in a position to become the leading CPU supplied in 2026, and there are plenty of early adopters that have started using Vera, such as OpenAI, Anthropic, SpaceX, & Perplexity. The CPU is being deployed across various supercomputer centers and cloud data centers, and will see full-fledged support from countless ecosystem partners. Optimized from the ground up for irregular, bandwidth-hungry, and concurrency-intensive AI workflows, Vera not only sustains higher per-core throughput under full load but also sets a new standard for efficient, scalable processing in AI datacenters. Follow Wccftech on Google to get more of our news coverage in your feeds.
[13]
NVIDIA Vera CPU Is Architected For The Agentic AI Era, as It Delivers Max Single-Core & Single-Thread Performance Versus x86; Full Architectural Breakdown Shows
The NVIDIA Vera CPU is designed for Reinforcement Learning and Agentic AI systems, offering 88 Olympus cores for max single-core performance. NVIDIA's x86-Killer, Vera CPU, Is Now Available, Built Purely With Agentic AI In Mind Surge in agentic AI workloads and reinforcement learning emerging as a key scaling mechanism for improving model capability have made CPUs a vital component for AI systems. Without a fast CPU (execution, evaluation, orchestration), a GPU won't be able to be supplied with context, and that leads to major drawbacks and inefficiencies. This is why NVIDIA designed Rubin, a CPU that is purpose-built for agents, and a chip that eliminates these bottlenecks, delivering the throughput and efficiency that the industry demands. Today, we will give you a full rundown of Vera, NVIDIA's latest and greatest CPU for AI, based on custom Arm cores. The Insides of The Vera CPU On a high-level, NVIDIA's Vera CPU is based on the Olympus Core architecture, which utilizes the custom Armv9.2 IP. It is based on a monolithic compute die and comes with technologies such as 2nd Gen NVIDIA Scalable Coherency Fabric (SCF), Confidential Computing (TEE-I/O Capable), and x16 PCIe Gen6 CXL 3.1. Key benefits of the NVIDIA Scalable Coherency Fabric include: * Low-latency communication between CPU clusters * Efficient access to shared system-level cache resources * High-bandwidth connectivity to memory controllers * Native integration with NVLink-C2C and accelerator infrastructure The Brand Prediction features Neural & Special-Purpose Predictors with 2 taken branchers per cycle. It has a 10-Wide Instruction Decode, a Novel Graph and Advanced Prefetcher, 96 KB of L1D, 64 KB of L1I, 2 MB of L2, and 164 MB of L3 cache. Each core has access to 14 GB/s of bandwidth. The NVLink C2C interconnect offers up to 1.8 TB/s of coherent CPU-GPU bandwidth. In total, there are 88 Olympus "high-performance" cores and 176 NVIDIA Spatial Multi-threading threads on Vera. The chip integrates a wide, high-IPC microarchitecture that is designed to offer maximum single-threaded performance, achieving efficient concurrency across thousands of software environments. Olympus - The AI Peak Olympus, the core powering Vera, is referred to as the compute engine of the CPU. It is built from the ground up with one goal: to maximize IPC (Instructions Per Cycle) for irregular, branch-heavy, latency-sensitive, and highly concurrent software. Versus Grace, Vera achieves a 50% gain in IPC. There are four key components of the Vera CPU architecture: * Front End: Optimized Branch Predictors * Mid Core: Critical Path Acceleration * Execution Engine: Complex Vector Optimization * Cache Subsystem: Latency-Optimized Instruction and Data Paths The following is the block diagram of the Olympus core: Olympus Neural Branch Predictor Starting with the dissection of the Olympus core, we first have the Neural Branch Predictor, which is designed to improve prediction accuracy on complex, branch-heavy software. The NBP learns longer-range control-flow behavior and statistically biased branch patterns that are common in large runtime environments, interpreters, databases, compilers, and agent workloads. The branch prediction improvements allow Olympus's front end to be supplied with useful instructions, which in turn enable higher instruction throughput and more consistent performance. Olympus Mid-Core The Mid Core on Olympus turns decoded instructions into executable operations while maxing instruction-level and memory-level parallelism. The Olympus Mid Core is designed to identify and execute independent work hidden behind these bottlenecks, enabling higher utilization of execution resources and improved overall IPC. Mid Core comes with a wide rename and allocation engine that is supported by a large reorder buffer and extensive physical register resources. The Mid Core Register renaming eliminates false dependencies between instructions, allowing the processor to expose additional parallelism and execute instructions as soon as their true operands become available. A large instruction window enables Olympus to look significantly further ahead in the instruction stream, finding useful work while other instructions wait on memory or branch resolution. Olympus also houses several technologies to break dependencies, such as: * Memory Renaming: Accelerates store-to-load dependency chains by allowing dependent instructions to execute before a load completes when the data relationship can be predicted or determined. This is particularly beneficial for pointer-heavy software, graph traversal, runtime frameworks, and complex object-oriented workloads. * Value Prediction: Identifies stable dependency chains and predicts future values before they are produced, allowing dependent instructions to execute speculatively while correctness is verified later. This capability is especially effective for repetitive software patterns and sequential data structures. * Move elimination and value optimization: Techniques to achieve higher utilization. Olympus Execution Engine The execution engine within the Olympus core is designed to sustain high instruction throughput across various workloads, including Agentic AI runtimes, RL environments, databases, and more. After the instructions pass through the rename and scheduling stages, they are dynamically dispatched to specialized execution resources optimized for integer arithmetic, branch processing, vector computation, and memory operations. The combination of a wide execution backend and deep out-of-order engine allows Olympus cores to continue making progress even when some parts of the workloads are slowed down by memory access, branch resolution, or long dependency chains. So for the Execution Engine, NVIDIA packs 8 simple Integer ALUs, 2 Complex ALUs, and 4 dedicated branch execution units. The simple ALUs handle common arithmetic and logical operations, while the complex ALUs accelerate higher-latency operations such as multiplication, division, CRC, and shift-intensive workloads. The dedicated branch units work in conjunction with the advanced branch prediction subsystem to rapidly resolve control-flow decisions and minimize pipeline disruption. This combination is particularly beneficial for modern agentic workloads, which often exhibit irregular control flow and frequent branching as agents reason, evaluate state, and execute actions. For vector and AI-oriented computation, Olympus integrates six 128-bit SVE vector execution units, including two crypto-capable vector pipelines, and supports FP8 precision. These units accelerate vectorized workloads commonly found in data processing, machine learning preprocessing, compression, encryption, scientific computing, and analytics frameworks. Support for Arm Scalable Vector Extension (SVE) enables efficient execution of both traditional HPC workloads and emerging AI data-processing pipelines. The Execution Engine is tightly integrated with a high-bandwidth memory subsystem through four load pipelines and 2 store pipelines. Together, these resources enable Olympus to maintain high utilization across both compute-intensive and memory-intensive workloads. Vera's Spatial Multi-Threading Offers Higher Single-Threaded Throughput Than Traditional SMT Approaches NVIDIA is making use of a new multi-threading solution in Vera CPUs called Spatial Multi-Threading. Compared to traditional SMT (Simultaneous Multi-Threading), the NVIDIA approach makes it so that resources can be partitioned across two hardware threads inside the Olympus core instead of relying primarily on a narrower core. Traditional SMT achieves this by allowing both hardware threads to opportunistically share the same core resources. In densely threaded sandbox environments, that sharing can create contention across branch prediction, decode, execution, load/store, cache, and memory resources. This creates a "noisy neighbor" effect, where activity on one thread can reduce the performance of the other thread sharing the same core within a sandbox. For agentic AI, reinforcement learning, and cloud services -- where many short-running tasks execute concurrently and tail latency matters -- this can increase variability and make performance less predictable. With this, Vera offers a high-throughput single-threaded core when needed, while the thread duo handles management. With 88 cores and 176 threads, NVIDIA's Vera is a flexible CPU that offers both high single-threaded performance and dense vGPU-style deployments. Spatial Multithreading improves determinism, isolation, and quality of service compared to traditional SMT approaches. The result is a CPU architecture that can run large numbers of concurrent agent tasks while maintaining more consistent latency and throughput. The Vera Memory Subsystem - LP5X Lowers Power Input Versus DDR DIMMs NVIDIA went with LPDDR5X memory for its Vera CPUs by fusing the chip with eight controllers, offering up to 1.2 TB/s of bandwidth. The LPDDR5X memory comes with data center-class ECC (Error-Correcting Code) and enables the platform to significantly reduce power consumption vs traditional DDR4-based servers. The chip enables up to 1.5 TB of LPDDR5X capacity. The use of LPDDR5X memory in the SOCAMM2 form factor also plays a vital role in driving up power efficiency, while the three primary objectives for this memory type include: * High Bandwidth - Deliver the memory bandwidth required to sustain modern AI, analytics, and HPC workloads. * Power Efficiency - Maximize bandwidth per watt using LPDDR5X technology to reduce overall platform power. * Enterprise Serviceability - Provide a modular, field-replaceable memory architecture suitable for hyperscale and enterprise deployments. The LPDDR5X memory on Vera operates at 9600 MT/s, offering up to 14 GB/s of memory bandwidth to each Olympus core. The memory capacities are fully scalable from 256 GB up to 1.5TB. Now, the use of LP5X also has a major reason, and that is to drive down the power envelope. Traditional DRAM standards such as DDR5 (DIMM, RDIMM, MRDIMM) do offer higher bandwidth, but they also use much higher power. RDIMM platforms can consume over 100W, while MRDIMM solutions can use over 200W of power in bandwidth-heavy workloads. SOCAMM2, on the other hand, achieves up to 1.2 TB/s speeds while consuming just 30-40W, and that's for a fully-loaded system. The IOs - PCIe 6.4, 2nd Gen NVLink, Dual-Socket Scale, and Confidential Computing On the IO, Security and Scale-Up front, NVIDIA Vera CPUs offer lots of technological innovations such as the following: * Second-Generation NVLink-C2C - High-bandwidth coherent interconnect for CPU-to-GPU and CPU-to-CPU communication. * Two-Socket Scale-Up - High-speed coherent processor-to-processor connectivity for scalable dual-socket systems. * PCI Express 6.4 - High-performance I/O connectivity for storage, networking, and accelerator devices. * Compute Express Link (CXL) 3.1 - Extends the PCIe infrastructure to support memory expansion, pooling, and composable system architectures. * Confidential Computing - Secure scale-up environment across 2-socket and NVL rack. Starting with PCIe, Vera leverages the latest PCIe 6.4 standard with each chip offering up to 88 PCIe lanes, while dual-socket configurations offer 176 lanes. These operate at speeds of up to 64 GT/s, doubling the throughput of the PCIe 5 standard, and ensure enough bandwidth for AICs and storage solutions. Vera also comes with a wide array of PCIe lane bifurcation, as detailed in the table below: Further extension to CXL 3.1 enables next-gen memory expansion and composable infrastructure with the support of CXL Type-3 memory devices beyond the standard LPDDR5X solution. Vera supports Arm CCA, RME-DA, and RME-CDA, providing the foundation for confidential VM isolation and secure device assignment. Per-VM encryption keys enable cryptographic isolation between tenants running on the Vera CPU, while support for TDISP-capable devices extends the trusted execution environment beyond the CPU. As the industry's first Confidential Computing CPU to enable TDISP for coherent devices such as GPUs, Vera provides high-performance, coherent access without bounce buffers or unnecessary data copies while maintaining strong workload isolation. On the security, reliability, and availability front, Vera offers the following: NVIDIA's Vera CPU represents a groundbreaking advancement in processor design, purpose-built to power the next era of agentic AI and reinforcement learning workloads. Leveraging the innovative Olympus core architecture with custom Armv9.2 technology, Vera delivers exceptional single-threaded performance through a high-IPC microarchitecture featuring advanced neural branch prediction, sophisticated mid-core parallelism techniques like memory and value prediction, a powerful execution engine with vector capabilities, and an optimized cache subsystem. Its Spatial Multi-Threading approach ensures superior determinism and isolation compared to traditional SMT, while the integration of high-bandwidth LPDDR5X memory, second-generation NVLink-C2C, PCIe 6.4, CXL 3.1, and comprehensive confidential computing features enables seamless CPU-GPU orchestration, massive scalability, and enhanced security. By eliminating traditional bottlenecks in execution, evaluation, and data movement, Vera achieves up to 50% better IPC over its predecessors, offering unprecedented efficiency, throughput, and low-latency performance tailored for complex, concurrent AI environments -- positioning it as a true x86 alternative and a cornerstone for upcoming AI systems. Follow Wccftech on Google to get more of our news coverage in your feeds.
[14]
Nvidia Unveils Vera to Speed Expansion in Server Processor Market
With Vera, Nvidia is betting on an architecture optimized for artificial intelligence applications, particularly autonomous agents that require rapid communication between CPUs and GPUs. The company says its processor delivers up to 50% higher performance than competing x86 chips on these workloads, thanks to an architecture that prioritizes per-core speed, memory bandwidth and lower latency. Vera will be sold as a standalone chip or integrated into different compute racks, notably alongside GPUs from the Vera Rubin platform. The push comes as the server CPU market regains importance with the rise of artificial intelligence. Nvidia is seeking to broaden its footprint beyond graphics processors, even as AMD and Intel retain strong legacy positions in the segment. Several analysts believe Vera could form a new category of processors geared to the most demanding AI uses, but its success will hinge on adoption by major cloud service providers, a market Nvidia is now targeting with a fully integrated offering. Nvidia shares are up 1.6% today.
Share
Copy Link
Nvidia released detailed specifications for its Vera CPU, featuring 88 custom Olympus cores designed specifically for agentic AI workloads. The chip is part of the company's Vera Rubin platform and represents its first custom core design for data centers. Early customers including OpenAI, Microsoft, and Oracle are already receiving systems, marking Nvidia's aggressive push into a CPU market traditionally dominated by AMD and Intel.
Nvidia has released comprehensive technical details about its next-generation Vera CPU for AI, marking a significant shift in the company's strategy beyond its GPU dominance
1
. The Vera CPU features 88 custom Olympus cores built on a monolithic die, distinguishing it from the chiplet-based approaches favored by AMD and Intel2
. Each core incorporates a 64KB four-way L1 instruction cache, 164MB of unified L3 cache, and supports 176 physical threads across the chip5
.
Source: Wccftech
The Olympus architecture integrates several advanced features, including a 10-wide decode engine processing up to ten fused instructions per cycle and a neural branch predictor that reduces wrong-path execution
5
. At the heart of the design lies the Scalable Coherency Fabric, which delivers 3.4 TB/s of core-to-core bandwidth and up to 1.2 TB/s of aggregate LPDDR5X memory bandwidth3
. Ian Buck, Nvidia's vice president of accelerated computing and CUDA architect, emphasized the company's commitment to innovation: "We're on a roadmap to crank out new architectures, not just GPUs but CPUs"1
.Nvidia shared internal SPEC CPU 2026 benchmark results comparing the Vera CPU against AMD Epyc Turin 9755 systems in dual-socket configurations
2
. The data shows Vera achieving a 3% advantage in overall SPECrate integer suite scores despite having fewer threads than the competing AMD Epyc system2
. In specific agentic AI workloads such as code compilation and interpretation tasks, Nvidia claims Vera delivers up to 1.8 times the performance of AMD's Epyc Turin 9755 systems5
.These benchmarks remain unofficial since Vera was tested in reference systems ahead of broad availability, scheduled for the second half of 2026
2
. Nvidia representatives confirmed that Vera chips were delivered to clients including OpenAI, Anthropic, and SpaceX in June4
. During a data center lab tour, Nvidia executives revealed that OpenAI already has one Vera Rubin rack in operation1
.The Vera CPU forms the centerpiece of Nvidia's Vera Rubin platform, which combines CPUs and GPUs into integrated systems for AI data centers
1
. Each Vera Rubin NVL72 system pairs 36 Vera CPUs with 72 Rubin GPUs in a single liquid-cooled rack containing approximately 1.3 million components3
. The platform processes ten times as many tokens per watt compared to the predecessor Grace Blackwell superchip1
.
Source: Wccftech
Nvidia has significantly reduced cabling requirements in multi-rack configurations, marketing Vera Rubin as "cable-free compute" with hot-swappable capabilities
1
. Installation time per rack theoretically drops from hours to minutes, according to statements from Buck and Andrew Bell, Nvidia's senior vice president of hardware engineering1
. The system uses 100 percent liquid cooling, reducing energy consumption compared to air-cooled alternatives1
. Early customers for the platform include Microsoft, OpenAI, and Oracle1
.Related Stories
The rise of agentic AI workloads has fundamentally altered hardware requirements in AI data centers, shifting ratios from eight GPUs per CPU to one-to-one configurations in some deployments
3
. CPUs orchestrate data flows, networking, and software tasks for AI agents that run independently with minimal human input4
. This transition has benefited CPU makers, with AMD and Intel stock prices rising 128% and 149% respectively in 2026, compared to Nvidia's 8% gain4
.
Source: SiliconANGLE
Nvidia estimates the data center CPUs market represents a $200 billion total addressable market opportunity, exceeding broader industry forecasts of $120 billion by 2030
3
. Morgan Stanley projected in April that agents could add as much as $60 billion to the AI server market3
. Buck revealed that Nvidia has "shipped hundreds of thousands of Grace standalone servers," building on the company's disclosure of over 2.5 million Grace CPUs shipped by May3
. A partnership with Meta for deploying standalone Grace servers demonstrates demand for Nvidia's CPU offerings beyond integrated GPU systems3
.Nvidia's push into data center CPUs represents a direct challenge to AMD and Intel, companies with deep relationships among hyperscalers and cloud providers
4
. The company's vertical integration strategy aims to control more AI chip infrastructure components within complete rack systems rather than selling individual chips4
. Nvidia's marketing push for Vera coincides with AMD's annual conference, where the company revealed details about its competing Helios AI chip rack1
.The monolithic design of Vera contrasts sharply with the chiplet approaches from AMD and Intel, dedicating substantial die area to fabric connectivity rather than maximizing core count
3
. Buck acknowledged the architectural trade-offs: "One of the reasons we don't have 128 cores is because we've dedicated so much of the die area toward the fabric"3
. Nvidia maintains that Vera targets expanding AI infrastructure markets rather than legacy cloud workloads where AMD and Intel remain entrenched3
. The company reportedly told Chinese customers that standalone Vera CPUs could be ready as soon as August1
.Summarized by
Navi
[2]
[3]
26 May 2026•Technology
08 Jul 2026•Technology

26 Jan 2026•Technology
1
Technology

2
Science and Research

3
Technology

1
AI Agents Escape Safety Tests, Start Turf Wars and Hack Real Systems in Alarming Security Incidents

2
DeepMind's AI weather model gives forecasters an extra day to prepare for deadly tropical cyclones

3
Google Unveils Pixel 11 Series With Gemini AI, New Pixel Tag Tracker and Watch 5 at Made by Google 2026
