2 Sources
[1]
Nvidia's Rubin CPX Back From the Dead? Ming-Chi Kuo Says AI Chip Back With Major Redesign - NVIDIA (NASDA
Nvidia Corp's (NASDAQ:NVDA) Rubin CPX appeared to have disappeared from the company's AI roadmap earlier this year. Now, analyst Ming-Chi Kuo says the chip giant has revived the accelerator with a substantially redesigned architecture and plans to begin production in the first quarter of 2027. Ming-Chi Kuo Says Nvidia Has Revived Rubin CPX "Just as the market had come to believe that Rubin CPX had been dropped from Nvidia's product roadmap," Kuo said Monday in a post on X, adding that his latest industry checks indicate that Nvidia has "revived the program." Kuo said the new version is designed to deliver stronger prefill performance, with major changes to both its GPU specifications and rack architecture. Rubin CPX is designed to accelerate the prefill stage of AI inference -- the process of reading and processing a model's input before it generates a response. Kuo said the accelerator will feature 168GB of HBM4 high-bandwidth memory per GPU, compared with 288GB of HBM4 memory in Nvidia's Rubin GPUs and 128GB of GDDR7 memory in the earlier CPX design. Tech Ming-Chi Kuo Challenges Report of $1 Billion in Apple Chips Stuck at TSMC: 'Tight Memory Supply is Real,' But This Doesn't Add Up On Monday, Ming-Chi Kuo disputed the $1 billion Apple chip backlog report, while confirming a real DRAM shortage. 4 min read Read this article Nvidia Rubin CPX Gets a Major Rack Redesign The revived CPX will move into a standalone MGX ETL rack instead of sharing a rack with Rubin GPUs, according to Kuo. Trending Customers could configure systems with 64, 128, 192 or 256 CPX GPUs. Within each 64-GPU module, eight compute trays would each house eight CPX GPUs, alongside a switch tray. Kuo also said Nvidia will use NVLink to connect the eight GPUs within each tray, while Ethernet will handle communication between trays and rack modules. Rubin CPX Will Work Alongside Vera Rubin Rather than replacing Rubin, CPX is expected to work alongside Nvidia's Vera Rubin NVL72 systems. Kuo said Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs, with CPX handling prefill and generating the KV cache before transferring that information to Rubin over Ethernet RDMA for the decode stage. Kuo's statement came after Nvidia appeared to remove CPX from its roadmap at GTC 2026, where the company instead highlighted Groq 3 LPUs and LPX racks. Nvidia did not immediately respond to Benzinga's request for comment. Nvidia's Revenue Surges as Vera Rubin Ramps Up In August, Nvidia reported $96.22 billion in second-quarter revenue, marking a 106% increase from a year earlier and surpassing the Street consensus estimate of $92.18 billion. The company also said Vera Rubin was ramping into full production, with CoreWeave, Nebius, Microsoft Azure, Google Cloud and Oracle Cloud among its partners. Price Action: Nvidia shares closed at $220.50 on Monday, up 1.36%. In after-hours trading, the shares are down 0.10%, according to Benzinga Pro. According to Benzinga Edge Rankings, Nvidia ranks in the 98th percentile for growth and maintains positive price-trend ratings across the short-, medium- and long-term time frames. Disclaimer: This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Markets Michael Burry Says This AI Chip Startup Is 'Serious Competition' for Nvidia Michael Burry points to Etched, which hires former Nvidia engineers, as posing a threat to Nvidia's AI dominance. 3 min read Read this article Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[2]
Nvidia revives Rubin CPX chip program with 2027 launch planned, analyst says By Investing.com
Investing.com -- Nvidia Corp (NASDAQ:NVDA) has restarted its Rubin CPX product program, which the market believed had been canceled, according to industry research by Ming-Chi Kuo of TF International Securities. Production is scheduled to start in the first quarter of 2027. "Just as the market had come to believe that Rubin CPX had been dropped from Nvidia's product roadmap, my latest industry checks indicate that Nvidia has revived the program, with production expected to begin in 1Q27," Kuo stated in a post on X. The updated Rubin CPX includes modifications to GPU specifications and rack design. Each CPX GPU delivers computing performance similar to the Rubin GPU and has a maximum power rating of 2,300 watts. The memory configuration uses 168 GB of HBM4, compared to 288 GB on Rubin and 128 GB of GDDR7 in the earlier CPX version. The new design uses a separate MGX ETL rack instead of sharing rack space with Rubin. Customers can select configurations with 64, 128, 192, or 256 CPX GPUs. Each group of 64 GPUs forms a module with eight compute trays containing eight GPUs each and one switch tray. NVLink provides connectivity for scale-up among eight CPX GPUs within each tray, offering 1 to 1.5 TB per second of NVLink bandwidth per CPX, compared to 3.6 TB per second for each Rubin GPU. Communication between trays uses Spectrum-6 Ethernet with copper links, while cross-module connections use OSFP optical links. Nvidia requires CPX to work alongside Vera Rubin NVL72 in a 1:1 ratio. The CPX handles prefill tasks and creates the KV cache, which transfers to Rubin through Ethernet RDMA for decode processing. Kuo noted that more than 50% of current AI inference workload involves processing input context and building KV cache. Each eight-CPX tray contains about 1.34 TB of HBM4 memory. This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.
Share
Copy Link
Nvidia has brought back its Rubin CPX accelerator after it appeared to vanish from the company's AI roadmap earlier this year. Analyst Ming-Chi Kuo reports the AI chip features a substantially redesigned architecture optimized for AI inference workloads, with production scheduled to begin in the first quarter of 2027.
Nvidia has revived its Rubin CPX chip program after the market believed the accelerator had been dropped from the company's product roadmap. Analyst Ming-Chi Kuo from TF International Securities revealed that his latest industry checks indicate Nvidia restarted the program with production expected to begin in the first quarter of 2027
1
2
. The AI chip features major changes to both GPU specifications and rack architecture, designed specifically to accelerate the prefill stage of AI inference—the process of reading and processing a model's input before generating a response.The revived Nvidia Rubin CPX will feature 168GB of HBM4 memory per GPU, positioning it between the 288GB of HBM4 memory in Nvidia's standard Rubin GPUs and the 128GB of GDDR7 memory in the earlier CPX design
1
. Each CPX GPU delivers computing performance similar to the Rubin GPU and has a maximum power rating of 2,300 watts2
. This configuration reflects Nvidia's strategic focus on optimizing AI inference workloads, particularly for handling the computationally intensive prefill tasks that account for more than 50% of current AI inference workload2
.The Rubin CPX chip program introduces a significant shift in deployment strategy with a standalone MGX ETL rack instead of sharing rack space with Rubin GPUs
1
. This modular design allows customers to configure systems with 64, 128, 192, or 256 CPX GPUs1
. Within each 64-GPU module, eight compute trays would each house eight CPX GPUs alongside a switch tray1
. Each eight-CPX tray contains approximately 1.34TB of HBM4 memory2
, providing substantial memory capacity for KV cache creation.Nvidia will use NVLink to connect the eight GPUs within each tray, offering 1 to 1.5TB per second of NVLink bandwidth per CPX, compared to 3.6TB per second for each Rubin GPU
2
. Communication between trays uses Spectrum-6 Ethernet with copper links, while cross-module connections rely on OSFP optical links2
. Ethernet will handle communication between trays and rack modules1
, creating a hybrid connectivity approach that balances performance and scalability.Related Stories
Rather than replacing Rubin, the CPX is expected to work alongside Nvidia's Vera Rubin NVL72 systems in a complementary role. Ming-Chi Kuo stated that Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs, with CPX handling prefill and generating the KV cache before transferring that information to Rubin over Ethernet RDMA for decode processing
1
2
. This division of labor optimizes AI inference by assigning each accelerator to the workload stage where it performs best, potentially improving overall system efficiency for cloud providers like Microsoft Azure and Google Cloud that are already deploying Vera Rubin systems.The revival of the Rubin CPX chip program comes as Nvidia continues to dominate the AI chip market. In August, Nvidia reported $96.22 billion in second-quarter revenue, marking a 106% increase from a year earlier and surpassing the Street consensus estimate of $92.18 billion
1
. The company confirmed that Vera Rubin was ramping into full production, with partners including CoreWeave, Nebius, Microsoft Azure, Google Cloud, and Oracle Cloud1
. The 2027 launch timeline positions Nvidia to maintain its competitive edge as AI inference workloads continue to grow and evolve, particularly as the industry shifts focus from training to deployment and inference optimization.Summarized by
Navi
[1]
09 Sept 2025•Technology

23 Aug 2025•Technology

10 Jun 2025•Technology
