Nvidia Revives Rubin CPX AI Chip With Major Redesign for 2027 Launch

2 Sources

Share

Nvidia has brought back its Rubin CPX accelerator after it appeared to vanish from the company's AI roadmap earlier this year. Analyst Ming-Chi Kuo reports the AI chip features a substantially redesigned architecture optimized for AI inference workloads, with production scheduled to begin in the first quarter of 2027.

Nvidia Rubin CPX Returns With Redesigned Architecture

Nvidia has revived its Rubin CPX chip program after the market believed the accelerator had been dropped from the company's product roadmap. Analyst Ming-Chi Kuo from TF International Securities revealed that his latest industry checks indicate Nvidia restarted the program with production expected to begin in the first quarter of 2027

1

2

. The AI chip features major changes to both GPU specifications and rack architecture, designed specifically to accelerate the prefill stage of AI inference—the process of reading and processing a model's input before generating a response.

Enhanced Memory Configuration and Performance Specifications

The revived Nvidia Rubin CPX will feature 168GB of HBM4 memory per GPU, positioning it between the 288GB of HBM4 memory in Nvidia's standard Rubin GPUs and the 128GB of GDDR7 memory in the earlier CPX design

1

. Each CPX GPU delivers computing performance similar to the Rubin GPU and has a maximum power rating of 2,300 watts

2

. This configuration reflects Nvidia's strategic focus on optimizing AI inference workloads, particularly for handling the computationally intensive prefill tasks that account for more than 50% of current AI inference workload

2

.

Standalone Rack Design With Flexible Configurations

The Rubin CPX chip program introduces a significant shift in deployment strategy with a standalone MGX ETL rack instead of sharing rack space with Rubin GPUs

1

. This modular design allows customers to configure systems with 64, 128, 192, or 256 CPX GPUs

1

. Within each 64-GPU module, eight compute trays would each house eight CPX GPUs alongside a switch tray

1

. Each eight-CPX tray contains approximately 1.34TB of HBM4 memory

2

, providing substantial memory capacity for KV cache creation.

Advanced Connectivity Architecture

Nvidia will use NVLink to connect the eight GPUs within each tray, offering 1 to 1.5TB per second of NVLink bandwidth per CPX, compared to 3.6TB per second for each Rubin GPU

2

. Communication between trays uses Spectrum-6 Ethernet with copper links, while cross-module connections rely on OSFP optical links

2

. Ethernet will handle communication between trays and rack modules

1

, creating a hybrid connectivity approach that balances performance and scalability.

Strategic Partnership With Vera Rubin NVL72 Systems

Rather than replacing Rubin, the CPX is expected to work alongside Nvidia's Vera Rubin NVL72 systems in a complementary role. Ming-Chi Kuo stated that Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs, with CPX handling prefill and generating the KV cache before transferring that information to Rubin over Ethernet RDMA for decode processing

1

2

. This division of labor optimizes AI inference by assigning each accelerator to the workload stage where it performs best, potentially improving overall system efficiency for cloud providers like Microsoft Azure and Google Cloud that are already deploying Vera Rubin systems.

Market Context and Nvidia's Continued Dominance

The revival of the Rubin CPX chip program comes as Nvidia continues to dominate the AI chip market. In August, Nvidia reported $96.22 billion in second-quarter revenue, marking a 106% increase from a year earlier and surpassing the Street consensus estimate of $92.18 billion

1

. The company confirmed that Vera Rubin was ramping into full production, with partners including CoreWeave, Nebius, Microsoft Azure, Google Cloud, and Oracle Cloud

1

. The 2027 launch timeline positions Nvidia to maintain its competitive edge as AI inference workloads continue to grow and evolve, particularly as the industry shifts focus from training to deployment and inference optimization.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved