CoreWeave deployed the first multi-rack Nvidia Vera Rubin NVL72 cluster, connecting hundreds of Rubin GPUs into a single scale-out system for agentic AI workloads. The company also introduced cross-region write acceleration and Archive tier in CoreWeave AI Object Storage, delivering 8x faster latency and up to 7 GB/s throughput per GPU.

CoreWeave Leads with First Multi-Rack Nvidia Vera Rubin Deployment

CoreWeave became the first AI cloud provider to validate and deploy a multi-rack Nvidia Vera Rubin NVL72 cluster on CoreWeave Cloud, marking a significant advancement in AI cloud infrastructure

1

2

. The deployment connects hundreds of Rubin GPUs into a single scale-out cluster specifically designed for agentic AI workloads. Following the announcement, CoreWeave shares gained 2.9% in premarket trading Wednesday

1

. Chen Goldberg, executive vice president of product and engineering at CoreWeave, stated, "CoreWeave was the first AI cloud provider to validate and bring up a Vera Rubin NVL72. With multi-rack Vera Rubin, we are connecting hundreds of Rubin GPUs as a single scale-out cluster"

2

.

Technical Architecture Powers Massive Scale

A single Nvidia Vera Rubin NVL72 rack combines 72 Rubin GPUs with 36 Vera CPUs, Nvidia NVLink 6, Nvidia ConnectX-9 SuperNICs, and Nvidia BlueField-4 DPUs

1

2

. CoreWeave unifies multiple racks of hundreds of accelerators using Nvidia Spectrum-X Ethernet networking into a coordinated system

3

. Each Rubin GPU includes two Nvidia ConnectX-9 SuperNICs, delivering 1.6 Tb/s of connectivity per GPU and supporting approximately 128,000 GPUs per rail in a non-blocking fabric

2

3

. The multi-rack Vera Rubin NVL72 clusters enable training and inference jobs to run across hundreds of Rubin GPUs, delivering the capacity to train larger models, serve more demanding inference workloads, and run reinforcement learning at scale

3

.

Enhanced Storage Capabilities Reduce Latency

CoreWeave introduced two new capabilities in CoreWeave AI Object Storage: cross-region write acceleration and a new Archive tier

1

2

. The cross-region write acceleration feature allows data to be written at local latency while being migrated to a second remote region in the background, enabling an agent's intermediate state, retrieved context, and outputs to move as fast as its reasoning

1

3

. The Archive tier offers lower-cost storage with no retrieval, early deletion, or reading fees

1

3

. CoreWeave AI Object Storage LOTA delivers reads at local NVMe speeds and reduces latency by 8x compared to reading from a traditional storage cluster, providing up to 7 GB/s of throughput per GPU

1

2

. Cécile Robert-Michon, director of internal infrastructure at Cohere, noted, "CoreWeave AI Object Storage gives us a unified dataset footprint across regions with reads cached locally, so nothing waits on the network"

2

.

Automated Systems Enable Rapid Deployment

CoreWeave brings multiple racks up as a single system through automating rack lifecycle control via Mission Control, which coordinates hardware detection, firmware updates, validation, power, and cooling

3

. The Rack LifeCycle Controller works with Racky for rack control and Valvey for executing cooling actions. CoreWeave combines Nvidia field diagnostics with full-rack workload testing, comparing every result before a rack goes into production to ensure every GPU performs at its best

3

. This deployment positions CoreWeave to compete more aggressively in the growing market for agentic AI workloads, where low-latency data migration and massive computational scale are critical. CoreWeave completed its public listing on Nasdaq in March 2025

2

, and the company has demonstrated industry-leading performance through record-breaking MLPerf benchmark results in inference and training

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved