NVIDIA Open-Sources cuFile API and Launches Storage-Next Initiative With 40+ Industry Partners

4 Sources

Share

NVIDIA announced it is open-sourcing its cuFile API at the Future of Memory and Storage conference, enabling direct GPU-to-storage data access in microseconds. The company also launched Storage-Next, bringing together over 40 storage vendors including DDN, Kioxia, and Micron to standardize GPU-driven storage solutions for AI workloads.

News article

NVIDIA Open-Sources cuFile API for Direct GPU Storage Access

NVIDIA announced at the Future of Memory and Storage conference that it is open-sourcing its cuFile API and the vertical storage software stack underneath it, marking a significant shift in how AI storage infrastructure operates

1

. The cuFile API, a core component of NVIDIA GPUDirect Storage, enables GPUs to read from and write to storage directly without CPU intervention, accessing data from storage in just microseconds

1

3

. Google, Intel, and Meta have joined as inaugural maintainers of this open-source initiative, creating a new home for APIs that can be optimized across various software and hardware platforms

1

. Intel's participation is particularly notable, as a leading supplier of x86 silicon in storage controllers has now signed on to maintain software designed to remove that silicon from the I/O path

2

.

Storage-Next Initiative Unites 40+ Industry Leaders

NVIDIA launched Storage-Next, an industry-wide initiative bringing together over 40 leading storage and flash vendors including DataDirect Networks, Kioxia, and Micron

1

3

. The initiative aims to align storage makers, controller vendors, thermal design, cooling and orchestration operators, and standards bodies on how GPU-driven storage should behave, then transform these advancements into interoperable, open industry standards

1

. DDN Chief Technology Officer Sven Oehme stated, "AI success will be defined not by how much infrastructure organizations own, but by how productively they use it"

2

3

. Partner systems from DDN, Dell, HPE, IBM, VAST Data and WEKA are due in the second half of 2026, while Kioxia's XL-Flash drives built for 512-byte access are in development under Storage-Next

2

.

SCADA Framework Addresses AI Workloads Bottleneck

The Storage-Next initiative is grounded in NVIDIA's SCADA framework—short for scaled, accelerated data access—which lets massively parallel GPUs pull only the data necessary for applications directly from storage into their own high-speed memory

1

. The SCADA framework moves the storage control path onto the GPU, allowing parallel GPUs to construct and complete their own storage requests and absorb per-operation latency

2

. This architecture prevents GPU starvation, where expensive accelerators sit idle waiting for data to be ready

3

. DDN is integrating SCADA with Infinia, its software-defined, AI-native data intelligence platform built to eliminate storage bottlenecks at scale

1

. SCADA also maintains security by splitting access across different jobs: user applications receive raw speed while staying outside the encrypted, secure computing base, while a separate privileged component configures protected access following Linux protocols

3

.

Solving the 512-Byte Problem for AI Inference

The push for GPU-initiated storage addresses a critical bottleneck in AI inference performance. Most enterprise SSDs have been tuned around 4KB random reads for nearly a decade to match virtualization and databases, but AI inference access patterns are considerably smaller—embeddings run a few hundred bytes, and KV cache blocks sit well under a kilobyte

2

. Serving 512-byte requests from 4K-tuned drives imposes approximately eightfold read amplification, and at tens of terabytes of small objects, that amplification determines whether flash works as a memory tier at all

2

. NVIDIA is focusing on 512-byte IOPS rather than raw bandwidth because inference performance depends heavily on serving high volumes of small random reads efficiently

2

. NVIDIA's roadmap calls for Gen7 SSDs sustaining 100 million IOPS each

2

.

Accelerated Computing Transforms Storage Economics

Benchmarks show that the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline

1

. With accelerated computing, storage stops being a passive place to keep data and becomes an active part of the data path, fundamentally changing the economics of when data belongs in memory versus on a storage drive

1

. The tradeoff between memory and storage access, first framed 40 years ago when measured in minutes, now plays out in microseconds on modern GPUs paired with AI storage solutions

1

. Using Direct Memory Access (DMA), cuFile moves data directly between NVMe storage and GPU memory without routing transfers through conventional CPU-managed system-memory buffers

4

. This creates a more efficient path for loading model weights, datasets, and other information that cannot remain permanently resident in GPU memory, though HBM and GDDR remain considerably faster with lower latency

4

. The development is immediately relevant to AI servers handling retrieval-augmented generation, agentic AI models, and large inference datasets, where even small reductions in idle periods can improve the economics of large GPU clusters

4

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved