NVIDIA Open Sources cuFile API to Enable Microsecond GPU Data Access for AI Storage

2 Sources

Share

NVIDIA announced it is open sourcing its cuFile API at the Future of Memory and Storage conference, enabling GPUs to read and write directly to storage in microseconds. The company also launched Storage-Next, an initiative uniting over 40 storage vendors including DDN, KIOXIA and Micron to address mounting AI demands on memory and develop GPU-driven storage solutions.

NVIDIA Tackles AI Demands on Memory with Open Source cuFile API

NVIDIA announced at the Future of Memory and Storage conference that it is open sourcing its cuFile application programming interfaces (APIs), a move designed to revolutionize how GPUs access storage systems

1

. The open-source cuFile API, a core component of NVIDIA GPUDirect Storage, enables GPUs to read from and write to storage directly without CPU intervention, achieving data access in just microseconds

2

. This development addresses the escalating AI demands on memory as applications consume massive datasets that exceed system memory limits

1

.

The technology leverages hundreds of thousands of GPU threads and fast high-bandwidth memory to enable secure, direct GPU-to-storage data access

1

. By allowing GPUs to bypass the CPU and system main memory through direct memory access (DMA), cuFile moves data directly from storage devices like NVMe drives into GPU memory, drastically reducing latency

2

. Making cuFile openly available supports interoperability between GPUs and data while helping make security context and storage accessible at the speed AI-powered defenses require

1

.

Storage-Next Initiative Unites Industry Leaders for GPU-Driven Storage Solutions

Alongside the cuFile announcement, NVIDIA launched Storage-Next, a large-scale industry initiative bringing together over 40 leading storage and flash vendors to optimize memory and storage solutions for accelerated computing

1

. The Storage-Next initiative includes prominent partners such as DataDirect Networks (DDN), Kioxia and Micron, each contributing to next-generation AI storage technologies

2

.

The initiative unites storage makers, controller vendors, thermal design experts, cooling providers, orchestration operators and standards bodies to align on how GPU-driven storage solutions should function

1

. "AI success will be defined not by how much infrastructure organizations own, but by how productively they use it," said Sven Oehme, chief technology officer at DDN

2

. DDN is integrating SCADA with Infinia, its AI-native data intelligence platform built to eliminate storage bottlenecks at scale

1

.

SCADA Framework Enables Massively Parallel GPU Data Access

At the core of Storage-Next lies SCADA—scaled, accelerated data access—a framework that enables massively parallel GPUs to pull only necessary data directly from storage into their high-speed memory

1

. SCADA provides the groundwork for high-bandwidth, extremely low-latency data access essential for modern AI training and inference workloads

2

.

Source: NVIDIA

Source: NVIDIA

The framework addresses a critical challenge in AI infrastructure: GPU starvation, where GPUs sit idle waiting for data, causing bottlenecks in distributed arrays

2

. With high-speed storage operating at microsecond speeds, massive datasets required for retrieval-augmented generation and agentic AI models can run faster through a low-latency connection between deep storage and GPU memory

2

.

Accelerated Computing Transforms Storage Economics

The pressure on AI infrastructure is intensifying as AI agents consume massive amounts of data and GPUs now initiate storage requests directly, generating thousands of concurrent operations

1

. Storage systems must continuously encrypt, compress, verify and reconstruct data—operations that become bottlenecks when thousands of agents access storage simultaneously

1

.

Benchmarks show that the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline

1

. With accelerated computing, storage transforms from a passive repository into an active part of the data path, fundamentally changing the economics of when data belongs in memory versus on storage drives

1

. What was once measured in minutes 40 years ago now plays out in microseconds on today's GPUs paired with AI storage solutions

1

.

Security-First Architecture Balances Speed and Protection

SCADA enables high-speed storage layers to operate securely by splitting access across different jobs: user application components receive raw speed while staying outside the encrypted, secure computing base, while a separate privileged computing component configures protected access following Linux protocols

2

. This design prevents data clobber—when two processes attempt to write to the same data simultaneously—while maintaining the security safeguards necessary for encrypted and privileged access

2

.

Fast, secure access to data and storage serves as a foundational element for powering preventive and detective cybersecurity measures

1

. The open-source approach supports initiatives such as the Open Secure AI Alliance, with Google, Intel, NVIDIA and Meta as inaugural maintainers optimizing technologies for use across various software and hardware platforms

1

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved