NVIDIA released CUDA Toolkit 13.4, marking the first time CUDA officially extends native Arm development from Linux to Windows. AI developers can now compile CUDA applications directly on Windows Arm64 systems and gain early developer access to the Vera Rubin architecture ahead of RTX Spark's October launch.

News article

NVIDIA Extends CUDA Development to Windows on Arm64

NVIDIA has released CUDA Toolkit 13.4, introducing native Windows on Arm64 development support for the first time in CUDA's history

1

2

. This marks a significant expansion as CUDA officially extends native Arm development from Linux to Windows platforms. AI developers can now compile CUDA applications directly on Windows Arm64 systems, eliminating previous platform limitations. The release also supports cross-compilation from x86-64, allowing existing development systems to target Arm-based Windows hardware without requiring immediate hardware upgrades

1

.

The timing aligns strategically with NVIDIA's RTX Spark platform, scheduled to launch in October. CUDA support for Windows-on-Arm helps developers prepare Windows Arm64 applications now and allows CUDA to run on existing Windows on Arm machines

2

. This preparation window gives AI developers critical time to port applications, validate dependencies, and test CUDA software paths before RTX Spark systems reach the market.

Early Developer Access to Vera Rubin Architecture

CUDA Toolkit 13.4 provides early developer access to NVIDIA's Vera Rubin architecture through compute capability 10.7

1

2

. This preview support enables developers and software-library maintainers to begin porting, compiling, and testing applications before Rubin hardware becomes widely available. NVIDIA describes the Rubin architecture as powering the era of agentic AI, positioning it as the next-generation GPU computing platform

2

.

The early access proves particularly important for maintainers of key libraries who need extended lead time to prepare for future NVIDIA platforms. Core math libraries in CUDA Toolkit 13.4 now have functional support for the Rubin GPU architecture, allowing developers to validate their code paths well ahead of general availability

2

.

RTX Spark Platform Preparation

The update directly supports NVIDIA's upcoming RTX Spark platform, described as a consumer PC processor featuring a 20-core Grace CPU and a Blackwell GPU with 6,144 CUDA cores

1

. NVIDIA claims up to 1 PFLOPS of FP4 AI performance and local operation with AI models containing up to 120 billion parameters. With RTX Spark laptops arriving in October, CUDA 13.4 ensures developers have the necessary tools to optimize and validate their applications for the platform at launch

2

.

Platform-Wide Updates Across CUDA Ecosystem

CUDA Toolkit 13.4 updates several areas of the platform, including the compiler, PTX ISA, CUDA C++, CUDA Tile programming, runtime APIs, software libraries, and developer tools

1

2

. CUDA Python's cuda.core component reaches version 1.1.0, adding texture and surface programming, NUMA-aware managed memory, and Python type stubs. The cuda.compute 1.1 update adds support for precompiling CCCL algorithms for multiple GPU architectures

1

.

Multi-Process Service V3 introduces a modernized control layer with a scriptable command-line interface, named server instances, TOML configuration, and cgroup-integrated GPU memory limits for precise GPU partitioning in containerized environments

1

2

. CUDA Compute Fabric Transport adds named logical endpoints and asynchronous operations for data transfers across NVLink-connected systems, offering a transport-centric API for communication-library developers

1

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved