Microsoft to deploy AMD Helios across Azure as chip rivalry with Nvidia intensifies

Reviewed byNidhi Govil

11 Sources

Share

Microsoft announced it will integrate AMD Helios, a rack-scale AI accelerator featuring 72 Instinct MI455X GPUs and 31TB of HBM4 memory, into Azure infrastructure. The system ships later this year, joining Meta, OpenAI, and Oracle as major customers. AMD's first rack-scale system directly challenges Nvidia's Vera Rubin NVL72, offering 2.9 exaFLOPS of FP4 inference compute and built on open standards to avoid vendor lock-in.

Microsoft Expands AI Infrastructure with AMD Helios Deployment

Microsoft will deploy AMD's Helios rack-scale AI system across its Azure cloud infrastructure, marking a significant expansion of the partnership between the two tech giants. The announcement positions AMD Helios as a direct competitor to Nvidia's dominant position in AI compute, with Microsoft CEO Satya Nadella stating the move will "give customers the performance, scale and choice they need to build and run the next generation of AI applications"

1

4

. While Microsoft and AMD didn't disclose exact financial terms or the precise scale of the deployment, the commitment represents another major win for AMD as it seeks to capture data center GPU market share from Nvidia

1

.

The new Helios system will power frontier model inference for Microsoft Azure, support Azure AI services, and provide managed compute for enterprise customers deploying AI workloads through Microsoft Foundry

1

. Microsoft joins Meta, OpenAI, Oracle, and others in adopting AMD's rack-scale AI system, with AMD reporting that eight of the top 10 AI companies now run workloads on its Instinct GPUs

2

.

AMD's Answer to Nvidia: Technical Specifications and Competitive Edge

Source: Wccftech

Source: Wccftech

AMD Helios represents the company's first rack-scale AI accelerator, designed to compete directly with Nvidia's Vera Rubin NVL72 system when it ships later this year. The system integrates 72 next-generation Instinct MI455X GPUs with an aggregate of 31.1TB of HBM4 memory capacity across the entire rack

1

3

. These GPUs deliver up to 1.4 exaFLOPS of FP8 compute and 2.9 exaFLOPS of FP4 for AI models, providing substantial processing power for both training and inference workloads

1

.

The architecture uses 18 compute trays, each holding four MI455X accelerators built on the new CDNA 5 architecture and one sixth-generation Epyc Venice CPU

3

. Each MI455X GPU carries 432 gigabytes of HBM4 memory with 19.6 TB/s of bandwidth

3

5

. AMD targets 260 TB/s of scale-up bandwidth within the rack, matching Nvidia's Vera Rubin NVL72, and 43 TB/s of scale-out bandwidth using UALink over Ethernet—approximately twice that of Vera Rubin, though real-world performance of UALink over Ethernet remains to be demonstrated

1

.

Open Standards Strategy Against Vendor Lock-In

AMD's architecture bet centers on open standards rather than proprietary technologies. AMD's AI-optimized Helios racks use UALink for scale-up interconnect between GPUs within the rack, Ultra Ethernet Consortium specifications for scale-out networking between racks, and the OCP Open Rack Wide form factor

3

. This contrasts sharply with Nvidia's competing NVL72, which relies on proprietary NVLink technology. AMD is betting that data centers seeking to avoid vendor lock-in will value the flexibility of open standards

3

.

AMD Pensando DPUs handle networking with programmable hardware and UEC-ready RDMA capabilities

3

5

. These Pensando data processing units are optimized for infrastructure management tasks such as coordinating storage equipment and encrypting network traffic, running such workloads more efficiently than CPUs and lowering costs while freeing up CPU capacity for customer applications

5

. Microsoft will leverage its existing deployment of AMD Pensando DPUs to integrate this hardware into Azure Boost offerings to accelerate networking and storage processing operations

1

.

The ROCm software stack supports PyTorch, TensorFlow, and JAX, meaning developers don't need to rewrite code to migrate from Nvidia's CUDA ecosystem, at least in theory

3

. Whether AMD can close the software gap that has kept it behind Nvidia in AI compute remains the critical question that AMD Helios is designed to answer

3

.

New Azure Virtual Machine Series for Specialized Workloads

Source: Guru3D

Source: Guru3D

Beyond deploying AMD Helios, Microsoft Azure will introduce three new virtual machine series powered by AMD silicon. The ND MI455X v7 series will run on Helios systems and is optimized for inference workloads such as artificial intelligence agents and search tools

5

. Microsoft also announced two new VM series built on AMD's upcoming sixth-generation Epyc Venice CPUs: the Azure HDv2 series aimed at agentic AI and data pipelines, and the Azure HXv2 series designed for semiconductor design workflows

1

4

.

The HDv2 series is optimized for tasks that AI applications carry out using CPUs rather than GPUs, including the process of preparing datasets for analysis by AI agents. Each instance includes up to 500 Epyc Vulcan cores, four terabytes of memory, and 32 terabytes of flash storage

5

. The HXv2 series represents an improved version of an existing Azure instance series optimized for electronic design automation applications that engineers use to design chips. Each HXv2 virtual machine features 176 Epyc Vulcan cores with clock speeds exceeding 5GHz, with each core featuring 50% more cache than previous-generation hardware

5

.

Timeline and Market Implications

Source: Tom's Hardware

Source: Tom's Hardware

Engineering samples of AMD Helios ship in the second half of 2026, with mass production beginning in Q2 2027

3

. AMD will start shipping to customers, including Microsoft, later this year

2

. This timeline positions AMD to compete more aggressively with Nvidia as demand for AI compute continues to grow, particularly in the wake of frontier-class open models that require massive computational resources for fine-tuning and serving

1

.

The 31TB of HBM4 memory in a single rack represents enormous materials cost that only hyperscaler and sovereign compute budgets can absorb, reflecting how the bottleneck in AI workloads has shifted from raw compute to memory capacity and bandwidth for frontier model training and long-context inference

3

. For Microsoft, which needs as much compute as possible as it ramps up its own model development and allocates more computing capacity to research and development, the partnership addresses critical infrastructure needs. The company announced seven models built in-house in June and continues to expand its AI efforts across products like 365 Copilot and GitHub Copilot

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved