Huawei Pulls Ascend 960 AI Chip Launch Forward by Three Quarters, Doubling Performance Targets

14 Sources

Share

Huawei announced at its Connect 2026 conference that its next-generation Ascend 960DT AI chip will launch in Q1 2027, three quarters ahead of schedule. The accelerated timeline comes as China races to build domestic semiconductor capabilities despite U.S. restrictions, with the chip delivering 4 FP4 PFLOPS and 288GB of memory to compete with Nvidia in AI computing.

Huawei Accelerates AI Chip Timeline to Challenge Nvidia

Huawei has moved the launch of its next-generation Ascend 960DT AI chip to the first quarter of 2027, three quarters earlier than originally planned, as announced at the Huawei Connect conference in Shanghai

1

. David Wang, Huawei's rotating and acting chairman, revealed the updated AI accelerator roadmap during his keynote, signaling the company's aggressive push to compete with Nvidia in the AI computing market

2

. The Ascend 960PR, optimized for inference workloads, will follow in Q3 2027, also ahead of its original schedule

4

.

Source: The Next Web

Source: The Next Web

The announcement comes just days before U.S. President Trump and Chinese President Xi Jinping are scheduled to meet on September 24 in Washington, DC, highlighting the geopolitical tensions surrounding semiconductor technology

1

. China tech analyst Rui Ma noted that U.S. semiconductor restrictions are unlikely to stop China's pursuit of self-sufficiency, stating the stakes are "just too high at this point"

1

.

Ascend 960 Specifications Double Performance Expectations

The Ascend 960DT delivers 2 FP8 PFLOPS and 4 FP4 PFLOPS, doubling the performance of Huawei's current Ascend 950 series

2

. The chip carries 288GB of HiZQ memory with 9.6TB/s bandwidth and features a 2.2TB/s interconnect

2

. What's more remarkable is that Huawei increased the FP4 performance of the Ascend 960PR to 8 FP4 PFLOPS, twice higher than originally announced last year

2

.

Source: New Atlas

Source: New Atlas

These next-generation Ascend NPUs represent a significant leap in China's domestic AI capabilities. While the 960DT offers similar memory capacity to Nvidia's B300 family, it delivers about half the FP8 and a third of the FP4 compute compared to American GPUs

3

. However, Nvidia cannot sell its most advanced chips in China, making Huawei's offerings the best available option for Chinese AI developers

3

.

Peerium Computing Architecture Scales to 4,096 Chips

Huawei unveiled its Peerium Computing Architecture, which relies on UnifiedBus technology to link processors with memory, storage, and networking hardware into large-scale AI computing systems

1

. The company's Atlas 960E SuperPoD, built on this architecture, can connect up to 4,096 neural processing units and delivers 8 EFLOPS at FP8 precision with one petabyte of HBM memory per pod

5

.

Source: Wccftech

Source: Wccftech

However, the scale has raised questions. China tech analyst Rui Ma pointed out that Huawei had previously stated its Atlas 960 SuperPoD would scale to 15,488 Ascend 960 chips, while this week's announcement referenced a system with 4,096 chips

1

. "The chip itself is coming WAY earlier, but the SuperPoD they announced is much smaller than what they originally laid out," Ma wrote

5

.

Near-Packaged Optics Cut Power and Boost Reliability

The Atlas 960E SuperPoD incorporates near-packaged optics through Huawei's Hi-ONE optical engine, which moves 7.2Tbit/s per unit

5

. By deploying 5,500 Hi-ONE units, Huawei eliminated 48,000 800G pluggable optical modules from the pod, cutting power consumption by more than 550 kilowatts and doubling system uptime to 99.8%

3

. The company has filed an implementation agreement on this technology with the Optical Internetworking Forum

5

.

Huawei's UnifiedBus architecture consolidates more than ten protocols into one, increasing bandwidth from hundreds of gigabytes per second to terabytes while reducing latency from seven microseconds to two

5

. In conventional 100,000-NPU clusters, as little as 20% of compute goes to the model, with the rest sitting idle during data movement. Huawei's lab simulations show pods of 4,000 NPUs reach 2.75 times the model FLOPs utilization of eight-NPU servers

5

.

Roadmap Extends Through 2029 with Ascend 970 and 980

David Wang announced that Huawei will maintain a one-generation-per-year cadence for its AI accelerators starting with the Ascend 960 series

2

. The Ascend 970 is scheduled for 2028 with 3.6 FP8 PFLOPS, 14 FP4 PFLOPS, 288GB of memory providing 14.4TB/s, and 4.4TB/s of interconnect bandwidth. The Ascend 980 follows in 2029 with 7.2 FP8 PFLOPS and 28 FP4 PFLOPS, along with 384GB of memory reaching 38.4TB/s and an 8TB/s interconnect

2

.

Wang credited the Tau Scaling Law for the accelerated cadence, noting that compute specifications will continue to double with huge improvements in memory bandwidth, memory capacity, and interconnect bandwidth

2

. Huawei has also revealed plans to scale compute clusters to half a million or more NPUs, with theoretical configurations reaching up to one million processors

5

.

Supply Constraints Remain Critical Challenge

Despite the technical advances, Huawei rotating chairman Eric Xu acknowledged that production capacity remains insufficient to satisfy domestic demand

2

. Executives told Bloomberg that Huawei now holds a bigger AI chip market share than Nvidia inside China but cannot manufacture enough units

5

. Xu stated the company is prioritizing Chinese customers and has no broad global expansion plan, projecting China will catch up with its own hardware demand by 2030

5

.

Huawei recently raised the price of the Ascend 950DT by 60% citing tight component supply

5

. DeepSeek alone plans to deploy at least 160,000 of those chips, more than Huawei can currently fulfill

5

. Wang revealed that Huawei has shipped more than 1,000 UnifiedBus-based systems to more than 370 customers

4

, though the Atlas 950 SuperPoD adoption does not appear to be proceeding rapidly, possibly due to insufficient supply or the all-new architecture requiring major software redesign

2

.

Software Ecosystem Gains Ground with CANN Framework

Huawei's CANN framework, which serves a similar role to Nvidia's CUDA for Ascend processors, has moved to community-driven open source

5

. External developers now make up 61% of the CANN community, outnumbering Huawei's own developers for the first time

5

. Ascend supports more than 90 third-party open-source projects and has become an official PyTorch accelerator backend, making it the first Chinese compute platform listed on the PyTorch website

5

. The Kunpeng ecosystem claims 4.16 million developers and 7,200 partners

5

, signaling progress in closing the software gap that has historically favored Nvidia's dominance.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved