Skip to main content

Impressive! Huawei Ascend AI Chips Surpass NVIDIA in China Market Share

Recently, Huawei's Rotating Chairman Xu Zhijun stated that Ascend chips have surpassed NVIDIA in market share in China.

Huawei is focusing on its SuperPoD solution, centered around the Peerium computing architecture and UnifiedBus high-speed bus technology, which can support collaboration among millions of processors.

厉害了!华为昇腾AI芯片中国份额已超过英伟达

According to institutional statistics and the latest data released at the Huawei Connect conference, Huawei's Ascend series AI chips have reached a 50% market share in China, surpassing NVIDIA's 8%.

Compared to the situation in 2022 where NVIDIA held 95% of the Chinese market share and Ascend only accounted for 3%, the market landscape has undergone a significant reversal over the past four years.

Xu Zhijun admitted that, when looking at individual chips, Huawei's bottleneck lies in the limited availability of advanced process nodes; however, by interconnecting a large number of chips to form supercomputers, Huawei has achieved a global leading position.

Huawei's new Peerium computing architecture and UnifiedBus (Lingqu Bus) technology can support up to one million processors in peer-to-peer interconnection, breaking the traditional master-slave design paradigm and significantly improving Model FLOPs Utilization (MFU) for large model training.

Regarding the software ecosystem and chip cluster scale, Dr. Liao Heng, Chief Scientist at HiSilicon, pointed out that two years ago, the biggest obstacle for Ascend was the software barrier formed by PyTorch and CUDA.

However, with the increase in model mathematics and programming complexity over the past 18 months, along with the emergence of new compiler languages, frontier labs are no longer overly reliant on CUDA, and the software gap is narrowing significantly or even reaching equilibrium.

Furthermore, regarding the Atlas 950 SuperPoD super node with 256,000 computing cards currently being deployed and tested by Huawei, Liao Heng revealed that this scale is designed to support the training of foundation models with 10 to 40 trillion parameters for six to seven frontier AI labs in China over the next two years. It is also a rational choice considering the capacity of data center power grids and electricity transmission.