Recently, Xiaomi's Xuanjie O100 AI acceleration chip and Apple's new Mac mini and Mac Studio were unveiled one after another. Both have pushed memory bandwidth to the TB/s level, signaling that the battle to feed GPUs with data has begun.
This is no coincidence; this trend aligns perfectly with NVIDIA's RTX 5090 released in early 2025—whether it's desktop SoCs, edge NPUs, or discrete GPUs, all are pushing memory bandwidth to new heights to solve the bottleneck of data movement during the operation of large AI models.
Memory bandwidth is the width of the channel between computing cores and data. During the autoregressive generation phase of large language models, every token output requires a complete read of hundreds of billions of parameter weights from memory.
No matter how powerful the compute units are, they spend most of their time waiting for data to arrive. This "memory-bound" phenomenon has become a key bottleneck restricting AI performance.
Laura Metz, Apple's Director of Mac Product Marketing, stated that the M6 series has increased memory bandwidth several times over, as models must reside in unified memory and be ready at all times when running AI Agents.
The three-tier configuration of Apple's M6 series precisely marks three thresholds for AI hardware memory bandwidth. The standard version offers 170 GB/s, the Pro version 307 GB/s, and the top-tier Ultra version breaks through to 1.2 TB/s.
Xiaomi's Xuanjie O100 uses wafer-level vertical stacking technology, placing DRAM wafers directly on top of the NPU, Achieving 1.22 TB/s bandwidth.
NVIDIA's RTX 5090 achieves this through a 512-bit bus width combined with GDDR7 VRAM, Reaching nearly 1.8 TB/s.
All three companies are pushing bandwidth to the TB/s range. Although their approaches differ, they point to the same conclusion: The speed at which data is fed to computing cores determines the true performance of AI hardware.
This race has made memory bandwidth a more critical AI performance metric than core count. Whoever gains an advantage in data transfer speed will take the lead in the next computing era.

