3DMGame

The War to Feed GPUs Begins: Xiaomi, Apple, and Nvidia Race to Boost Memory Bandwidth

Summary

Recently, Xiaomi's Xuanjie O100 AI accelerator chip and Apple's new Mac mini and Mac Studio were unveiled, both pushing memory bandwidth to TB/s levels, signaling that the battle to "feed the GPU" has begun. This trend aligns with Nvidia's RTX 5090 released in early 2025—whether in desktop SoCs, on-device NPUs, or discrete GPUs, memory bandwidth is being pushed to new heights to solve the data movement bottleneck in running large AI models. Memory bandwidth is the width of the channel between the compute core and data. During autoregressive generation, large language models must read the weights of hundreds of billions of parameters from memory for every token output. No matter how powerful the compute unit is, most of the time it is waiting for data to arrive. This "memory-bound" phenomenon has become a key bottleneck limiting AI performance. Laura Metz, Apple's Mac product marketing director, stated that the M6 series has multiplied memory bandwidth to better handle AI models. This is not just a hardware spec war; it deeply impacts the future of gaming and AI applications. For players, higher memory bandwidth means smoother ray tracing, faster high-resolution texture loading, and more complex physics simulations. For the industry, co-optimizing compute and bandwidth is becoming a core chip design focus. As AI models continue to grow, the memory bandwidth race will only…

Steam

Read original