Chinanews
Other爱范儿Thu, 27 Aug 2026 10:15:44 +0000

Why Apple and Xiaomi's AI Computers Are Obsessing Over the 'Memory Wall' | AI Artifacts

Why Apple and Xiaomi's AI Computers Are Obsessing Over the 'Memory Wall' | AI Artifacts

Smartphones have dominated the digital ecosystem for over a decade. They are attention black holes and our most intimate personal belongings. But from their very inception, phones were designed for "humans staring at them"—all their logic terminates at the screen.

AI demands the exact opposite: it requires continuous perception of the physical world—seeing what you see, hearing what you hear, and being constantly present, rather than waiting for you to unlock the screen to wake up.

When AI truly becomes a foundational capability, it will inevitably break out of the screen and find a form of its own. This will be a long process of exploration and evolution.

This is where the "AI Artifacts" column comes in. ifanr wants to continuously observe with you: How will AI change hardware design? How will it reshape human-computer interaction? And more importantly—what form will AI take as it enters our daily lives?

This is the 23rd article in the "AI Artifacts" series.

The new Macs Apple recently released have once again reset performance expectations for "AI computers." The compact Mac mini has been officially positioned as a productivity tool capable of running Agents around the clock, while the professional-grade Mac Studio is offered with up to 512GB of unified memory.

However, their true "strength" lies hidden in a parameter that few people cared about in the past—memory bandwidth.

The unified memory bandwidth of the standard M6 reaches 170GB/s, the M5 Pro version pushes it to 307GB/s, and the top-tier M5 Ultra directly crosses the TB threshold, hitting 1.2TB/s.

To complement these new products, Apple even filmed a "bullish" promotional video for the Mac mini, depicting the small silver box growing muscular arms.

Coincidentally, just one day before the new Mac Studio was released, Xiaomi announced its on-device AI accelerator chip, the Xuanjie O100, which also achieved a memory bandwidth of 1.22TB/s. Meanwhile, Nvidia's latest consumer flagship graphics card, the RTX 5090, leverages 32GB of GDDR7 VRAM to push bandwidth to nearly 1.

8TB/s.

The architectural paths of desktop SoCs, dedicated on-device NPUs, and discrete GPUs are vastly different. Yet, whether they make computers, phone chips, or graphics cards, everyone is unanimously doing the exact same thing—making the data in memory run faster.

AI Computes Faster, but Increasingly Goes "Hungry"

Chipmakers are putting so much effort into fighting for bandwidth fundamentally because today's AI can compute more than ever, but it is also increasingly easily "starved" by memory.

Many people's first reaction here is often: Isn't memory just about capacity? Bigger is better!

In reality, capacity and bandwidth are entirely different things. Capacity determines how large a model this computer can hold, while bandwidth determines how fast these data can be fed to the computing cores per second.

Once a large model starts running, the most resource-consuming part is often not the pure computation, but the continuous "data moving."

Image | LLM Prefill-Decode Illustration | NVIDIA

When we see words popping out one by one in the chat box, during the autoregressive generation (Decode) phase, every time the computing unit generates a new token, it must completely read through billions or tens of billions of parameter weights in memory.

It's like a chef in the kitchen with extremely fast hands; stir-frying only takes two minutes, but the prep cook takes forever to bring a plate of chopped ingredients. No matter how fast the chef cooks, he spends most of his time just holding the spatula, waiting at the stove.

Whether it's a CPU, GPU, or NPU, they are all stuck at this hurdle. Computing power is soaring, but the widening speed of memory channels cannot keep up. When the model is small, it can still manage, but once it scales to tens or hundreds of billions of parameters, spitting out a single word requires swallowing and spitting out massive amounts of data.

No matter how capable the computing units are, they can only waste a lot of time waiting for memory to feed them. This phenomenon is known in computer architecture as memory-bound.

Image | Compiled from public internet information

Faster Memory is Also More Expensive

When many people look at memory prices, their intuition is often "32GB is definitely more expensive than 16GB, and 512GB is naturally more expensive than 32GB." But if we look into the semiconductor supply chain, the real rule is: larger memory is more expensive, but faster memory gets exponentially more expensive.

To make massive data run faster, you can't just rely on stacking chips; you need wider bit widths, higher-frequency interfaces, and extremely complex advanced packaging and vertical stacking. Take HBM (High Bandwidth Memory), standard for today's AI servers: it vertically drills and stacks multiple layers of DRAM chips (TSV), and its manufacturing difficulty and chip area far exceed traditional memory.

This extreme thirst for ultra-high bandwidth is triggering a global "capacity migration" upstream.

According to TrendForce data, due to the surge in demand from AI data centers for HBM and high-spec server DDR5, the three memory giants—Samsung, SK Hynix, and Micron—are going all out to tilt limited advanced wafer capacity toward enterprise-grade products.

This directly squeezes the supply space left for consumer PC DRAM. Even in the traditional off-season, contract prices continue to rise month-over-month.

At the end of last year, Micron officially announced it would gradually exit the Crucial consumer memory business, pivoting supply chain resources entirely toward larger-scale, higher-margin strategic customers. Recently, the termination of operations by SK Hynix's third-party authorized brand stores in China also reflects the upstream original manufacturers' cold shoulder toward consumer channels.

Image | Micron

When ordinary PC assemblers hesitate over prices, they look up and realize: their budget to add a memory stick is invisibly competing with global AI data centers for the same batch of advanced wafers.

To prevent computing power from getting stuck on data moving,