The hardware behind autonomous machines


Sparse Change: SemiAnalysis Says GLM-5.3 Trims Attention Math, Not the Memory It Fills

SemiAnalysis’s GLM-5.3 teardown finds top-k sparse attention cuts bandwidth at the attention step while the full context sits in HBM and spills to host DRAM.

Sparse Change: SemiAnalysis Says GLM-5.3 Trims Attention Math, Not the Memory It Fills

SemiAnalysis has published a serving-level teardown of Z.ai's GLM-5.3 with a conclusion for anyone pricing HBM demand off model architecture: "Sparse attention reduces KV cache memory and bandwidth requirements at the SDPA operation, but it does not reduce the overall memory capacity usage."

The mechanism is DeepSeek Sparse Attention, which GLM-5 inherits from DeepSeek V3.2. A lightning indexer scores every token in the context and picks the top 2,048 for each query; below that length attention stays dense. The saving is in what the attention operation reads, not in what the system must hold. "Concretely, the top-k selection operation typically requires the full context to be in HBM, so sparse attention doesn’t eliminate the memory capacity bottleneck." Translated, the indexer cannot pick the relevant tokens without seeing all of them, so the full KV state still has to sit somewhere fast.

Where it sits is the DRAM story. SemiAnalysis walks through HiSparse, the SGLang team's hierarchical cache that treats HBM as a least-recently-used tier and spills KV entries to host DRAM, prefetching layer N while layer N-1 computes. Its B200 measurements, from SemiAnalysis's own InferenceX benchmark rows rather than a vendor deck, show the shift: as concurrency rises from 8 to 16 requests, the share of prompt tokens reused from GPU memory falls from 90.3% to 54.8% while reuse from host memory rises from 6.0% to 40.3%. "Much of the decline in GPU cache is offset by reuse from host memory, keeping the overall cache hit rate above 95% at all concurrency levels." In capacity terms the demand steps down one rung, from the scarcest memory to the next scarcest: a DRAM order rather than a cancelled one.

That is the counterweight to the argument this desk carried two weeks ago, when DeepSeek claimed V4.1-Flash needed a quarter of the HBM of its predecessor and Samsung and SK hynix fell more than 3% the next session. A model can cut what attention touches per token and still leave the server holding the same context, plus a host-memory tier it did not need before.

The hardware inference is the piece specialists will argue over. GLM-5 is a 744 billion parameter mixture-of-experts model with 40 billion active, routing each token through 8 of 256 experts plus one shared expert, and it runs 64 query heads where DeepSeek V3.2 runs 128. SemiAnalysis derives the attention step's arithmetic intensity from those dimensions and notes that DeepSeek's 128 heads land on the H800's ridge point while GLM-5's 64 land near 121 FLOP per byte. "Thus, we suspect GLM-5 is optimized for Moore Threads MTT S4000, which is about 128 FLOP/B." That is SemiAnalysis's inference, offered as suspicion and backed by Moore Threads' day-zero support for GLM-5.3-Flash, not a Z.ai statement; if it holds, a Chinese frontier lab is shaping attention geometry around a domestic GPU rather than Nvidia's.

Z.ai has also been cutting the indexer's own overhead. IndexShare, introduced in GLM-5.2, has every four sparse-attention layers share one indexer and cache one set of top-k indices. The figures SemiAnalysis relays come from Z.ai's IndexCache work on a 30 billion parameter test model, not from GLM-5.3 itself: "This design reduces indexer cache by 75%, reduces indexer FLOPs by 75%, and boosts throughput by 1.5x to 1.8x at different context lengths and prefill/decode."

On serving cost, SemiAnalysis's modeled estimates, built on its own assumptions of $2.31 per GB300 GPU hour and $1.86 for GB200, put GB200 near $0.044 per million total tokens at 150 tokens per second against $0.049 for AMD's MI355X on the ATOM engine, with the ranking flipping at 100 tokens per second and again under a two-second first-token limit. Neither vendor holds a uniform cost lead.

The piece's cybersecurity section sits behind the paywall and was not read. The memory conclusion is above it: sparse attention changes the shape of the memory bill, moving part of it from HBM to the DRAM behind the CPU, and none of the system tricks SemiAnalysis documents make the context smaller than it was.

Leave a Reply

Discover more from Autonomy Magazine

Subscribe now to keep reading and get access to the full archive.

Continue reading