▶ See it in 3D The Memory Stack region
On-chip scratch paper the GPU scribbles on so it can avoid the slow trip out to main memory. Crucially, it has almost stopped getting cheaper each generation — which is a big reason the industry leans so hard on HBM instead.
L1/shared (256 KB/SM) and a 50 MB L2 shared across the die. FlashAttention exists entirely to keep data here and avoid touching HBM.
SRAM does not scale with Moore's law anymore — the reason designers lean ever harder on HBM bandwidth. Again captured by TSMC + the GPU designers.

Makes the GPU die AND the CoWoS packaging that bonds HBM to it — and now the logic base die under SK Hynix's HBM4 too. CoWoS is heading for ~130k wafers/month by end-2026 and is still sold out into 2027 (NVIDIA alone holds ~60%); it is farming 240–270k wafers/year of front-end steps out to Amkor and SPIL. A Taiwan disruption stops the whole stack.
Buys the most HBM, NVLink and optics on earth and pre-buys supply to lock rivals out: ~60% of CoWoS, $2B each into Coherent and Lumentum (Mar 2026), $2B into Marvell. B300 (288 GB HBM3E) is the shipping reference GPU; Rubin (288 GB HBM4, 22 TB/s) samples Q4 2026, volume Q1 2027.
MI-series GPUs with leading HBM capacity; the anchor of UALink. Zen 6 EPYC "Venice" launched July 2026 with roughly a year before Intel answers; MI400 carries the first UALink production silicon (H2 2026).