The Memory AtlasText edition▶  Open the 3D atlas
Text edition / The Memory Wall / KV-Cache Economics
why memory is the wall

KV-Cache Economics

why memory is the wall

▶  See it in 3D   The Memory Wall region

On the journey · stop 2 of 30

AI is bottlenecked by memory, not math. To write each word, a GPU re-reads the entire model from memory — and sits idle ~1,200× waiting on it. That one fact is what this whole map is about.

Decode
memory-bound
H100 starve
~1,200× wait
Fix #1
more HBM
Fix #2
CXL + photonics
In plain English

Here's why everything else on this map exists: when an AI writes an answer, it has to re-read its entire memory of the conversation for every single word it produces. That memory (the "KV cache") balloons with longer chats and more users, and it has to live in expensive HBM. More memory pressure → more demand for every company in this atlas.

What it is

LLM decode reads every weight plus the entire KV cache from HBM for each token. The GPU computes ~1,200× faster than it can fetch — so it idles on memory. Longer context and bigger batches explode the cache (drag the sliders).

The business

This is why the whole supply chain exists. It converts straight into demand for HBM (SK Hynix/Micron), CoWoS (TSMC), pooled CXL (Astera) and optical disaggregation (Marvell/Celestial). Every doubling of context is a doubling of memory demand.

Bottleneck severity

Severe
Where the chokepoint comes from

A demand chokepoint rather than a supply one: every doubling of context length or concurrent users doubles the memory a GPU must hold, and it has to be HBM-fast. That pressure lands directly on the three HBM makers and TSMC's packaging line — the tightest links in the whole chain.

Who makes it

NVIDIANVDA

GPUs, NVLink, NVSwitch, photonic switches · USA (fab TSMC)

Buys the most HBM, NVLink and optics on earth and pre-buys supply to lock rivals out: ~60% of CoWoS, $2B each into Coherent and Lumentum (Mar 2026), $2B into Marvell. B300 (288 GB HBM3E) is the shipping reference GPU; Rubin (288 GB HBM4, 22 TB/s) samples Q4 2026, volume Q1 2027.

Rubin ramp timingsupply lockupsNVLink Fusion breadth
↗ Investor relations & news

SK Hynix000660.KS

HBM3E / HBM4 stacks · South Korea

#1 HBM maker (~62% share). Qualified HBM4 at NVIDIA first and holds an estimated 60–70% of Rubin's HBM4 allocation, with a TSMC-made logic base die. Sold out through 2027; its CEO calls 2027 potentially the tightest supply year the industry has seen. The bellwether of the whole cycle.

Rubin HBM4 shareHBM4E / 16-Hi samplesM15X + Yongin ramp
↗ Investor relations & news

MicronMU

HBM3E / HBM4, DDR5, CXL modules · USA (fab Japan/Taiwan)

#3 HBM, the only US-HQ DRAM maker. Qualified on Rubin HBM4; sold out through 2027 and meeting only ~50–65% of what customers ask for. New Idaho capacity doesn't land until 2027–28 — so the squeeze is structural, not seasonal.

HBM bit-sharefill rate vs requestsHBM gross margin
↗ Investor relations & news

Astera LabsALAB

Retimers, CXL controllers, fabric switches · USA (fab TSMC)

The only large-cap pure-play on AI connectivity — Aries retimers, Leo CXL controllers, Scorpio fabric switches. Q2 2026 revenue $392M (+104% YoY) as Scorpio X entered volume; Q3 guided $540–560M. Building NVLink Fusion hybrid racks for 2027. Risk: one customer is most of revenue.

Scorpio X volumeQ3 guide $540–560Mcustomer concentration
↗ Investor relations & news

MarvellMRVL

Optical DSP, custom XPU, Photonic Fabric · USA (fab TSMC)

~70% optical-DSP share + custom AI silicon + Celestial AI photonic memory fabric ($3.25B). Most complete optical stack in public equities.

DSP sharecustom-silicon winsCelestial run-rate
↗ Investor relations & news

Metrics that move this layer

Where it comes from

—a demand signal, not a place — but it pulls Korea HBM + Taiwan CoWoS

Connected to

HBM3E / HBM4CXL MemoryNetwork / FabricThe Memory WallInside the KV CacheHow the KV Cache Works

Further reading

Epoch AI — AI chip supply-chain constraints ↗SemiAnalysis — Co-Packaged Optics ↗