The Memory AtlasText edition▶  Open the 3D atlas
Text edition / The KV Cache / How the KV Cache Works
attention · step by step

How the KV Cache Works

attention · step by step

▶  See it in 3D   The KV Cache region

Per new word
1 K/V pair
Without cache
recompute ALL
Prompt
prefilled in 1 pass
Growth
linear with context
In plain English

Think of the newest word as someone walking into a room asking "who here matters to me?" Everyone already in the room holds up a card (their key) and an offer (their value). Those cards never change — so the model keeps them in a stack instead of asking everyone to rewrite them for every new arrival. That stack is the KV cache: one new card per word, forever growing, living in the GPU's fastest memory.

What it is

When a model writes a word, that word becomes a query — it asks "who here matters to me?" Every earlier word answers through its key (how relevant am I?) and contributes through its value (what do I offer?). Those K/V vectors never change once computed — so the model stores them. Each new word computes exactly one new pair and appends it; the prompt itself is "prefilled" in a single pass. Without the cache, every single step would recompute K and V for the entire conversation — the same numbers, again and again.

The business

This one engineering trick is why long conversations are possible — and why they are expensive. The cache turns quadratic recomputation into linear growth, but that linear growth lands in HBM, the scarcest memory on earth. Every extra word of context is a permanent tenant in the GPU's memory until the chat ends. Open The Wall to see what that does to GPU counts.

Bottleneck severity

Moderate
Why it's not a chokepoint

Not a supply chokepoint — a demand engine. The cache is pure memory pressure: it must live in HBM (nothing slower keeps up with decode), it grows with every word and every user, and it can't be shared between chats. It is the single biggest reason inference wants more memory every year.

Who makes it

NVIDIANVDA

GPUs, NVLink, NVSwitch, photonic switches · USA (fab TSMC)

Buys the most HBM, NVLink and optics on earth and pre-buys supply to lock rivals out: ~60% of CoWoS, $2B each into Coherent and Lumentum (Mar 2026), $2B into Marvell. B300 (288 GB HBM3E) is the shipping reference GPU; Rubin (288 GB HBM4, 22 TB/s) samples Q4 2026, volume Q1 2027.

Rubin ramp timingsupply lockupsNVLink Fusion breadth
↗ Investor relations & news

SK Hynix000660.KS

HBM3E / HBM4 stacks · South Korea

#1 HBM maker (~62% share). Qualified HBM4 at NVIDIA first and holds an estimated 60–70% of Rubin's HBM4 allocation, with a TSMC-made logic base die. Sold out through 2027; its CEO calls 2027 potentially the tightest supply year the industry has seen. The bellwether of the whole cycle.

Rubin HBM4 shareHBM4E / 16-Hi samplesM15X + Yongin ramp
↗ Investor relations & news

MicronMU

HBM3E / HBM4, DDR5, CXL modules · USA (fab Japan/Taiwan)

#3 HBM, the only US-HQ DRAM maker. Qualified on Rubin HBM4; sold out through 2027 and meeting only ~50–65% of what customers ask for. New Idaho capacity doesn't land until 2027–28 — so the squeeze is structural, not seasonal.

HBM bit-sharefill rate vs requestsHBM gross margin
↗ Investor relations & news

Metrics that move this layer

Where it comes from

—an algorithm, not a place — but its appetite lands on Korean HBM and Taiwanese packaging

Connected to

KV-Cache EconomicsHBM3E / HBM4The Memory Wall

Further reading

KV cache, explained intuitively — Saad Ahmed ↗DeepSeek-V2 paper — MLA compressed KV ↗A visual guide to attention variants — Sebastian Raschka ↗