▶ See it in 3D The KV Cache region
Think of the newest word as someone walking into a room asking "who here matters to me?" Everyone already in the room holds up a card (their key) and an offer (their value). Those cards never change — so the model keeps them in a stack instead of asking everyone to rewrite them for every new arrival. That stack is the KV cache: one new card per word, forever growing, living in the GPU's fastest memory.
When a model writes a word, that word becomes a query — it asks "who here matters to me?" Every earlier word answers through its key (how relevant am I?) and contributes through its value (what do I offer?). Those K/V vectors never change once computed — so the model stores them. Each new word computes exactly one new pair and appends it; the prompt itself is "prefilled" in a single pass. Without the cache, every single step would recompute K and V for the entire conversation — the same numbers, again and again.
This one engineering trick is why long conversations are possible — and why they are expensive. The cache turns quadratic recomputation into linear growth, but that linear growth lands in HBM, the scarcest memory on earth. Every extra word of context is a permanent tenant in the GPU's memory until the chat ends. Open The Wall to see what that does to GPU counts.
Not a supply chokepoint — a demand engine. The cache is pure memory pressure: it must live in HBM (nothing slower keeps up with decode), it grows with every word and every user, and it can't be shared between chats. It is the single biggest reason inference wants more memory every year.
Buys the most HBM, NVLink and optics on earth and pre-buys supply to lock rivals out: ~60% of CoWoS, $2B each into Coherent and Lumentum (Mar 2026), $2B into Marvell. B300 (288 GB HBM3E) is the shipping reference GPU; Rubin (288 GB HBM4, 22 TB/s) samples Q4 2026, volume Q1 2027.

#1 HBM maker (~62% share). Qualified HBM4 at NVIDIA first and holds an estimated 60–70% of Rubin's HBM4 allocation, with a TSMC-made logic base die. Sold out through 2027; its CEO calls 2027 potentially the tightest supply year the industry has seen. The bellwether of the whole cycle.
#3 HBM, the only US-HQ DRAM maker. Qualified on Rubin HBM4; sold out through 2027 and meeting only ~50–65% of what customers ask for. New Idaho capacity doesn't land until 2027–28 — so the squeeze is structural, not seasonal.