Pick a model class, a conversation length and how many people are chatting at once. The calculator fills one B300 (288 GB) with the model's weights, then adds each chat's KV cache — and tells you when you're buying whole extra GPUs.
▶ See the tanks fill in 3D Why memory is the wall
Weights are fixed per model class (fp16, fp8 for the 1T MoE class). The KV cache is bytes per token × context × chats — and bytes per token is the number that architecture decides: a GQA model like the 405B class stores ~504 KB per token, while an MLA model like the 1T MoE class compresses that to ~69 KB. A GPU is "full" at 288 GB; anything over spills into another GPU at roughly $45k each. Speed is memory-bound: a decode step must re-read weights plus the whole cache, so ms/token scales with what's resident.
The same numbers drive the 3D Wall — watch the tanks overflow, or read how the KV cache works word by word.