Why does a GPU with free memory still refuse new requests?
the paper →Efficient Memory Management for LLM Serving with PagedAttentionKwon et al., SOSP 2023 ↗The paper and the system are the same artifact — this is vLLM. Its waste breakdown is the number this page is built on.Reserved memory is spent memory.
A server with two hundred gigabytes free will tell you it is full. Not because the memory is used — because it is promised. Every request that arrives is handed room for the longest reply it could ever produce, then writes four hundred words and leaves. The rest sits reserved, holding nothing, for the life of the request. PagedAttention is what happened when someone stopped booking the whole hotel for every guest.
On the machineThis lands on HBM, station 3 of 10 on the path the constraint took through the hardware. Spends the pool on tokens that exist rather than on lengths a request was permitted. See it on the drawing →Read as far as you want. Each level assumes the one above it and nothing more.
A sequence’s keys and values grow one token at a time, but a contiguous allocation has to be sized before the first token exists. So it is sized for the longest sequence the server permits, and the gap between what was reserved and what is used is dead memory — usually most of it. PagedAttention gives up contiguity: the cache lives in fixed sixteen-token blocks scattered anywhere in the pool, with a per-sequence table saying where they are. The allocation becomes the size of the sequence instead of the size of the promise.