Why is attention slow when the GPU is barely doing any arithmetic?
the paper →FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessDao et al., NeurIPS 2022 ↗The original. Section 3 is the whole idea — tiling with an online softmax so the N×N matrix never exists in HBM.Moving data can cost more than computing with it.
Imagine doing long multiplication on a whiteboard the size of a football field, when the working would have fitted on a napkin. You would spend the day walking, not multiplying. That is what standard attention does: it builds a table with a hundred million entries, writes the whole thing to the far side of the chip, walks back to read it, and then throws it away. FlashAttention does the same arithmetic without ever writing the table down.
On the machineThis lands on HBM, station 3 of 10 on the path the constraint took through the hardware. Removes an N×N table from HBM entirely, so prefill stops competing for bandwidth. See it on the drawing →Read as far as you want. Each level assumes the one above it and nothing more.
Attention never needs its N×N table of scores as an output — only the weighted sum of values that comes out the far side. A standard implementation writes that table to main memory anyway, then reads it back, twice. FlashAttention computes the same result in tiles small enough to stay on the chip, carrying a running softmax so no tile ever needs to see its neighbours. The arithmetic is identical, and slightly increased. The traffic is not.