Why this paper matters
Twelve papers that moved where the constraint in AI sits, from the Transformer in 2017 to the models of 2025. Each one opens with the question it answered.
trying to answer →What becomes scarce when intelligence becomes cheap?
A paper matters when it changes what is scarce. Each one below solved a real problem, and each solution moved the cost somewhere else, usually into memory. Read them in order and you watch the same few constraints get attacked again and again.
- 1Vaswani et al. · 2017Attention Is All You NeedWhy did transformers replace RNNs?Parallelism is what turns money into capability.3-minute brief · Transformer
- 2Shazeer · 2019Fast Transformer Decoding: One Write-Head is All You NeedWhy does a long conversation get slower and cost more than a short one?Previously computed state is worth more than recomputing it.explainer, runs in the page · lands on HBM
- 3Fedus, Zoph and Shazeer · 2021Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityWhy do some AI models have experts?What you can afford to hold and what you can afford to run stopped being the same number.3-minute brief · Mixture of experts
- 4Dao et al. · 2022FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessWhy is attention slow when the GPU is barely doing any arithmetic?Moving data can cost more than computing with it.explainer, runs in the page · lands on HBM
- 5Gu & Dao · 2023Mamba: Linear-Time Sequence Modeling with Selective State SpacesWhy does a model reread the whole conversation instead of just remembering it?A fixed memory can only keep what it chose to keep.explainer, runs in the page · lands on HBM
- 6Kwon et al. · 2023Efficient Memory Management for LLM Serving with PagedAttentionWhy does a GPU with free memory still refuse new requests?Reserved memory is spent memory.explainer, runs in the page · lands on HBM
- 7DeepSeek-AI · 2024DeepSeek-V3 Technical ReportHow did DeepSeek train a frontier model so cheaply?A constraint you cannot buy your way past becomes a research agenda.3-minute brief · DeepSeek
- 8Lieber et al., AI21 Labs · 2024Jamba: A Hybrid Transformer-Mamba Language ModelWhy does Jamba matter?AI21’s open-weight model, and the hybrid mamba.left predicted before it existed: one attention layer for every seven Mamba layers, with mixture-of-experts layered on top, holding 256k of context in a fraction of the memory a pure-attention model would need.a note · Jamba
- 9Ma et al., Microsoft Research · 2025BitNet b1.58 2B4T Technical ReportWhy does BitNet matter?Microsoft’s open-weight model, trained from scratch at ternary precision —a note · BitNet
- 10DeepSeek-AI · 2025DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningCan a small AI model learn to reason like a huge one?Reasoning ability can be trained against a checkable reward, not copied from a hand-written example of the reasoning itself.3-minute brief · DeepSeek-R1
- 11Kimi Team, Moonshot AI · 2025Kimi K2: Open Agentic IntelligenceWhy does Kimi K2 matter?Moonshot AI’s open-weight model —a note · Kimi K2
- 12Gemma Team, Google DeepMind · 2025Gemma 3 Technical ReportHow do models handle huge context windows without the memory bill exploding?Most words do not need to see the whole conversation to be predicted correctly.3-minute brief · Sliding-window attention
Every page opens with the question the paper answered, not the paper’s own title. Four are full explainers you can run, three levels deep, each pointing at the part of the machine it binds on in How electricity becomes intelligence. The rest are three-minute briefs, and a few are still only notes, marked as such.