Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Why this paper matters · 1 of 12

Why did transformers replace RNNs?

Transformer: Parallelism is what turns money into capability.

the paper →Attention Is All You NeedVaswani et al., NeurIPS 2017 ↗

The idea that let a model read a whole sentence at once instead of one word at a time. It is why training got fast enough to build ChatGPT and Claude — and every conversation since has come with a growing memory bill nobody has found a way to erase.

The room it came out of

In 2016 Google had just rebuilt Translate on eight stacked layers of recurrent networks. It translated better than anything before it and it was brutally expensive, because a recurrent network reads one word at a time and a warehouse of chips can only wait. Eight people set out to make that cheaper. The paper is benchmarked on English-to-German and English-to-French; there is not a word in it about chatbots.

How it works

Before 2017, a model that read text read it the way you do: one word, then the next, carrying a summary forward. That is a recurrent network, and it has a fatal property for anyone with a warehouse full of chips — step ten cannot begin until step nine has finished. You can own ten thousand processors and the sequence will still be walked single file.

The Transformer's move was to delete that dependency. Instead of passing a summary forward, every word looks at every other word directly, all at once, and the model works out how much attention each one deserves. Nothing waits for anything. A sequence that took a thousand sequential steps to read now takes one very wide step, and a very wide step is precisely what a GPU is for.

That is the whole reason the last decade happened. Not that attention is a cleverer way to read — it is that attention is a *parallel* way to read, and parallelism is the only thing that converts money into capability. Every scaling law, every hundred-million-dollar training run, every argument about compute budgets sits downstream of a decision to make the arithmetic wide instead of deep.

The bill came due at the other end. If every word must see every other word, then the model has to hold something about every word it has already read — and that state grows with the conversation, without limit, forever. In 2017 nobody minded: sequences were a few hundred words and the state was a rounding error. It is not a rounding error now, and almost everything else in this atlas is a response to it.

So the honest summary is not "the Transformer made models better." It is that the Transformer traded a constraint nobody could pay for a constraint everybody could — and then spent ten years discovering how expensive the second one turned out to be.

What it traded

gave up
the ability to process a sequence one step at a time, cheaply
got
the ability to train on the whole sequence at once

What exists now that didn’t before

Every word now has to see every other word — so remembering a conversation means holding a slice of it in memory for every word already said, and that pile only grows.

Asked, and answered

What did a paper about translating sentences have to do with chatgpt?

Everything — it's the same architecture. The 2017 paper was about making translation faster to train, not chatbots; the trick that did it (read the whole sentence at once instead of one word at a time) turned out to be what made every model since possible.

Read next