Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 1 Aug 2026

AI Has Passed Every Exam. It Has Never Had an Idea.

The machine has passed every exam we can build and discovered nothing.

First published on Substack, 1 Aug 2026.

The machine has passed every exam we can build and discovered nothing. Both are true. What dissolves the contradiction is the one thing no test has ever measured in machines, or in us.

Here are two facts about AI in 2026 that you have almost certainly seen, but never had to hold in the same hand.

The first: the machine passes everything. Humanity’s Last Exam, a wall of graduate-level questions built specifically to be too hard, it clears. Gold at the International Math Olympiad. The bar exam, the medical boards, competition-grade code. We keep building harder tests, grandly named, designed by PhDs to find the ceiling, and the machine keeps stepping over them. On any exam a human can be given, it is now roughly the best test-taker alive. Frontier models gained 30 percentage points in a single year on Humanity’s Last Exam alone.

The second is harder to state cleanly than it was even a year ago. AI systems now genuinely contribute to scientific discoveries. AlphaFold, which won the 2024 Nobel Prize in Chemistry, GNoME, which predicted 2.2 million crystal structures of which outside labs have physically synthesized 736, FunSearch and others have shown that much. But notice the shape of those discoveries. In every case, a human decided the question was worth asking. The machine answered it, often spectacularly. I could not find an example where it independently decided which question mattered. No AI has posed a question nobody thought to ask, found an anomaly nobody was looking at, or made the kind of leap that turns a field inside out. It answers our questions superbly. It has never, on its own, decided which question was worth asking.

Both true. Sit in the discomfort for a second, because the discomfort is the whole essay. The best test-taker in history, and it has never had an idea. If intelligence were one thing, that shouldn’t be possible.

So what mechanism would reliably produce this outcome, over and over, regardless of how good the underlying system gets?

The easy dissolves don’t work

Everyone reaches for a comfortable way out of that contradiction, and both exits are blocked.

The first exit: it’s not really thinking, it’s just autocomplete. That feels good and explains nothing, because “just autocomplete” doesn’t clear Humanity’s Last Exam. Whatever it’s doing, dismissing it doesn’t survive the scoreboard.

The second exit: give it time and scale, discovery is coming. Maybe. But that’s a promise, not an explanation, and it dodges the actual question which is why the gap has this particular shape. Why superhuman on every answer and silent on every question? Scale doesn’t explain a structural asymmetry. It just promises to wash it away later.

Around halfway through writing this essay, Tom Zahavy at Google DeepMind published a position paper called LLMs Can’t Jump. For about ten minutes I thought I’d been beaten to the idea. Then I realized we were asking different questions. His argument runs through Peirce: models handle induction and deduction, but not the abductive leap that invents an explanation the data never contained. He has since clarified that this is a personal position rather than DeepMind’s, and that scaling might prove him wrong.

Zahavy asks whether the machine can make the leap. I found myself asking something slightly stranger: suppose it did. How would our benchmarks know?

So there’s a hidden variable, the way there always is when two careful observations point in opposite directions. And the variable isn’t in the machine. It’s in the tests.

A benchmark is a question someone already chose

Here is the thing that dissolves the paradox, and once you see it you can’t unsee it.

Every benchmark, every exam, every one of those grandly-named tests, has the same structure: someone writes the question, and the machine finds the answer. The question is given. It arrives pre-selected, pre-formatted, flagged as important, with a known answer sitting in a locked drawer so the thing can be graded.

And that means every test we have ever built measures exactly one half of intelligence, the finding of answers and structurally cannot measure the other half: the choosing of the question. A test can’t measure question-choosing, because a test is a chosen question. The container can’t weigh the thing that decides what goes in the container.

That sounds like wordplay until you look at where the great leaps actually came from, and notice that the choosing was always the hard part.

Einstein didn’t answer the question. He found it.

In 1887, two physicists named Michelson and Morley ran an experiment expecting to measure how the Earth’s motion changed the speed of light. They got nothing. Light moved at the same speed no matter which way you chased it. A clean, baffling null result and they published it.

Then it sat there. For eighteen years, in the open, in the literature, available to every physicist alive. The anomaly wasn’t hidden. The data was on the shelf.

In 1905 a patent clerk picked it up and did the thing nobody else had done. He didn’t find a cleverer answer to the question everyone was asking how does the Earth’s motion affect light? He decided that question was wrong. He found the existing picture of absolute time and space intolerable in a way his colleagues, staring at the same data, did not. The leap wasn’t the mathematics of special relativity; the math is undergraduate now. The leap was upstream of the math, the decision that simultaneity itself was the thing to doubt. Everyone had the answer-shaped problem. Einstein found the question-shaped one.

The physicist David Deutsch puts the general version cleanly: the people who changed science didn’t extrapolate from what was known. They stepped outside the space of existing explanations and proposed one that wasn’t logically waiting there to be derived. Darwin, Turing, Einstein none of them was doing better test-taking. They were deciding, against the grain of everyone around them, which question deserved a life.

The test we built for genius, and what we had to hand it

Now the part that should make you put the coffee down.

Researchers at DeepMind have proposed what they call the Einstein test the sharpest benchmark yet for real machine genius. The design: feed an AI everything known before the breakthrough, and see if it can independently produce the leap. Give it the physics of the era and see if relativity falls out.

It’s a beautiful idea. And look at what we had to do to build it.

We curated the dataset around the discovery. We pre-selected the era, the field, the anomaly, and pointed the machine straight at it: here is everything before 1905, now derive relativity. We built a test for the one man whose genius was deciding what to look at and to make it gradeable, we did the looking for it. I don’t think this is a flaw in the benchmark. I think it’s unavoidable. The moment you need a score, someone has to decide what counts as the question.

The Einstein test skips the part that made Einstein. It measures whether the machine can answer the question once a human has already found it. Which is the same thing every other benchmark measures, wearing a lab coat. Even our test for genius is, underneath, an answer key, because an answer key is the only kind of test we know how to write.

The measurement cannot contain the thing it’s trying to measure.

Why every test we build has this hole

There’s a deeper reason our benchmarks all share this blind spot, and it’s the quiet assumption running the entire debate.

We have graded answer-finding for three thousand years. Exams, vivas, boards, olympiads, we are extremely good at scoring whether someone can produce the right answer to a set question, because that’s the part of the mind we long ago learned to make legible. We have never once built a reliable test for curiosity, for the decision that a boring answer is intolerable, that a settled question isn’t settled, that this problem and not that one deserves a decade. We can’t score it because we’ve never figured out how to make it legible, even in each other.

The closest anyone has come is instructive MEDIQ strips patient information out of medical question-answering so a model has to decide at each turn whether it knows enough or needs to ask. Prompting models to ask questions dropped diagnostic accuracy by 11.3 percent. CLAMBER found models fail at clarifying questions because they cannot assess the boundaries of their own knowledge. Curiosity isn’t just unmeasured. It’s penalized.

So our benchmarks inherit our own blindness. They measure the half of the mind we can see and are silent on the half we can’t and then we’re startled when the machine sails past us, without noticing that it’s sailing past us on the only stretch we ever learned to measure. The machine looks like it’s overtaking human intelligence. It’s overtaking the articulable part of human intelligence, which was always the smaller, cheaper part.

What struck me wasn’t the benchmark itself. The same optimization pressure reappears after training. In deployment, reliability is the product, so inference is made as deterministic as possible, even at significant computational cost: Thinking Machines Lab traced residual nondeterminism at temperature zero to GPU batching and fixed it at roughly 60 percent slower inference. In academia unstable creativity benchmark are treated as a problem to be engineered away with tighter rubrics. Different incentives, same direction. That’s probably an essay of its own.

Three independent systems end up selecting for the same thing. Benchmarks determine what gets measured. Training determines what gets optimized. Products determine what survives deployment. None of them explicitly sets out to eliminate question-finding. Together, they may.

Where I’ll stop short

Let me refuse the too-easy version, because it’s the mirror of the hype and I don’t believe it.

I am not telling you a machine can never ask a real question. Two honest problems with that claim. First, the Einstein test is leakier than it looks the anomaly is sitting in the pre-1905 data, so a sufficiently good search might surface relativity without any genuine leap, which means even this test doesn’t cleanly separate asking from answering. Second, machines already produce things nobody designed: novel protein structures, game moves no human would play, research systems that propose their own next experiments. Those are real. But so far they are a straight-A student’s answer to a question the rules already posed extraordinary answer-finding inside a space we defined, not the leap of deciding the space was wrong.

So the honest claim isn’t “the machine can’t.” It’s stranger and more unsettling than that: we have no instrument that would tell us if it could. We have never built a test that measures wanting-to-ask, because a test is a thing you’re handed. So on the single capability that would actually mark the arrival of a new kind of mind, we are flying completely blind and we’ve mistaken our own blindness for the machine’s limit.

The half we never learned to grade

There’s a tool hiding in all of this, and you can run it on your own life before you finish the page.

Ask of any test you’re inside a benchmark, a KPI, an exam, a performance review: did someone hand me this question, or did I decide it was the question? The first is answer-work, valuable, gradeable, and exactly the kind of work now being handed to machines. The second is question-work, and it has always been the scarce thing, in science, in a career, in a life. We just never learned to measure it, so we quietly built a civilization that rewards the half we could score.

We are now doing it again, at the scale of datacenters. We will keep building harder exams, and the machine will keep clearing them, and every time it does, someone will announce that genius has arrived. It hasn’t. It won’t show up on the scoreboard, because the scoreboard was only ever built to grade answers, and the thing we’re waiting for was never an answer.

We built a test for Einstein and handed it Einstein’s question. The genius was never in the derivation. It was in deciding the question was worth a life and that is the one exam we have never known how to write, for a machine or for ourselves.