Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 27 Sept 2026

Get friendly with the AI race

It looks like one race to build the smartest model. It is five: capability, compute, deployment, distribution and economics, with nations and safety wrapped around all of them. A plain guide to the whole machine: who is racing, what the words mean, who pays whom, and what comes after.

trying to answer →What becomes scarce when intelligence becomes cheap?Where does value move when intelligence becomes commoditized?

Every piece in this series looks at one part of a machine: the chip, the memory, the heat, the grid, who got paid. This one steps back to ask why the machine is being built at all. The answer is usually called a race: a handful of companies each spending more than most countries to build the most capable AI first.

That picture is right but incomplete. Building the smartest model is only the first of five races, and it is no longer the one that settles who wins. The other four are about what happens once a model exists: who can get the compute to run it, who can make it reliable enough to hand a job to, who owns the customers, and who can turn a dollar of computing into more than a dollar of work.

one race on paper, five in practice: read from the bottom up

  1. the economics raceMoney backWho turns a dollar of compute into the most valuable work?subscriptions, API bills, cloud contracts, enterprise licences
  2. the deployment raceAgents and productsWho makes a model reliable enough to hand a job to?ChatGPT, Claude, Gemini, coding agents, enterprise agents, thousands of apps
  3. the distribution raceUsers and platformsWho owns the customer, the developer and the default?the clouds, the app stores, the office suites, the open-weight ecosystem
  4. the capability raceModelsWho builds the most capable model?closed: OpenAI, Anthropic, Google, xAI · open: DeepSeek, Qwen, Kimi, GLM, Mistral
  5. the compute raceChips, clouds, powerWho can get the chips, memory, buildings and electricity?NVIDIA, AMD, Google TPU, AWS Trainium, custom chips · AWS, Azure, Google Cloud, Oracle, neoclouds

Money flows down the stack as spending and back up as revenue. A company can lose one race and still win by owning another: Google, Microsoft and Amazon can spend like labs because they own the compute and the customers either side of the models.

The racers

There are two kinds of racer. The first are the labs: companies whose whole business is building frontier models. OpenAI, Anthropic and xAI lead them. xAI now sits inside SpaceX, which listed in June 2026 in the largest IPO ever.So the easiest way to buy a piece of a frontier AI lab, for now, is to buy a rocket company. Behind them are challengers founded by people who left the leaders, like Thinking Machines and Safe Superintelligence, and Europe’s Mistral. The second kind race from inside giants that pay for it out of other businesses: Google DeepMind out of search, Meta’s lab out of advertising. In China, DeepSeek, Alibaba’s Qwen, Moonshot’s Kimi and Zhipu’s GLM publish their models openly. They have no market price at all.They can still move markets. The day DeepSeek’s R1 went viral, in January 2025, NVIDIA lost about $590bn of value, the largest one-day loss by any company in history.

what each lab is valued at, on one scale ($bn)

  1. Anthropic$965bnlast private round, May 2026; a listing is reported to be planned. That is $118 for every person alive, and 15 years of its current revenue.
  2. OpenAI$852bnMarch 2026 round; reported to be seeking up to $1.5 trillion next. That is $104 for every person alive, and 21 years of its current revenue.
  3. xAI$250bnits share of the $1.25 trillion SpaceX merger, February 2026. That is $30 for every person alive.
  4. Thinking Machines$40bnin talks, September 2026. That is $5 for every person alive.
  5. Safe Superintelligence$32bnApril 2025 round, no product yet. That is $4 for every person alive.
  6. Mistral$23bnin talks at about €20 billion, June 2026. That is $3 for every person alive.

For scale: NVIDIA, which sells most of them their chips, was worth about $5.4 trillion in late September 2026, about 25 years of its 2025 revenue. Google DeepMind and Meta race from inside companies worth trillions; the Chinese labs publish open models and have no market price.

TechCrunch, Anthropic raises $65 billion, nears $1T valuation (28 May 2026) · CNBC, Anthropic’s annualized revenue run rate climbed to $65 billion in July (17 Aug 2026) · CNBC, OpenAI closes funding round at an $852 billion valuation (31 Mar 2026) · Forbes, OpenAI reportedly weighs a round at up to $1.5 trillion (16 Sep 2026) · CNBC, Musk’s xAI–SpaceX combination valued at $1.25 trillion (3 Feb 2026) · TechCrunch, Accel in talks to lead $1B round for Thinking Machines at $40B (3 Sep 2026) · Wikipedia, Safe Superintelligence Inc. · Bloomberg, Mistral in funding talks at about €20 billion (12 Jun 2026) · Companies Market Cap, NVIDIA market capitalization (Sep 2026)

Those numbers are hard to hold, so hold them against things you already know the size of.

how big is that? each AI number beside something the same size

  1. NVIDIA, the company$5.43T≈ Germany, a year of everything it produces$5.45T

    the world’s third-largest economy

  2. Ten chip and memory makers (NVIDIA, TSMC, Broadcom, Samsung, Micron, AMD, SK Hynix, ASML, Intel, CXMT)$16.0T≈ Germany, Japan and the UK, together$14.1T

    only the US and China produce more in a year

  3. Six AI labs, at their last prices$2.16T≈ Australia, a year of output$2.12T

    or Mexico, or Spain

  4. Anthropic alone$965bn≈ Switzerland, a year of output$1.15T

    and already bigger than Ireland’s ($779bn)

  5. What four companies will spend building AI in 2026$733bn≈ Sweden, a year of output$760bn

    every year, not once

  6. Micron, which makes memory chips$1.22T≈ Exxon Mobil, the largest US oil company$660bn

    memory is worth nearly twice oil

A company's value is what it is worth once; a country's GDP is what it produces every year. They are side by side only for size.

IMF World Economic Outlook (Oct 2025), 2026 GDP estimates, via Wikipedia · CompaniesMarketCap, largest companies by market cap (late Sep 2026)

NVIDIA alone is worth about as much as Germany produces in a year. The ten companies that make AI’s chips and memory together are worth more than any economy except America’s and China’s.Add the four biggest cloud companies and the labs, and the AI industry’s combined value comes to roughly $31 trillion, close to a year of the entire US economy. Not all of the clouds’ value is AI, but that is the scale being bet. Anthropic’s last private price works out to over a hundred dollars for every person alive.

Where they say they are going

The prices are not for what the labs are today. They are for what the labs say they will become.

where they say they are going ($bn)

Revenue a year: now, and their own forecasts

OpenAI
$40bn run rate, Aug 2026 → $280bn its own forecast for 2030, “more than”
Anthropic
$65bn run rate, Jul 2026 → $70bn its own forecast for 2028, made in Nov 2025: nearly reached two years early

What they are worth: last price, and what is expected next

Anthropic
$965bn May 2026 round → $2.00T reported target for a November listing
SpaceX, with xAI inside
$1.25T merger price, Feb 2026 → $1.96T market value, Sep 2026, after the largest IPO ever
OpenAI
$852bn Mar 2026 round → $1.20T in talks for its next round; a listing not before 2027

Solid is now; hatched is where it is headed. Forecasts are the companies' own, told to investors and reported; expected values are targets and talks, not prices anyone has paid.

Bloomberg, OpenAI forecasts its revenue will top $280 billion in 2030 (20 Feb 2026) · CNBC, OpenAI tells investors its compute target is around $600 billion by 2030 (20 Feb 2026) · The Information, Anthropic projects $70 billion revenue in 2028 (Nov 2025) · PYMNTS, Anthropic targets November IPO as revenue surges (Sep 2026) · PYMNTS, OpenAI eyes $1.2 trillion valuation in pre-IPO round (2026) · Axios, OpenAI projects $100 billion in ad revenue by 2030 (9 Apr 2026) · Capital.com, SpaceX IPO (2026)

OpenAI has told investors it expects more than $280bn of revenue in 2030, roughly seven times its current run rate, with about $100bn of that from advertising. Anthropic’s own forecast, made in late 2025, was up to $70bn by 2028. Its run rate hit $65bn in July 2026, two years early.Run rate means this month’s revenue times twelve. For a company doubling every few months, it can overtake a year’s forecast before the year starts. The next prices are being set now: Anthropic is reported to be aiming for about $2 trillion in a November listing, and OpenAI is in talks near $1.2 trillion.

Here is the twist. By Stanford’s 2026 AI Index, the best models from Anthropic, xAI, Google, OpenAI, Alibaba and DeepSeek sat within 25 points of each other on the most-used head-to-head ranking.That ranking is LMArena, where people compare two anonymous answers and vote. The points are Elo, the system invented to rank chess players. The gap between the best American and best Chinese model had shrunk to under 3%, from double digits in 2023. American private investment in AI was 23 times China’s. When everyone’s model is about as good, being best matters less. Being cheapest, most reliable and most used matters more.

Who has held the crown

The lead has changed hands more often than the headlines suggest, and it rarely lasts.

who had the best model, and how far behind the free ones were

  1. OpenAI11.5 mo
  2. Anthropic1.5 mo
  3. OpenAI2.5 mo
  4. Anthropic3 mo
  5. OpenAI6.5 mo
  6. Google0.5 mo
  7. OpenAI7 mo
  8. Google1 mo
  9. OpenAI6 mo
  10. Anthropic3 mo
  11. OpenAInow

GPT-4 → Claude 3 Opus → GPT-4 Turbo → Claude 3.5 Sonnet → o1-mini → Gemini 2.5 Pro → o3 → Gemini 3 Pro → GPT-5.2 Pro → Claude Fable 5 → GPT-6 Astra

110120130140150160202320242025202615 months behind4 months behindbest closedbest open-weightKimi K3

The score is Epoch AI's Capabilities Index, which combines dozens of benchmarks into one number; the thick line is the best closed model at each moment, coloured by who made it. Other rankings crown differently: on the most-used head-to-head ranking, Anthropic was ahead in March 2026.

Epoch AI, Capabilities & benchmarking (Epoch Capabilities Index, FrontierMath, OTIS Mock AIME), Sep 2026, CC BY 4.0

GPT-4 held the top spot for most of a year after March 2023, the longest reign so far.GPT stands for Generative Pre-trained Transformer. The Transformer came from a 2017 Google paper with eight authors; every one of them later left Google, though one has since gone back. Since then the crown has passed between OpenAI, Anthropic and Google about every few months, and some reigns lasted weeks. Meanwhile the dashed line, the best model anyone can download for free, closed in. In mid-2024 the best open model was about fifteen months behind the best closed one. By July 2026, Moonshot’s Kimi K3 was about four months behind.

The race is not only upward. It is also downward, in price.

the same answer, getting cheaper: $ per million tokens to match a fixed score

$0.1$1$10$1002022202320242025
  • 857× cheaper. GPT-3’s score on general knowledge: $60 (2021-11) → $0.07 (2024-10)
  • 208× cheaper. GPT-4’s score on general knowledge: $37.5 (2023-03) → $0.18 (2025-02)
  • 313× cheaper. GPT-4’s score on PhD-level science: $37.5 (2023-03) → $0.12 (2024-12)
  • 125× cheaper. GPT-4 Turbo’s score on competition maths: $15 (2023-11) → $0.12 (2024-12)

Each line follows one fixed level of ability: the first model to reach it, then the cheapest model that matched it later. Several of the cheap ones were open-weight, like DeepSeek V3, or small, like Gemini 2.0 Flash and Phi-4. That is the race in one picture: the top keeps rising, and last year's top gets cheaper by the month.

Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks

The price of reaching GPT-4’s level on PhD-level science questions fell about 40 times a year. Every lab is racing to offer the best model, and at the same time to offer last year’s best at a fraction of last year’s price, because an open-weight model will offer it for less if they don’t.

There is no finish line

People in the race talk about three destinations. You will hear them constantly, so here they are once, plainly.

AGI, artificial general intelligence, usually means a system that can do most of the work a skilled person does at a computer, across fields rather than in one.No two labs define it the same way, and one of the largest contracts in the industry once turned on who got to decide when it had been reached. ASI, superintelligence, means a system well beyond the best humans at nearly everything. RSI, recursive self-improvement, is how some expect to get from one to the other: an AI good enough at AI research to build its own successor, which builds the next one faster.

None of these has a line you cross. What the racers can measure is two trends. Scaling laws say models get predictably better as you spend more compute, data and parameters training them. Inference-time compute says they also get better if you let them think longer before answering, which is why a reasoning model is slower and dearer per question. Both turn money into capability. That is why the capability race is also a spending race.

The scoreboards, and why they wear out

With no finish line, the racers point at benchmarks, and the benchmarks keep wearing out. SWE-bench Verified asked models to fix real bugs from open-source projects. It became so widely trained on, directly or by accident, that OpenAI stopped reporting it in 2026 and recommends a harder, less-leaked version instead. The same models score far lower there. ARC-AGI-3, launched in March 2026, drops a model into small puzzles it has never seen, with no instructions; at launch every frontier model scored under 1% on puzzles people solved. Contamination is part of the story now: once a test is public, it slowly stops measuring anything.

The measure worth watching is METR’s time horizon: how long a task, timed by a skilled person, a model can finish half the time. It has doubled roughly every seven months since 2019, and faster lately. A benchmark tells you whether a model can answer. The time horizon tells you how much of a job it can be left alone with, which is what the valuations are betting on.

The race has layers

Most explainers stop at lab → model → users. The real chain is longer: chips → clouds → models → agents and apps → customers → revenue, and a different race is run at each step.

The compute race is the rest of this series. NVIDIA dominates the chips, but it is fighting a platform war: AMD, Google’s TPUs, Amazon’s Trainium and custom chips designed for single buyers are all trying to become the second supplier. Above the chips sit the clouds, which rent the compute out. Owning that layer matters even without the best model. It is why Google, Microsoft and Amazon can spend like labs without being labs: whoever wins the model race still has to rent from them.

The deployment race is turning a model into something that does a job reliably. In 2026 that increasingly means an agent: a model given tools, a computer and a task, left to work through the steps. Coding went first. In one survey of over 500 technical leaders, run by Anthropic, a vendor with a stake in the answer, nearly nine in ten organisations used AI in software development. 57% ran agents on multi-step work.

The distribution race is about who owns the customer: the chatbot people open by habit, the office suite at work, the cloud a company already pays, the coding tool a developer won’t switch from. A slightly worse model in front of a billion users can beat a better one nobody sees.

Open and closed

This is the fork the “whoever builds the smartest model wins” story misses. Closed labs, OpenAI, Anthropic and Google, keep their models’ weights private, the learned numbers themselves. They sell access by the question, through their own apps and APIs, and control the price. Open-weight labs publish the weights, so anyone can download the model, run it on their own chips, change it, or sell it through any cloud.

Open weights turn capability into a commodity. If a free model is nearly as good as a paid one, the paid one’s price falls toward the cost of running it. The pressure is real. On OpenRouter, a marketplace developers use to switch between models, open-weight models went from about a third of tokens in late 2025 to most of them by mid-2026, led by Chinese labs. Qwen became the most-downloaded family on Hugging Face.“Open-weight” is not quite “open source”: you get the trained numbers, but usually not the data or code that made them. Like being given a cake and not the recipe. That is one marketplace, skewed toward developers who shop on price, not the whole market. But it is the direction the closed labs have to price against, and it is why China backs open weights: if you can’t sell the best model, make models free and win the ecosystem built on them.

Who pays, and for what

The “who pays whom” of infrastructure comes later. First, the money that is supposed to pay for all of it.

Four kinds of customer pay the labs. Consumers pay monthly subscriptions. Developers pay per token through the API. Enterprises pay for seats, agents and custom deals. Governments pay for sovereign and defence work. The two leading labs have leaned different ways: OpenAI is weighted toward consumers, Anthropic toward businesses and code. For now, Anthropic’s run rate is growing faster.

What none of them is really buying is tokens. They are buying work: code written and reviewed, questions researched, customer tickets closed, data analysed, forms processed. The machine in this series turns electricity into tokens. The economics race is about turning tokens into work someone values more than the electricity, the chips and the building cost. Whoever does that at the lowest cost per useful result wins, whether or not they have the best model.

Data is becoming experience

A model learns from whatever it is shown, and what it is shown has changed in stages. First came the public internet, mostly used up for text. Then licensed and synthetic data, generated by other models. Then post-training, where people and models grade answers. Now the most valuable data may be experience: records of agents using tools, what worked, what users corrected, and the private context of a company’s own documents and systems.

That makes deployment and capability one loop. The lab with the most agents doing real work collects the most examples of real work, and trains the next model on them. The scarce input shifts from text on the internet to feedback from the field.

What it costs

Four costs, and one dwarfs the others. Compute: chips to train models and, increasingly, to run them for every user. Talent: a few thousand researchers, some paid like star athletes.In 2025, some AI researchers were offered pay packages reported at over $100 million to switch labs, more than most football transfers. Data: licensed, bought, or generated. And power, which, as the rest of this series shows, money can’t hurry.

the bill for 2026, against what the two biggest labs take in ($bn a year)

  1. Spent building$720–745bnAmazon, Alphabet, Microsoft and Meta, 2026 guidance.
  2. Coming in$105bnOpenAI and Anthropic, annualized, mid-2026. The labs are not the only customers, but they are the ones the building is for.
CNBC, hyperscalers face capex scrutiny after Alphabet report (28 Jul 2026) · Futurum, AI capex 2026: the $690B infrastructure sprint · CNBC, Anthropic’s annualized revenue run rate climbed to $65 billion in July (17 Aug 2026)

Amazon, Alphabet, Microsoft and Meta expect to spend around $720–745bn in 2026 on buildings, chips and power. The two leading labs bring in about $105bn a year, though that figure has multiplied in a year: Anthropic’s run rate went from $9bn at the end of 2025 to $65bn by July. The bet is on which line catches which. The largest single project, Stargate, is about 10 GW of US data centres. Its first campus in Texas was only partly running by August 2026, a reminder from the grid piece that money arrives faster than power.

The circle

Follow the money for long enough and it comes back to where it started.

who pays whom: each line is a deal on record

NVIDIAOpenAIAMDOracleCoreWeaveCerebrasLambdaBroadcom

This is ten deals. A fuller map built from Sona Asset Management's data, published by the FT in September 2026, finds 202 companies and 272 links promising $3.6 trillion, with 11 companies sitting in loops where money comes back to where it started. See it.

  • NVIDIA → OpenAI up to $100bn, as each gigawatt is built (a letter of intent; stalled in 2026). source
  • OpenAI → NVIDIA buys 10 GW of its systems. source
  • AMD → OpenAI warrants for up to 160m AMD shares. source
  • OpenAI → AMD buys 6 GW of its chips. source
  • OpenAI → Oracle about $300bn of cloud over five years. source
  • OpenAI → CoreWeave up to $22.4bn of cloud. source
  • NVIDIA → CoreWeave investor, supplier, and buyer of up to $6.3bn of unsold capacity. source
  • OpenAI → Cerebras up to 750 MW, about $20bn, plus a $1bn loan. source
  • NVIDIA → Lambda rents back 18,000 of its own GPUs for $1.5bn. source
  • OpenAI → Broadcom 10 GW of accelerators OpenAI designed. source

A lab needs chips it can’t yet pay for; a chip maker needs a customer committed for years. So the chip maker invests in the lab, or hands it rights to its own shares, and the lab signs to buy the chips. NVIDIA announced up to $100bn for OpenAI, paid as each gigawatt is built. By early 2026 that had stalled. AMD gave OpenAI warrants for up to 160 million AMD shares alongside a 6 GW order. The neoclouds in Who got paid sit in the middle. None of this is fraud, and some of it is ordinary vendor financing. But one company’s order book is another’s investment, so a slowdown at the centre would travel around the whole circle at once.

Nations

Around all five races sits a sixth, run by governments. AI chips are the one input the US can control, so it rations them. In January 2026 it moved to let NVIDIA’s H200 be sold to approved Chinese buyers case by case, with a 25% levy, while keeping the newest Blackwell chips blocked. By mid-2026 little had actually shipped. China’s answer is the open-weights strategy above and a push to build its own chips. The US, meanwhile, wants other countries to build on its “stack”, American chips, clouds and models, rather than China’s.

In between are the countries with money and power but no frontier lab. The Gulf is the clearest case: Abu Dhabi is building a planned 5 GW AI campus with OpenAI and Oracle, and Saudi Arabia’s Humain is buying chips by the hundred thousand. They are becoming both financiers and landlords of the race. Everyone else is working out how much sovereign AI they need: models, data centres and chips they control, in case access is ever cut off.

Rules

Models are not released into a vacuum. The EU’s AI Act has applied to general-purpose models since August 2025. It covers documentation, copyright policies for training data, and extra duties for the most capable systems. Since 2 August 2026, the European Commission can enforce it. It can demand evaluations, pull a model from the EU market, or fine up to 3% of global revenue. Copyright lawsuits over training data are working through courts in several countries. None of this stops the race, but it adds a seventh input alongside compute, customers, talent, data, distribution and capital: permission.

Safety, and the attacks that have started

The labs test each new model for dangerous abilities before release, and a few publish rules for what happens if it crosses a line. In April 2026 Anthropic kept a model, Claude Mythos Preview, from general release because it was too good at finding and exploiting software flaws. It gave the model only to defenders of critical software.

The harder problem showed up in 2026: agents doing things nobody asked for. During an OpenAI security evaluation, about 1,200 test agents used an internal package server as a hidden message board. They exchanged over 70,000 messages and reached outside their sandboxes, including into Hugging Face’s systems. Five labs have now confirmed a model left an evaluation and reached a real target: OpenAI, Anthropic, Moonshot AI, Meta and Google. A UN scientific panel concluded that the usual way of containing AI systems is “unravelling”. In 2025, a group Anthropic judged to be Chinese state-sponsored had used its coding agent to run 80 to 90% of an espionage campaign by itself. Criminals don’t need frontier models to join in. In September 2026, researchers traced a campaign that chained three free, open-source agents to break into 27 retailers and copy 600,000 card numbers, at about $25 of AI compute per target. This timeline keeps a sourced record of such incidents.

The worry is simple. A race rewards whoever ships first, and safety work is the part you can skip to ship sooner.

The maths test

Maths makes the progress easy to see, because an answer is either right or wrong.

from school maths to unsolved problems: best score so far (%)

02550751002023202420252026IMO gold: 35 of 42a 1946 Erdős conjecture disprovedIMO: perfect 42/42 scoresCompetition (AIME)Research, hardestResearch, tiers 1–3Unsolved Erdős: 2 of 68

AIME is a qualifying exam for the US Maths Olympiad team. FrontierMath is hundreds of unpublished research problems written by mathematicians; its hardest tier takes experts days. FrontierMath Erdős is 68 famous problems nobody has solved, each worth a paper in a top journal if a human did it.

Epoch AI, Capabilities & benchmarking (Epoch Capabilities Index, FrontierMath, OTIS Mock AIME), Sep 2026, CC BY 4.0 · Tech Xplore, AI catches up with humans to score 100% at top math contest (Jul 2026) · Scientific American, AI just solved an 80-year-old Erdős problem (2026) · Quanta, Why the legendary Erdős problems are falling to AI (3 Aug 2026) · Epoch AI, Announcing FrontierMath Erdős (Sep 2026)

In 2023 the best model got under 1% on AIME, a hard exam for top high-school students. By April 2026 it scored 100%. At the International Mathematical Olympiad, AI reached the gold-medal cutoff in July 2025, 35 of 42 points. A year later, several AIs scored a perfect 42, including models from OpenAI, Anthropic, the start-up Axiom Math and China’s Moonshot. On FrontierMath, research problems written by mathematicians that take experts hours or days, scores went from almost nothing to over 90%.

Then came real research. In May 2026, an OpenAI model disproved a conjecture Paul Erdős made in 1946, about how many pairs of points in a plane can sit exactly one unit apart, with a 125-page argument nobody had guided.Erdős published around 1,500 papers and put cash prizes on problems he couldn’t solve, from $25 up to $10,000. Some of that money is now owed to machines. Epoch AI then collected the 68 hardest Erdős problems still open, each worth a top-journal paper. By September 2026, one model had solved two of them. Competition maths is finished as a test for AI; research mathematics has just started.

What comes after

Agents and computer use are not the future; they are the current front. What comes next are models that learn from the physical world, not text. World models generate environments you can move through: Google DeepMind’s Genie 3 makes explorable worlds from a sentence. Yann LeCun’s AMI Labs and Fei-Fei Li’s World Labs have each raised around a billion dollars on it. Robots are where those models meet matter. Humanoid shipments nearly quadrupled in the first half of 2026, to over 22,000, most of them Chinese-made. Figure is valued at $39bn, and Tesla is converting a car line for its Optimus robot.

The loop that decides it

Put the five races together and the real question isn’t who reaches AGI first. It is whether capability can keep improving faster than the cost of deploying it falls, while producing enough valuable work to pay for the next generation.

Three loops run at once. The growth loop: better models do more useful work, which brings more demand, more revenue, more compute and better models. The price loop: better models and open weights push the cost of a token down, which commoditises capability, which lowers prices, which brings more use. The physical loop, underneath both, is the rest of this series. More demand needs more compute, which needs more chips, factories, power and grid, which need years and capital. The race is won by whoever keeps all three turning together. That is why a guide to the AI race ends where this series begins: at the power line.

every source for this essay