Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew Why
Big Tech turned AI usage into a metric and a 124-year-old bounty scheme in colonial Hanoi explains exactly what happened next.
The Tailless Rats of Hanoi
In 1902 the French administration in Hanoi had a rat problem, and what looked like a tidy solution. Pay a bounty for every rat killed. To save officials from handling thousands of carcasses, just collect the tails as proof. The bounties went out. Tails came in by the sackful. And then people started spotting rats around the city with no tails.
The locals had read the incentive more carefully than the people who wrote it. The government wasn’t paying for dead rats. It was paying for tails. So you caught a rat, cut off the tail, and released it to go make more rats with tails. A few enterprising types skipped the sewers and just bred rats on the edge of town.
I’ve been thinking about those tailless rats, reading the news out of Uber, Amazon, and Microsoft.

The metric ate the strategy
Here’s what happened in 2026, with the press releases removed. The big tech companies decided AI adoption was the future, which is fine. But “are we an AI-first company” is almost impossible to measure, so they grabbed something they could count instead. How many tokens are people burning. What share of engineers touch the tool each week. Then they put it on leaderboards and wired it into performance reviews.
You already know the rest, and so did the engineers. Amazon staff started calling it tokenmaxxing: pointing autonomous agents at make-work just to climb the internal board. Reddit threads and internal forums started filling up with advice on how to maximize token usage without setting off alarms: point autonomous agents at make-work, generate endless reports, route trivial tasks through the most expensive models. Uber ranked its engineers on a leaderboard by how much they used the tool, then blew its entire full-year AI budget in four months. When the company’s COO went looking for what all that spending had bought, he couldn’t find it. The link between the tokens and the value, he admitted, “is not there yet.”
The tails were coming in by the sackful. The rats were fine.
A measure becomes a target
The economist Charles Goodhart gets his name on this, though the clean version came later: once a measure becomes a target, it stops being a good measure. The idea underneath is almost too simple to respect. A metric is a stand-in. It points at something you care about but can’t watch directly. Reward people for the stand-in and you’ve told them, without saying it, that the stand-in is the actual job. People are extraordinarily good at doing the actual job.
What makes this hard to catch is that it never feels like a blunder while you’re committing it. It feels like rigor.

We have done it constantly. The Soviets ran nail factories on output quotas and got mountains of tiny useless nails; they switched to weight and got a few comically enormous ones. During Vietnam, Robert McNamara, the original numbers man, made enemy body count the measure of progress, and his commanders duly inflated kills and filed civilians as combatants while the war went sideways underneath the rising graph. Before 2008, banks measured danger with Value-at-Risk, which told you the worst you’d lose on 99 days out of 100; traders built books that were spotless on those 99 days and ruinous on the hundredth, the one the model couldn’t see.
My favorite is the NHS, because it’s so physical you can picture it. Hospitals were told no patient could wait more than four hours on a trolley before admission. Administrators unscrewed the wheels and reclassified the trolleys as beds. The clock stopped. The target turned green. The patient was still in the corridor, now lying on a piece of furniture that had been promoted.
Every one of these was run by clever people. That’s the part I’d ask you to sit with for a second.
Why the clever people miss it
The lazy read is that the executives were fools, and it’s wrong. These are among the best-credentialed problem-solvers on the planet. The failure isn’t IQ. It’s a specific trap, and it has a name.
Psychologists call it surrogation. Asked to judge something genuinely hard is the strategy working? the mind quietly swaps in an easier question it can actually answer: is the number going up? You don’t feel the swap happen. That’s what makes it dangerous. The dashboard stops being a window onto the strategy and becomes the strategy. The map eats the territory.
The historian Jerry Muller put a sharper edge on this in The Tyranny of Metrics. His claim is that organizations don’t reach for metrics despite their flaws but because of one specific feature: a number lets leadership skip the work of judgment. It looks objective. It fits in a board deck. It spares you from walking down to where the work happens and forming an opinion about whether it’s any good. Measurement quietly becomes a replacement for understanding.
And then consulting industrializes the whole thing. In late 2023 almost too perfectly McKinsey published a piece called “Yes, you can measure software developer productivity,” and a good chunk of the engineering world, Kent Beck included, set it on fire. The complaint was clean: the proposed metrics counted activity, not value. Lines of code, deployments, story points. Pay an engineer for lines of code and you get more lines of code. You get the exact bloat you were trying to kill. That advice didn’t stay in the PDF. Two years later it was the unspoken philosophy behind every mandate to push AI usage up and to the right, with nobody checking whether the usage made anything.
Why this one is actually worse
I don’t want to land on the comfortable note same old mistake, nothing new under the sun because mechanically there is something new here, and it’s the part that should bother you.
Each of those needed human labor to game. The factory still had to forge the giant nail. McNamara’s officers still had to go run the pointless mission. Someone still had to walk over and unscrew the trolley wheels. Faking output took work, and work is a speed limit. There’s only so fast a person can waste a budget by hand.
The agent has no speed limit.
It doesn’t run one query and stop. It loops. Queries a database, reads what came back, critiques itself, calls another tool, patches its own errors, goes again. Aim one at a meaningless task to pad your usage stats and you’ve kicked off an automated loop that can spend thousands of dollars while you’re getting coffee. Output tokens cost a multiple of input tokens, so every needless report it spits out is more expensive coming out than going in. For the employee, the cost of faking productivity has fallen to nothing. For the company it’s gone vertical.
There’s a tempo here the older cases never had. Each of those pathologies ripened slowly Soviet industry rotted over generations, the leverage that broke the banks built across years and even then the reckoning arrived only at human pace. What cost the Soviets generations, Uber ran in four months. Take away the human speed limit and you don’t just waste money faster; you watch a metric’s whole pathology run in fast-forward, months where it used to take careers.
A telemetry firm called Entelligence put a number on the damage that I haven’t been able to shake. Across thousands of companies, they reckoned that of every dollar spent on AI tokens, about eighteen cents turned into something a user could actually use. The rest went to fixing the bugs the AI wrote, reworking its code, and clearing the review pileup it caused. We built the most efficient waste machine ever made and then gave it a leaderboard.
There’s a stranger layer underneath all this. When gaming a metric was expensive, the cost at least stayed inside the company that set the metric the giant nail sat on the factory floor, useless, but nobody outside got rich off it. Token waste doesn’t sit on a floor. It leaves the building as a payment. Every pointless agent loop, every make-work report run to nudge a usage number upward, is somebody’s revenue. Which puts the AI vendors closer to Hanoi’s rat farmers than to the government that set the bounty: not the party defrauded by the system, but the party that quietly understood the system paid for tails and built a business on supplying them. The buyer calls it productivity. The seller books it as growth. The fastest-growing revenue line in the industry may turn out to be, in part, a lot of companies paying to learn the wrong lesson from their own dashboards.
What to do, and what it’s really about
The fix isn’t a cheaper model or a faster chip. It’s the thing all of these stories have been waving at for a hundred years: pay for the dead rat, not the tail. Measure the outcome instead of the exhaust cost per resolved ticket rather than tokens spent, product shipped rather than pull requests opened. Route the trivial work to cheap models automatically, so the call isn’t left to a nervous engineer reaching for the most powerful thing on the shelf in case someone’s watching. And, as Muller would say, let people use judgment again, since that’s the thing the number was always a bad copy of.
But the operational fixes aren’t the lesson that stays with me. The lesson is that the urge to swap judgment for a number is bottomless, and it always arrives wearing the costume of discipline. The Soviet planner thought tonnage was discipline. McNamara thought the enemy body count was discipline. The bank thought VaR was discipline. The executive watching the AI usage dial climb thinks the dial is discipline.
It never is. A metric is a finger pointing at something worth looking at, and most of management is the long history of people admiring the finger.