Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 13 Jun 2026

The Wall That Wasn’t Yours

Fable 5 is Anthropic’s best model, but Friday’s letter showed why 'best' doesn’t mean valuable

First published on Substack, 13 Jun 2026.

On June 9, Anthropic’s fable announcement described a Stripe team that pointed Fable 5 at a fifty-million-line Ruby codebase and ran a full migration in a single day. The same task would have taken a whole team over two months by hand.

That is a vendor-sourced single case study, and it should be treated as one. But it describes the specific phenomenon worth paying attention to: a task that costs $250,000 in human time and $5,000 in AI time doesn’t just cut costs it unlocks projects that wouldn’t exist at $250,000. That is the uncapped demand version of this story. It is not the only version.

The demand problem is not what the bears think

The bear case on frontier AI spending has two versions. Version one: the productivity gains are fake, it’s all Goodhart’s Law, engineers gaming token leaderboards with make-work agents while COOs go looking for value and find none.

But that version describes a specific failure mode: metrics that reward usage rather than outcomes. It doesn’t describe Stripe running a genuine migration, or Anthropic’s run rate climbing from $9 billion a year ago to $47 billion, a trajectory Ben Carlson documented in June 2026 that doesn’t emerge from fake engagement.

The distinction that matters is whether the task is capped or uncapped. Routine coding work the marginal feature, the iterative bug fix may be capped. Microsoft, by one account, pulled Claude Code licenses from an entire division when the compute cost exceeded the cost of the people it was meant to help. Capped market, margin compression, no new demand created. The Stripe migration is different: it unlocked a project that wouldn’t have been approved at human-labour cost. That is genuine new demand.

The question is what fraction of real AI enterprise spending looks like Stripe and what fraction looks like tokenmaxxing. That ratio is genuinely contested, and anyone telling you the answer with confidence is selling something.

The business model is closing an obvious incoherence

On June 23, Fable 5 was scheduled to leave Anthropic’s flat-rate subscription plans. Want the best after that date: pay metered API credits, $10 per million tokens in, $50 out.

For two years, flat-rate subscribers on the $200 Max plan were extracting $600 to $1,500 per month in token-equivalent compute. The head of Claude Code called the economics “really hard to do sustainably.” GitHub made the same correction on June 1, shifting Copilot from flat-rate to usage-based billing; one developer’s projected monthly cost moved from roughly €67 to €966.

A subscription business monetises engagement shallowly the cost of each interaction is variable while the price is fixed. A metered API business monetises the Stripe migrations directly. The transition is closing an obvious incoherence. Whether it re-accelerates Anthropic’s revenue trajectory depends on how much of that $47 billion run rate was built on genuine uncapped demand versus subsidised usage a number the IPO filings will eventually force into view.

The open source ceiling

In April 2026, a free downloadable model called GLM-5.1 climbed to the top of SWE-bench Pro, the benchmark Fable 5 now leads. It held the crown for nine days before Anthropic shipped Opus 4.7 and the lead was gone.

This is the tempo of the market. Ramp’s enterprise spend data shows the average cost per million tokens fell from roughly $10 to $2.50 in a single year. Economist Luis Garicano, makes the structural case that intense competition and near-zero switching costs mean most surplus flows to customers not to model providers while upstream bottlenecks like TSMC and Nvidia hold the scarcity rents. Switching providers is a configuration change. No installed base. No data gravity. No retraining cost.

The model provider in the middle gets squeezed from below by open source and from above by the infrastructure layer. This was the bull case for pure model plays: build the moat fast enough to escape that squeeze before the ceiling descends.

The wall Anthropic built

Fable 5 ships with safety classifiers that specifically reroute distillation queries to Opus 4.8. Distillation queries are prompts engineered to extract a frontier model’s reasoning patterns — the exact mechanism by which smaller open-weight models are trained to approximate expensive frontier ones. The locked-down twin, Mythos 5, sits behind individual vetting in Project Glasswing, inaccessible to any public API. You cannot distil what you cannot access.

This is a deliberate attempt to break the commoditisation pipeline at its source. The wall slows imitation — it does not prevent it, because open-source labs have routes through reinforcement learning, synthetic data, and model merging that don’t require Anthropic’s outputs. But it buys time. And the real moat may not live in the weights at all: enterprise integration depth, compliance certifications, security audits, workflow dependencies. AWS and Salesforce became enormous businesses without unassailable technology. What they held was operational embeddedness switching costs that rose not because the product was irreplaceable, but because it was structurally woven in.

Anthropic might win the margin race even if open-source closes the capability gap. The question is whether it uses the current frontier window to embed deeply enough before the pricing floor arrives.

That was the thesis as of Thursday.

A letter arrived at 5:21 PM

At 5:21 PM on Friday, June 13, a directive arrived from Howard Lutnick’s Commerce Department. It required Anthropic to immediately suspend all access to Fable 5 and Mythos 5 for any foreign national whether inside or outside the United States, including Anthropic’s own foreign national employees.

Anthropic had no reliable way to verify which users were US citizens and which weren’t. So they turned it off for everyone.

Not degraded. Not restricted. Gone.

There is credible reporting that this is partially political retaliation. The Pentagon had labeled Anthropic a “supply chain risk” after Anthropic refused to remove its acceptable use restrictions on autonomous lethal weapons and mass domestic surveillance. A federal judge in San Francisco temporarily blocked that designation, finding Anthropic likely to succeed on its First Amendment retaliation claims and writing that the government’s measures “appear designed to punish Anthropic.” Friday’s Commerce directive looks like the administration finding a different lever.

The political context matters but doesn’t change the structural point. Anthropic’s safety narrative gave the government a workable justification: when you spend months publicly arguing your model is uniquely dangerous, you provide the framing that a national security directive requires. Anthropic’s own pushback noted that the demonstrated jailbreak found vulnerabilities “widely available from other models including GPT-5.5” their strongest legal argument, and simultaneously a damaging self-disclosure about the product’s claim to singular danger.

The third variable

This is the variable that was missing from every bull and bear case on frontier AI.

The question was whether model providers capture Nvidia-like margins or get competed down to airline economics. Friday introduced a third option: the government turns it off.

This is not unprecedented export controls on 40-bit cryptography in the 1990s, the PlayStation 2’s GPU export restrictions, ITAR applied to dual-use software. In each case: real capability, real controls, controls eventually obsolete, delay real in the meantime.

For enterprises, what changed on Friday is categorical. A key vendor changing pricing or going out of business is a known risk category. What became empirical on Friday is different: a US government letter, with four hours’ notice, can zero out production infrastructure. Anthropic’s own statement warns that “if this standard was applied across the industry, we believe it would essentially halt all new model deployments for all frontier model providers.”

For investors the implications compound. Pure model plays now face squeezed margins from open source and a regulatory overhang that closed-model providers can’t price away. Infrastructure plays the hardware layer, the energy layer, the open-weight labs look structurally cleaner.

Large enterprises will solve the compliance problem they have lawyers and compliance teams for exactly this, and multinationals operating across dozens of countries will find ways to route around a citizenship-verification requirement that is practically unenforceable at the query level. What they won’t do is un-see the precedent. Sovereign AI goes from cautious CIO talking point to line item in the next procurement cycle. The immediate disruption gets resolved. The strategic calculation doesn’t.

The revised question

There is a version of Friday that resolves quickly. Anthropic fights in court, the administration walks it back, Fable returns behind some form of compliance framework, and the metered transition happens on a delayed schedule. That is probably the most likely outcome.

But the precedent is set regardless. You cannot un-demonstrate that a US government directive can shut down a frontier model in four hours. Smart players were already hedging. This makes hedging table stakes.

Frontier AI isn’t purely a market good anymore. It’s strategic infrastructure with government oversight.

Anthropic built the anti-distillation wall. They embedded the classifiers. They gated Mythos behind individual vetting. They built the most elaborate moat in the industry around capabilities they had described as uniquely dangerous.

The letter arrived anyway.

The wall they built was real. It just wasn’t the only wall in the room and they didn’t get to choose which one enclosed them.