Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Visualization · 26 Sept 2026

The bets against the wall

Every startup attacking AI's bottlenecks is betting that a wall will still be standing when its product ships. Whether it's right about the wall decides what it builds. Who has to buy it tends to decide how that gets sold, and how the story goes.

trying to answer →What becomes scarce when intelligence becomes cheap?What is a moat in an AI world?

Building AI runs into one wall after another: the chip, then the memory next to it, then the package that joins them, then the heat, then the power. This map holds 35 companies founded to get around one of those walls. One keeps the whole AI model on the chip so it never waits for memory. One sends data between chips as light instead of down copper. Several are building small nuclear reactors.

Each is a thesis, and a thesis about a wall has three parts. Which wall: will it still be standing by the time you ship? What you give up to get around it. And who has to buy you. The first two are engineering, and they decide what a startup builds. The third is strategy, and it has done more than the other two to decide how the product gets to market. So far there are four ways that has gone. Each card on the map below keeps two things apart that are easy to blur: how the company’s story has gone, and whether anyone is actually using its product. An investor, a strategic partner or a pilot is not a paying customer, and the map doesn’t count one as a customer.

Sold or licensed to the wall. When one rich company owns the wall, beating it has often meant selling to it. Groq built a chip that never waits for memory; NVIDIA, whose chips do, agreed to pay about $20bn for a non-exclusive licence and hired Groq’s founder. Groq was not bought. It kept its company, raised $650m in 2026 and now runs an inference cloud, and the Justice Department is looking at the deal. NVIDIA paid over $900m to license Enfabrica’s technology and hire its team; Enfabrica had made tens of thousands of chips act as one. Marvell completed its purchase of Celestial AI in February 2026, for about $3.25bn plus earnouts. AMD has agreed to acquire Taalas, which is due to close by the end of the year. The wall bought the battering ram, or rented it.

Bought by the industrials. Cooling went the same way, one level down. One after another, the leading liquid-cooling companies were bought by companies that already sell equipment to buildings: Flex took JetCool, Schneider took 75% of Motivair, Trane took LiquidStack, Eaton paid $9.5bn for Boyd Thermal and Ecolab about $4.75bn for CoolIT. These were not bets on paper. Motivair’s cooling is part of more than $290m of kit at TeraWulf’s Lake Mariner campus, and LiquidStack had a 300 MW order from a US operator it didn’t name. When a fix slots into someone’s existing catalogue, the catalogue’s owner tends to end up owning the fix.

Backed by the customer. Power is different, because nobody owns the grid’s wall. There is no incumbent to sell to, so the customers finance the attack themselves. Google signed to take up to 500 MW from Kairos’s molten-salt reactors, the first due in 2030, and 396 MW from Fervo’s geothermal wells in Utah from 2028. Amazon anchored X-energy’s funding and plans more than 5 GW with it. Meta agreed to prepay for power from Oklo’s planned 1.2 GW campus in Ohio. The same thing happens in chips when a buyer wants a second supplier badly enough. OpenAI signed for up to 750 MW of Cerebras compute, about $20bn by Cerebras’s own IPO filing, and also lent it $1bn and took warrants. When nobody stands on the wall, the buyer becomes the financier.

Still independent. Most of the map, and “independent” covers very different positions. Some already have customers you can name. Jane Street has Etched’s first rack, and Etched claims more than $1bn in contracts. Tenstorrent’s servers run at ai& in Japan, Virtu Financial and Cirrascale. The inference cloud Parasail is deploying d-Matrix’s chips. JPMorganChase picked SambaNova to run AI on its own premises, and India’s NxtGen has Akash’s diamond-cooled servers. Some have partners but no paying customer on record: Ayar Labs is designing racks with Wiwynn, and Lightmatter joined NVIDIA’s NVLink Fusion ecosystem. Emerald AI, whose software eases a data centre off when the grid is strained, has pilots with Silicon Valley Power and Digital Realty, and demonstrations with a long list of utilities, clouds and grid operators. Those are partners, not customers. And some are simply early: MatX, Positron, Wafer and Substrate have no public customer yet, and Anthropic’s reported talks with Fractile are just that. Several of these companies are young enough that the question hasn’t come up yet.

Look under the four and there are five situations, and each one tends to produce its own kind of deal:

  • An incumbent controls the bottleneck. The startup gets licensed, acquired or folded into the incumbent’s plans.
  • The fix is physical kit and the buyer is an industrial. An industrial company buys it.
  • The bottleneck is generation. The customer pays through offtakes, prepayments and strategic money.
  • The fix is software that makes something more efficient. It sells through partnerships and cloud deployments.
  • Nobody owns the wall at all. The startup may have to become the infrastructure itself. Crusoe is the clearest case: it pulls the site, the power and the building together, and so has nobody to wait for. Groq running its own cloud is another.

Which suggests the scarce thing is often not the bottleneck itself. It is control of the interface between the bottleneck and the buyer. NVIDIA owns the relationship with everyone buying AI chips, so a better chip idea is worth more inside NVIDIA than against it. Eaton and Schneider own the relationship with the building. Google, Amazon and Meta own the demand for power, so they get to choose which reactor gets built. A startup that fixes the bottleneck without reaching the buyer tends to end up selling its fix to whoever does.

Two places are nearly empty, and both tell you something. Only three small bets have gone against advanced packaging, the step that joins chips to each other and to their memory in one package, though it has been one of the tightest walls of all. Two of the three lean on government money. TSMC holds this step so completely that a newcomer struggles to get qualified. And almost nobody is betting against the assembly of racks, because many companies can do it, the margins are thin, and the big assemblers already have the customers. A wall with no startups might be unbreakable. It might also be owned so thoroughly by an incumbent that nobody can reach its buyers, or be a market too thin to be worth attacking. Those are three different findings, and the empty space alone can’t tell you which one you’re looking at.

the whole map: 35 companies against 11 of the 13 steps, from the chip factory to the grid. open any one to read the bet

  1. Demand for AI

    Startups sell to demand; they do not bet against it.

  2. publicCoreWeavefounded 2017 · any AI workRent out NVIDIA’s chips faster than the big clouds can, to whoever cannot wait.Still independentIn productionwho pays for it: Meta, OpenAI and Anthropic, among others, on multi-year contracts.

    The bet, in detail. A cloud built only for AI: it buys NVIDIA systems as soon as they ship, puts them in rented data halls, and leases the capacity to the labs on multi-year contracts. A bet that speed to deploy the newest chips is worth more than the breadth of services the big clouds sell.

    The money. Listed on Nasdaq in March 2025 at $40 a share, raising about $1.5bn; NVIDIA is an investor.

    What it gives up. Everything on debt secured against the chips, with a few very large customers and chips that age in three to five years.

    Backers and partners, not customers. NVIDIA is its supplier, an investor, and the buyer of up to $6.3bn of any capacity it cannot sell.

    What happened. Public. Revenue doubled to $2.58bn in the second quarter of 2026, with a backlog of about $104bn.

    CoreWeave, second quarter 2026 results (11 Aug 2026) · Motley Fool, CoreWeave’s $6.3 billion backstop deal with Nvidia (Oct 2025)

    privateLambdaany AI workBuy chips on credit and rent them to the clouds, including the company that made them.Backed by its customerIn productionwho pays for it: Microsoft, in a multi-billion-dollar deal, and NVIDIA, which rents back 18,000 of its own GPUs for $1.5bn over four years.

    The bet, in detail. An AI cloud that grew from selling GPU workstations to researchers. It buys NVIDIA systems, runs them in liquid-cooled US data centres, and rents them out. A bet that the big clouds will keep renting capacity from someone else while they build their own.

    The money. $1bn of private debt in August 2026 to buy chips for Microsoft; reported to be raising up to $3bn before an IPO targeted for 2027.

    What it gives up. It is financed by debt against the chips, and its biggest customers are also its biggest competitors.

    What happened. Private.

    TechCrunch, Neocloud Lambda secures $1B in debt to buy more chips (28 Aug 2026) · DCD, Nvidia signs $1.5bn deal to lease its GPUs back from Lambda

    publicNebiusany AI workA full AI cloud built from Europe outward, sold in five-year blocks to the biggest buyers.Still independentIn productionwho pays for it: Microsoft, for $17.4bn to $19.4bn over five years, and Meta.

    The bet, in detail. The AI cloud spun out of Yandex’s international business: its own data centres, its own server designs and its own software on NVIDIA chips. A bet that the hyperscalers will rent capacity rather than wait to build it.

    The money. Public on Nasdaq as NBIS.

    What it gives up. A few contracts make up most of the business, so one customer’s change of plan moves everything.

    What happened. Public. Aiming for $7bn to $9bn of annual run-rate revenue by the end of 2026, more than half of it already booked.

    Techzine, Nebius expands through billion-dollar deals with Microsoft and Meta · Motley Fool, Nebius signed $46 billion in AI cloud deals with Microsoft and Meta (Apr 2026)

    privateFluidstackany AI workOwn no power plant and almost no chips: turn other people’s sites into AI clouds, with a giant standing behind the lease.Backed by its customerIn productionwho pays for it: Anthropic, Meta and Mistral, among others.

    The bet, in detail. An AI cloud that leases converted crypto-mining sites and fills them for the labs. A bet that the scarce skill is stitching together a site, a lease and a customer, not owning any of them.

    The money. $1.5bn led by Jane Street at an $18bn valuation in September 2026.

    What it gives up. Its leases are long and its customers few; without its backer the leases would be hard to sign at all.

    Backers and partners, not customers. Google backstops its lease payments to TeraWulf and Cipher Mining, and took warrants in TeraWulf.

    What happened. Private. Hosting at TeraWulf’s Lake Mariner and Cipher Mining’s Texas sites.

    Forbes, Fluidstack hits an $18 billion valuation (3 Sep 2026) · DCD, Fluidstack signs additional lease with Cipher Mining, backed by Google

    Renting out the chips (neoclouds)

    The circular part of the boom: the chip maker invests in the clouds that buy its chips, and the clouds’ customers back them too.

  3. privateSubstratechip-makingA new way to print chips, with X-rays, around the one company that makes the machines today.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. Substrate is building lithography that uses X-rays from a compact particle accelerator instead of ASML's EUV light, and plans to run its own foundry with it. It is a bet against the gate behind this whole station: one company makes the scanners every leading chip is patterned with, and they ship in tens a year.

    The money. $100m in October 2025 at a valuation above $1bn, with Founders Fund, General Catalyst and In-Q-Tel among the investors.

    What it gives up. Nothing is in production. It aims for volume by 2028 against the most mature tool supply chain in the industry.

    What happened. Private and pre-revenue.

    DCD, Substrate raises $100m for lithography tools to challenge ASML (Oct 2025)

    Chip factories

    Almost all the money here is the incumbents’ own spending on new factories. One startup challenger is on record.

  4. privateEtchedfounded 2022 · for running modelsA chip that can only run today’s kind of AI model, and runs it faster for it.Still independentIn productionwho pays for it: Jane Street, which has its first rack; the company claims more than $1bn of contracts.

    The bet, in detail. Etched was founded by two Harvard dropouts on one wager: that the transformer would still be the architecture in use when their chip shipped. A chip that can only run transformers spends none of its area on anything else, so more of the die is arithmetic. It is this station taken to its end: if throughput per socket is what binds, remove everything that is not throughput.

    The money. $500m at about a $5bn valuation in January 2026, then a reported $300m at $10.3bn in July and a reported $700m at $21bn in August 2026.

    What it gives up. If the architecture changes, the chip cannot follow. It is still fed by HBM and still waits in the same TSMC queue.

    What happened. Private. Jane Street, its first named customer, received its first rack in July 2026; Etched says it has more than $1bn of customer contracts.

    DCD, Etched raises $500m for a $5bn valuation (Jan 2026) · Etched, raises $700M at a $21B valuation and completes first customer delivery to Jane Street (18 Aug 2026) · Startup Fortune, Etched raises $300M at a $10.3B valuation (Jul 2026) · Memeburn, Etched doubled its valuation to $21B in 26 days (Aug 2026)

    privateMatXfor training modelsA chip for large language models and nothing else, for training as well as answering.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. MatX was founded by Reiner Pope and Mike Gunter, who worked on Google's TPUs, to build a processor for large language models and nothing else, spending more of the die on the matrix arithmetic those models need and less on generality. It is the Etched argument applied to training as well as serving.

    The money. $500m Series B in February 2026, led by Jane Street and Situational Awareness, with Marvell among the investors.

    What it gives up. Anything that is not a large language model runs worse or not at all, and it waits in the same TSMC queue as everyone else.

    Backers and partners, not customers. Jane Street led its round, as an investor.

    What happened. Private. The MatX One is due to ship in 2027.

    Bloomberg, AI chip startup MatX raises $500 million to compete with Nvidia (24 Feb 2026) · SiliconANGLE, chip startup MatX raises $500M to speed up large language models

    privateWaferfounded 2025 · for running modelsNo chip at all: AI that rewrites the code chips run, so cheaper chips can keep up.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. Wafer makes no chip. It builds AI agents that rewrite and tune the low-level code GPUs run, so the same hardware does more work. Its sharpest claim is aimed at this station's moat: that tuned code brings AMD's MI355X to 80% of NVIDIA B200 throughput at under half the cost. If code can be made fast without years of hand tuning, the reason buyers pay NVIDIA's margin shrinks.

    The money. $40m Series A at over $200m in September 2026, co-led by Marathon and Chemistry, with AMD Ventures participating; its seed was $4m in April 2026.

    What it gives up. The gains are its own benchmarks on chosen models, and a year-old company of about ten people is betting against the largest software ecosystem in the industry.

    What happened. Private. Reported to have turned down acquisition offers.

    Wafer, raises $40M Series A (Sep 2026) · FinSMEs, Wafer raises $40M in Series A funding (Sep 2026)

    privateTenstorrentfounded 2016 · for training modelsAttack NVIDIA’s real moat, its software, by giving software away.Still independentIn productionwho pays for it: ai& in Japan, Virtu Financial, Cirrascale and Turiyam, and Equinix’s AI hub.

    The bet, in detail. Tenstorrent, led since 2023 by the chip architect Jim Keller, builds on the open RISC-V instruction set and publishes its software. It is a bet on which part of NVIDIA's position lasts: not the silicon, which others can match, but the software nobody wants to rewrite. Open software is how you attack that, and it is why the company also licenses its designs instead of only selling chips.

    The money. $693m Series D in December 2024 at a $2bn pre-money valuation, led by Samsung Securities and AFW Partners.

    What it gives up. It competes on price and openness against a moat measured in years of developer habit, and it sits in the same packaging and memory queues as everyone else.

    Backers and partners, not customers. Samsung and LG license its designs and invest.

    What happened. Private. Its Galaxy servers are deployed with ai& in Japan, its largest installation, and with Virtu Financial, Cirrascale and Turiyam; sixteen sit in Equinix’s Ashburn hub. Samsung and LG are licensees and investors, not the same thing as customers.

    Tenstorrent, closes $693M+ of Series D funding (Dec 2024) · Tenstorrent, enables AI at scale (Apr 2026)

    privateSambaNovafounded 2017 · for running modelsHold many AI models in memory at once and switch between them quickly.Still independentContract, order or pilotwho pays for it: JPMorganChase, for on-premises inference; SoftBank is its first deployment partner.

    The bet, in detail. SambaNova builds a dataflow processor with three tiers of memory: on-chip SRAM, HBM, and a large pool of ordinary DRAM. That lets many models sit in one system and be switched between quickly. The bet is that serving is limited by how many models and how much context you can hold, not only by bandwidth.

    The money. $1bn first close of a Series F at an $11bn valuation in July 2026, led by General Atlantic.

    What it gives up. Customers must adopt its own software stack instead of CUDA.

    What happened. Private. JPMorganChase chose it in July 2026 to run AI inference on its own premises, on SN40 and SN50 systems. SoftBank is its first deployment partner; it did not fund the round.

    TechCrunch, SambaNova raises $1B at $11B valuation (8 Jul 2026)

    The AI chip

    and if it works, where does the wall go?
    • Fixed-function systolic arrays (TPU, Groq) deployed then: The same memory wall as everyone else
    • Custom silicon for a single buyer deployed then: You are still in the same packaging queue
    • Optical matrix multiply research then: Getting data into and out of the optics
  5. licensedGroqfounded 2016 · for writing repliesKeep the whole AI model on the chip, so it never waits for memory.Sold or licensed to the incumbentIn productionwho pays for it: Developers on its own inference cloud. NVIDIA licensed the technology rather than buying chips.

    The bet, in detail. Groq was started by Jonathan Ross, who had begun Google's TPU project. Its chip keeps a model's weights in on-chip SRAM, with no HBM at all, and fixes the timing of every operation when the program is compiled. That only pays when generating tokens, where an accelerator spends its time waiting for weights to arrive from memory rather than doing arithmetic with them. An all-SRAM chip is a bet that decode stays limited by bandwidth.

    The money. NVIDIA agreed in December 2025 to pay about $20bn for a non-exclusive licence to its technology, and hired its founder and president. Groq then raised $650m in June 2026 for its cloud.

    What it gives up. SRAM holds hundreds of megabytes a chip, not tens of gigabytes, so one large model is spread across hundreds of chips and several racks.

    What happened. Now an inference cloud. Its chip lives on inside NVIDIA as the LP30 in the Groq 3 LPX rack, in production from August 2026, and the Justice Department opened an investigation into the deal in September 2026.

    Bloomberg, Nvidia reaches technology licensing deal with Groq (24 Dec 2025) · Groq, non-exclusive inference technology licensing agreement with NVIDIA · Groq, raises $650M to scale its AI inference cloud (22 Jun 2026) · Tom's Hardware, Nvidia presents Groq 3 LPX architecture at Hot Chips 2026 · Bloomberg, DOJ probes Nvidia’s license deal with Groq (10 Sep 2026)

    privated-Matrixfounded 2019 · for writing repliesDo the maths inside the memory, so less has to be moved.Still independentIn productionwho pays for it: Parasail, an inference cloud, is deploying it beside its GPUs.

    The bet, in detail. d-Matrix does the multiplication inside the memory array instead of fetching operands to separate arithmetic units, an approach it calls digital in-memory compute. It builds only for inference, because at small batch sizes inference spends most of its energy moving weights rather than using them. The company is a bet that moving weights, not multiplying them, stays the cost that matters.

    The money. $275m at a $2bn valuation in November 2025; $450m raised in total.

    What it gives up. Capacity per chip is small, and it does not train.

    Backers and partners, not customers. Gimlet Labs published tests of it.

    What happened. Private. Its Corsair accelerator entered full production in June 2026, and the inference cloud Parasail is deploying it beside its NVIDIA GPUs.

    DCD, d-Matrix raises $275m against $2bn valuation (Nov 2025)

    privateFractilefor running modelsPut the maths where the model is stored, so nothing has to be fetched.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. Fractile, in London, puts the arithmetic inside the memory on one die, so the weights are used where they are stored instead of being fetched across a bus. It aims at the same wall HBM was built to push back, and claims frontier models could run up to 100 times faster and ten times cheaper than on current GPU systems.

    The money. $220m in May 2026, co-led by Accel, Factorial Funds and Founders Fund.

    What it gives up. Its first commercial chip is not expected before 2027, and until then the claims are its own.

    Backers and partners, not customers. Anthropic was reported to be in talks about its chips; nothing confirmed.

    What happened. Private. Anthropic is reported to be in early talks to buy its chips.

    The Next Web, Fractile raises $220m to take its in-memory-compute chip into production (May 2026) · DCD, Fractile raises $220m to accelerate development of AI inference chips

    publicCerebrasfounded 2015 · for writing repliesBuild one chip the size of a dinner plate, so data never has to leave it.Backed by its customerIn productionwho pays for it: OpenAI, for up to 750 MW through 2028, put at about $20bn in its IPO filing.

    The bet, in detail. Cerebras builds one chip from an entire wafer instead of cutting it into dies, and keeps the weights in SRAM spread across it, so operands never leave the silicon on their way to the arithmetic. It answers the problem HBM answers, operands arriving too slowly, by refusing to go off-chip at all. Its strongest results are in generating tokens, the phase where that problem is worst.

    The money. Listed on Nasdaq on 14 May 2026 at $185 a share, raising $5.55bn, four months after OpenAI agreed to buy up to 750 MW of its compute through 2028. Reported at over $10bn in January, the agreement is put at about $20bn in its IPO filing, and OpenAI also lent it $1bn and holds warrants for about a tenth of it.

    What it gives up. A wafer-sized part has its own yield, power delivery and cooling problems, and a model larger than the on-wafer memory has to be streamed in from outside.

    Backers and partners, not customers. OpenAI is also its lender and holds warrants.

    What happened. Public as CBRS. It opened at $350 on its first day.

    TechCrunch, Cerebras raises $5.5B, then stock pops (14 May 2026) · Cerebras, pricing of initial public offering · CNBC, Cerebras scores OpenAI deal worth over $10 billion (14 Jan 2026) · Cerebras Systems, Form S-1 (Apr 2026)

    privatePositronfounded 2023 · for running modelsUse the cheap memory found in phones and laptops, and simply use a lot of it.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. Positron builds inference chips around memory capacity rather than peak arithmetic, and uses LPDDR5X, the commodity memory in phones and laptops, instead of HBM: its Asimov chip carries 288 GB to 2,304 GB. The bet is that serving large models is limited by how many bytes sit next to the chip, and that leaving the HBM queue is worth the lower bandwidth per byte.

    The money. $875m at a $5bn post-money valuation in September 2026, co-led by NEA, Atreides, Valor, Andra and SemiAnalysis Capital.

    What it gives up. LPDDR5X delivers far less bandwidth per chip than HBM, so it wins on capacity and supply rather than speed. Asimov has not taped out yet.

    What happened. Private. Asimov is due to tape out on TSMC N3P at the end of 2026, with production in the second half of 2027.

    SiliconANGLE, Positron nabs $875M to speed up inference with consumer-grade memory (10 Sep 2026) · Positron AI, raises $875 million at a $5 billion valuation

    acquiredTaalasfounded 2023 · for running modelsBake one AI model permanently into the chip: as fast as it gets, useless for any other model.Sold or licensed to the incumbentNo public customerwho pays for it: None named. AMD agreed to acquire it in August 2026.

    The bet, in detail. Taalas takes the idea to its end. It builds a chip for one model, with that model's weights stored in the silicon itself as read-only memory, so nothing is fetched from HBM at all. Its first chip, HC1, holds Meta's Llama 3.1 8B and generates around 14,000 tokens a second at about 200 W, by its own account.

    The money. Raised about $219m before AMD agreed in August 2026 to acquire it, terms undisclosed.

    What it gives up. Every new model needs a new chip, so it only pays for a model served unchanged, at volume, for long enough to pay for the mask set.

    What happened. AMD agreed to acquire it in August 2026; the deal is expected to close in the fourth quarter, subject to regulatory approval.

    AMD, acquires Taalas (6 Aug 2026) · DCD, Taalas raises $169m, unveils HC1 processor optimized for Llama 3.1 8B

    Memory next to the chip

    and if it works, where does the wall go?
    • More HBM stacks deployed then: Packaging supply
    • Faster HBM generation shipping then: Cooling the memory, not just the die
    • Deeper memory hierarchy deployed then: Scheduler and framework complexity
    • CXL memory expansion early then: Fabric bandwidth and latency
    • Taller stacks (12-high, 16-high) shipping then: Yield, and it bites hard
    • GDDR7 instead of HBM shipping then: Bandwidth, sooner than you wanted
    • Hybrid bonding, copper to copper early then: Bonder supply
  6. privateSilicon Boxchip-makingBuild new advanced-packaging capacity outside TSMC.Still independentIn productionwho pays for it: Chip designers who need chiplets packaged outside TSMC; none named, but it reports volume shipments.

    The bet, in detail. Advanced packaging for chiplets, done on large square panels rather than round wafers. A bet that packaging capacity outside TSMC can be built new, and built where governments want it.

    The money. A $150m Series B2, then SGD 100m of debt in June 2026; the European Commission approved about €1.3bn of Italian state aid for its planned plant in Novara.

    What it gives up. A new packaging house has to win qualification from chip designers one product at a time.

    What happened. Private. Over 250 million units shipped from its Singapore plant by early 2026, by its own account.

    Silicon Box, SGD 100M financing (8 Jun 2026) · Silicon Box, European Commission approves €1.3bn Italian state aid

    privateEliyanany AI workBetter links between the small chips that big chips are now built from.Still independentDesign or ecosystem partnerwho pays for it: None named yet.

    The bet, in detail. The links between chiplets: interface IP and chiplets for die-to-die and chip-to-chip connections. A bet that as designs split into many dies, the links between them bind.

    The money. $50m of strategic investment in January 2026 from AMD, Arm, Coherent and Meta, with Samsung Catalyst Fund and Intel Capital returning.

    What it gives up. It has to be designed into someone else’s chip; it cannot ship on its own.

    Backers and partners, not customers. AMD, Arm, Coherent and Meta are strategic investors; that is not the same as buying it.

    What happened. Private.

    Eliyan, $50 million in strategic investments (28 Jan 2026)

    privateSyentachip-makingA new way to make the fine copper wiring between chips inside one package.Still independentNo public customerwho pays for it: None named yet. It would sell to packaging houses and chip makers.

    The bet, in detail. Localized electrochemical manufacturing: a way of making the dense copper interconnects between chips in a package. A bet that the wiring of the package, not the chips, becomes the limit.

    The money. $26m Series A in April 2026, led by Playground Global and Australia’s National Reconstruction Fund; Pat Gelsinger, Intel’s former chief executive, joined its board.

    What it gives up. A new manufacturing process has to be adopted by packaging lines that change slowly.

    What happened. Private. Opening a US base in Arizona.

    DCD, Syenta raises $26m (Apr 2026)

    Advanced packaging (chip-to-chip)

    Little startup money for a wall this important. What there is leans on government support, and TSMC keeps expanding its own.

    and if it works, where does the wall go?
    • Bigger monolithic die deployed then: You stop at the reticle
    • 2.5D on interposer deployed then: Packaging capacity
    • Bridge (EMIB-style) shipping then: Still needs the substrate
    • 3D stacking (logic on logic) early then: Thermal, immediately
    • Panel-level packaging research then: Qualification time
  7. privateAyar Labsfounded 2015 · for training modelsSend data out of the chip as light instead of down copper wires.Still independentDesign or ecosystem partnerwho pays for it: No public customer yet.

    The bet, in detail. Ayar Labs puts optical input and output on the package itself, so data leaves the chip as light instead of running over copper to a pluggable module at the front of the rack. It exists because the links between accelerators now spend enough power, and use enough of the package edge, to limit how many accelerators can work as one.

    The money. $500m Series E in March 2026 at a $3.75bn valuation, with NVIDIA and AMD among its backers, then $150m more in September: $650m raised in 2026.

    What it gives up. Lasers fail more often than copper, and on the package they sit next to the hottest part in the building.

    Backers and partners, not customers. Wiwynn is building racks with it; NVIDIA and AMD invest.

    What happened. Private. Designing an optically connected rack with Wiwynn, the server maker. NVIDIA and AMD are investors; neither is a named customer.

    DCD, Ayar Labs closes $500m funding round (Mar 2026) · Ayar Labs and Wiwynn, co-packaged optics for rack-scale AI (11 Mar 2026) · Lightwave, Ayar Labs scales funding to $650M (Sep 2026)

    acquiredCelestial AIfounded 2020 · for training modelsDeliver data by light to anywhere on a chip, not just its edges.Sold or licensed to the incumbentNo public customerwho pays for it: None named before Marvell bought it.

    The bet, in detail. Celestial AI built what it calls a Photonic Fabric: optical links that can deliver data to any point on the die rather than only to its edge, where electrical input and output have to sit. The edge of the package had become a scarce resource, and this was a way around it.

    The money. Acquired by Marvell for about $3.25bn, completed February 2026, rising to as much as $5.5bn if revenue targets are met.

    What it gives up. Revenue was years away. The buyer paid for the time it would have taken to build the same thing.

    What happened. Part of Marvell, which expects first revenue from it in the second half of its fiscal 2028. In April 2026 Marvell cancelled the purchase orders Celestial had placed with POET Technologies, citing breaches of confidentiality.

    Marvell, completes acquisition of Celestial AI (2 Feb 2026) · Yahoo Finance, POET Technologies stock falls 46% after Marvell cancels orders (Apr 2026)

    privateLightmatterfounded 2017 · for training modelsConnect chips to each other with light. It began by trying to do the maths in light; moving the answers turned out to be the real problem.Still independentDesign or ecosystem partnerwho pays for it: No public customer yet.

    The bet, in detail. Lightmatter set out to do matrix multiplication in light and moved to what customers bought: a photonic interposer that connects chips to each other optically. The change of direction is the finding. Multiplying turned out to be the cheap part; moving the result between chips was the constraint.

    The money. $400m in October 2024 at a valuation above $4.4bn; $850m raised in total.

    What it gives up. The same thermal and reliability questions as any optics placed on the package.

    Backers and partners, not customers. Part of NVIDIA’s NVLink Fusion ecosystem.

    What happened. Private. Joined NVIDIA’s NVLink Fusion ecosystem in June 2026 and plans to sample its L20 optical engine late in 2026. No anchor customer disclosed.

    Sacra, Lightmatter valuation and funding · Lightmatter, joins NVIDIA NVLink Fusion (2 Jun 2026)

    licensedEnfabricafor training modelsOne networking chip that lets tens of thousands of AI chips work as one computer.Sold or licensed to the incumbentNo public customerwho pays for it: None named. NVIDIA licensed it and hired its team.

    The bet, in detail. Enfabrica, founded by veterans of Broadcom and Google, built networking silicon to connect tens of thousands of accelerators and the memory around them so that they work as one system. That is the problem this station is named for, and NVIDIA judged it worth more than $900m in September 2025.

    The money. Raised $260m in venture funding before NVIDIA paid more than $900m in cash and stock for a licence and its team.

    What it gives up. It never shipped at scale on its own.

    What happened. Licensed to NVIDIA. Its chief executive, Rochan Sankar, joined NVIDIA.

    CNBC, Nvidia spent over $900 million on Enfabrica CEO and technology (18 Sep 2025)

    Wiring between chips

    and if it works, where does the wall go?
    • Copper, for as long as it reaches deployed then: The size of one rack
    • Linear pluggable optics early then: Reach, again
    • Co-packaged optics early then: Serviceability
  8. Assembly into racks

    Almost no startup money here: many companies can assemble racks, so the margin is thin, and the big assemblers already own the relationships.

  9. privateCorintisfounded 2022 · any AI workCooling plates with tiny channels shaped to each chip’s hot spots.Still independentNo public customerwho pays for it: No public customer yet.

    The bet, in detail. Corintis designs cooling plates with networks of microscopic channels shaped to each chip's hot spots, so coolant reaches the heat where it is made instead of flowing through a uniform plate. In Microsoft's tests the approach removed heat up to three times better than today's common cold plates. It exists because the last millimetre above the die is now the largest resistance in the thermal path.

    The money. $24m Series A in 2025 led by BlueYard, then $25m more led by Applied Digital; $58m raised in total.

    What it gives up. Each design is specific to one chip, and channels that fine must not clog or leak over years of service.

    Backers and partners, not customers. Applied Digital, a data-centre operator, led its latest money as an investor.

    What happened. Private. Scaling toward more than a million cooling plates a year.

    DCD, Corintis raises $24m for microfluidic cooling systems · GGBa, Corintis secures an additional USD 25 million

    privateZutaCoreany AI workCool chips by boiling a liquid on them instead of pumping water past them.Still independentIn productionwho pays for it: Data-centre operators, across more than 75 deployments by its own account; none named here.

    The bet, in detail. Waterless two-phase cooling: a dielectric liquid boils on the processor and carries the heat away as vapour. A bet that as chips pass 4,000 W, water loops and their leaks become the limit.

    The money. $100m Series C in June 2026, with Mitsubishi Electric, Carrier Ventures and Samsung Ventures among the investors.

    What it gives up. A second fluid system with its own supply and handling, beside the water loops data centres already know how to run.

    Backers and partners, not customers. Mitsubishi Electric and Carrier, cooling incumbents, are investors.

    What happened. Private. More than 75 deployments, by its own account.

    SiliconANGLE, ZutaCore raises $100M (2 Jun 2026)

    privateAkash Systemsfounded 2017 · any AI workPut synthetic diamond under the chip: it moves heat several times better than copper.Still independentIn productionwho pays for it: NxtGen, in India, under a $27m contract; the first systems are delivered.

    The bet, in detail. Akash Systems began by growing gallium nitride on diamond for satellite radios, where heat in a small amplifier was the limit. It now puts synthetic diamond under GPUs in servers. That market only opened once liquid cooling had taken the bulk heat and the largest remaining resistance was the last millimetre above the die, which is exactly where diamond conducts several times better than copper.

    The money. A $27m contract to supply diamond-cooled servers to NxtGen in India, and a non-binding preliminary CHIPS Act award of $18.2m in direct funding.

    What it gives up. Diamond has to bond to silicon without the interface resistance eating the gain, and it costs far more than the copper it replaces.

    What happened. Private. It reports delivering the first diamond-cooled NVIDIA servers to NxtGen.

    Akash Systems, $27 million contract with NxtGen · Akash Systems, preliminary agreement for CHIPS Act funding

    Cooling

    Here the money went to acquisitions: Flex bought JetCool, Schneider Electric bought Motivair (75%), Trane bought LiquidStack, Eaton bought Boyd Thermal, Ecolab bought CoolIT Systems.

    and if it works, where does the wall go?
    • Vapour chamber deployed then: Rack ceiling stands
    • Direct-to-chip liquid deployed then: Junction-to-plate resistance
    • Immersion early then: The same junction resistance
    • In-die microfluidics research then: Yield and field service
    • Diamond spreader early then: Thermal boundary resistance
    • Negative-pressure loop early then: Pump capacity
    • Two-phase direct-to-chip early then: Fluid supply and its permits
  10. privateHeron Powerany AI workA transformer made of electronics instead of iron and copper.Still independentContract, order or pilotwho pays for it: Intersect Power and Crusoe, named as early customers; RWE signed a pilot in September 2026, for installation in April 2027.

    The bet, in detail. A transformer made of power electronics instead of iron and copper, connecting DC equipment (solar, batteries, AI hardware) straight to the medium-voltage grid. A bet that the transformer shortage is a design problem as much as a factory one.

    The money. $140m Series B in February 2026, co-led by Andreessen Horowitz and Breakthrough Energy Ventures; $183m raised in total.

    What it gives up. Power electronics have to prove the decades of service utilities expect of an iron transformer, and its field demonstrations only begin in mid-2026.

    What happened. Private. Building a 40 GW US factory, with partner installations in early 2027 and full production in the second half of 2027.

    TechCrunch, Heron Power raises $140M (18 Feb 2026) · Heron Power and RWE sign pilot agreement (21 Sep 2026) · DCD, Heron Power signs LOI with Crusoe for Stargate

    privateAmperesandany AI workTurn grid power straight into what AI racks need, skipping steps along the way.Still independentContract, order or pilotwho pays for it: PSA International, trialling it at the Port of Singapore, where its first commercial units went in early 2026. The hyperscalers it targets are not named.

    The bet, in detail. A medium-voltage solid-state transformer that converts utility distribution power (11 to 35 kV) directly to the DC voltage AI racks use, removing one or two conversion stages.

    The money. $80m Series A in November 2025, co-led by Walden Catalyst Ventures and Temasek; over $92m raised.

    What it gives up. It replaces equipment utilities and builders have trusted for a century, and has to be qualified at hyperscale.

    What happened. Private. Targeting 30 MW of commercial systems in 2026, aimed at hyperscale AI customers.

    DCD, Amperesand raises $80m (Nov 2025) · PSA International, PSA unboXed trials Amperesand solid-state transformers

    privateDG Matrixany AI workOne programmable box that routes power between the grid, batteries, solar and generators.Still independentDesign or ecosystem partnerwho pays for it: No named hyperscale customer found.

    The bet, in detail. A multi-port solid-state transformer that routes up to 2.4 MW between the grid, solar, batteries and generators under software control. A bet that a site fed from several sources needs one programmable box, not a room of switchgear.

    The money. $60m Series A in February 2026, led by Engine Ventures, with Mitsubishi Heavy Industries joining and ABB, an existing investor, putting in more.

    What it gives up. The same qualification problem as any new power equipment, in a market that buys on track record.

    Backers and partners, not customers. ABB, one of the incumbents it would displace, and Mitsubishi Heavy Industries invest.

    What happened. Private. Data centres are about 90% of its pipeline.

    DCD, DG Matrix raises $60m to scale solid state transformer for data center market (Feb 2026)

    Transformers and power gear

    and if it works, where does the wall go?
    • Legacy AC distribution deployed then: Cannot reach 100 kW a rack
    • Rack-level DC bus bar deployed then: The feeder outside
    • 800 V DC distribution early then: Same feeder, same queue
    • Behind-the-meter generation early then: Local permission, not physics
    • Site where power already is deployed then: There are only so many such sites
    • Solid-state transformer (SiC, GaN) research then: Wide-bandgap semiconductor supply
  11. privateCrusoeany AI workOwn the power, the land and the building, so there is nobody else to wait for.Still independentIn productionwho pays for it: The AI labs and clouds renting its capacity; its release does not name them.

    The bet, in detail. Integrate the site, the power and the infrastructure as one company: energy, land, campus and AI capacity, so the site starts where the power already is. A bet that turning power and land into compute quickly is itself the scarce skill.

    The money. $3.9bn Series F at a $30.9bn valuation in September 2026, co-led by Atreides, Mubadala Capital and Valor, with NVIDIA among the investors.

    What it gives up. Everything on one balance sheet: a data-centre company that owns its energy carries both businesses’ risks.

    What happened. Private. More than 6 GW of contracted capacity and 1 GW operating, by its own account.

    Crusoe, $3.9B Series F at a $30.9B valuation (17 Sep 2026)

    Buildings and land

    and if it works, where does the wall go?
    • Prefabricated megawatt blocks shipping then: Still the transformer
    • Brownfield: a dead smelter or coal plant deployed then: The old plant connection size, which is a hard ceiling
    • Build where the power already is deployed then: Long-haul fibre into somewhere nobody laid any
  12. privateKairos Powerfounded 2016 · power plantSmall reactors cooled by molten salt, with Google signed up as a customer before one exists.Backed by its customerContract, order or pilotwho pays for it: Google, under an offtake for up to 500 MW; the first unit is targeted for 2030.

    The bet, in detail. Kairos Power designs small reactors cooled by molten fluoride salt instead of water, so they run at low pressure. Google agreed in October 2024 to buy power from a fleet of them, up to 500 MW, with the first unit targeted for 2030. A buyer contracting for reactors that do not exist yet is the clearest sign the constraint had reached generation itself.

    The money. Private. The Google agreement is an offtake, not an investment.

    What it gives up. Nothing is built yet, and no reactor of this design is licensed for commercial power.

    What happened. First unit targeted for 2030 and all 500 MW by 2035, on the companies’ own schedule.

    DCD, Google signs nuclear SMR deal with Kairos (Oct 2024)

    publicOklofounded 2013 · power plantSmall reactors that sell electricity, not power plants, to data centres.Backed by its customerContract, order or pilotwho pays for it: Meta, which agreed to prepay for power from a 1.2 GW campus in Ohio.

    The bet, in detail. Oklo designs 50 MW fast reactors and proposes to sell the power rather than the plant, which is the form a data centre operator wants to buy. In December 2024 it signed a non-binding framework with Switch for 12 GW through 2044. It has not built a reactor yet; the first is targeted for 2027 at Idaho National Laboratory.

    The money. Listed in 2024 through a merger with a listed shell company.

    What it gives up. Everything rests on licensing a new reactor type and then building about 240 of them.

    What happened. Public as OKLO. No reactor operating. In January 2026 Meta agreed to prepay for power from a planned 1.2 GW Oklo campus in Pike County, Ohio, with the first phase targeted for 2030.

    CNBC, Oklo targets 12 gigawatts through agreement with Switch (18 Dec 2024) · Utility Dive, Oklo inks 12-GW supply agreement with Switch · Oklo, Oklo and Meta announce agreement for 1.2 GW in southern Ohio (9 Jan 2026)

    privateX-energyfounded 2009 · power plantSmall reactors that burn fuel sealed in ceramic pebbles, with Amazon paying ahead.Backed by its customerContract, order or pilotwho pays for it: Amazon, which anchored its round and plans more than 5 GW with it.

    The bet, in detail. X-energy designs small high-temperature gas-cooled reactors that burn fuel sealed in ceramic pebbles. Amazon anchored its financing in October 2024, and the two target more than 5 GW of new plants by 2039. Amazon is paying for the thing upstream of its own connection queue.

    The money. A $500m Series C-1 in October 2024, anchored by Amazon, later upsized to $700m.

    What it gives up. The fuel plant has to be built as well as the reactors, and the first units are years away.

    What happened. Private. More than 5 GW targeted with Amazon by 2039.

    X-energy, Amazon invests in X-energy (16 Oct 2024) · ESG Today, Amazon-backed X-energy raises $700 million

    publicFervo Energyfounded 2017 · power plantBorrow oil-drilling techniques to make geothermal power work almost anywhere.Backed by its customerIn productionwho pays for it: Google, which bought power from its Nevada pilot and signed for 396 MW in Utah from 2028.

    The bet, in detail. Fervo drills horizontal wells into hot dry rock and fractures it, using techniques from oil and gas to make geothermal work where there is no natural reservoir. Geothermal runs around the clock, which is what a data centre needs from its connection. Google bought power from its Nevada pilot from 2023, and in September 2026 agreed to take 396 MW from its Utah project, the largest enhanced geothermal contract on record.

    The money. Listed on Nasdaq on 13 May 2026, selling 70 million shares at $27.

    What it gives up. Drilling costs and how long a fractured reservoir keeps its heat are still being proven at commercial scale.

    What happened. Public as FRVO. The Utah contract starts delivering in 2028.

    Fervo Energy, pricing of its upsized initial public offering (May 2026) · Fervo Energy, Fervo Energy and Google sign 396 MW PPA (1 Sep 2026)

    Power plants

    and if it works, where does the wall go?
    • Behind-the-meter gas turbines deployed then: Gas supply and the permit itself
    • Restart a retired nuclear unit shipping then: There is no fifth one
    • Enhanced geothermal early then: Drilling capacity
    • Small modular reactors research then: Licensing, measured in years
  13. privateEmerald AIany AI workSoftware that lets a data centre ease off when the grid is strained, so the grid can take it on sooner.Still independentContract, order or pilotwho pays for it: Silicon Valley Power, under a flexible-connection pilot; Digital Realty, on a nearly 100 MW power-flexible AI factory.

    The bet, in detail. Software that flexes a data centre’s power draw with the grid’s condition, so the grid can connect more load than its worst hour allows. A bet that the scarce thing is not electricity but a place on the grid.

    The money. $150m Series A at a $1.05bn valuation in August 2026, co-led by Energize Capital and DCVC; over $220m raised.

    What it gives up. The data centre has to accept slowing or shifting its work when the grid is stressed.

    Backers and partners, not customers. Demonstrations with NVIDIA, EPRI, Oracle Cloud Infrastructure, Nebius, National Grid, SRP, APS, Dominion and PJM.

    What happened. Private. Five demonstrations at commercial data centres in Arizona, Illinois, Virginia, Oregon and London, and a pilot with Silicon Valley Power that gives flexible data centres faster grid access.

    Business Wire, Emerald AI raises $150M Series A at $1.05B (25 Aug 2026) · Silicon Valley Power and Emerald AI launch flexible data-centre pilot

    Plugging into the grid

    and if it works, where does the wall go?
    • Behind-the-meter generation early then: Local permission, not physics
    • Site where power already is deployed then: There are only so many such sites