Why do some AI models have experts?
Mixture of experts: What you can afford to hold and what you can afford to run stopped being the same number.
the paper →Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityFedus, Zoph and Shazeer, 2021 ↗This is why a model that behaves like it has hundreds of billions of parameters can still answer as fast and as cheap as one a tenth the size — most of it sits idle on any given word. Only a handful of experts fire per token, so it runs like a small model and has to be held in memory like a huge one.
The room it came out of
The idea dates to 1991 and sat unused for thirty years, because nobody had a reason to want a model with far more parameters than it runs. Google revived it in 2021 for a purely economic reason: the scaling laws wanted parameters and the budget refused to pay for the arithmetic, and this is the one architecture where those two can be separated.
How it works
A dense model puts every token through every parameter. A mixture-of-experts model does not: each layer holds many parallel sub-networks, a small router picks two or three of them per token, and the rest sit idle. A model with six hundred billion parameters might do the arithmetic of a thirty-billion one.
That sounds like a straightforward win and it is not, because the idle experts are still in memory. The router cannot know in advance which expert the next token will want, so all of them have to be resident and ready. Arithmetic scales with the experts you use; memory scales with the experts you have.
Which is why a mixture-of-experts model has two sizes and press releases quote the flattering one. "Active parameters" is what it costs to run a token. "Total parameters" is what it costs to hold the model at all, and it is the number that decides how many machines you need before you serve anybody. A reader who sees one figure for an MoE model is being shown whichever suits the argument.
There is a third cost that only appears at scale. If the experts are spread across machines — and at frontier size they must be — then every token's chosen experts have to be fetched across the network, every layer, in a pattern nobody can predict. That is the all-to-all traffic on the Networking layer, and it is a large part of why the fabric inside a rack became a product people argue about.
So the honest framing is not that sparsity made models cheaper. It is that sparsity moved the cost off arithmetic and onto memory and interconnect, which is the same shape as everything else in this atlas, and it happened to move it toward the two resources that were already the tightest.
What it traded
- gave up
- memory — every expert must be resident whether or not it runs
- got
- arithmetic — only a small fraction of the model touches any given token
What exists now that didn’t before
Every expert has to be resident whether or not it runs, so the saving in arithmetic arrived as a bill in capacity and interconnect.
Asked, and answered
How can a huge model answer as fast as a small one?
Only a handful of its internal 'experts' actually fire for any given word — the model has to be held in memory like a giant one, but it computes like a small one, because most of it sits idle on every single token.
Read next
- HBMWhere the idle experts sit, and why holding them is the expensive part.
- NVLink and InfiniBandWhat carries an unpredictable expert lookup across machines, every layer, every token.
- ParametersWhy the single number everyone quotes stopped meaning one thing.