Why does Jamba matter?
Jamba: AI21’s open-weight model, and the hybrid mamba.left predicted before it existed: one attention layer for every seven Mamba layers, with mixture-of-experts layered on top, holding 256k of context in a fraction of the memory a pure-attention model would need.
the paper →Jamba: A Hybrid Transformer-Mamba Language ModelLieber et al., AI21 Labs, 2024 ↗AI21’s open-weight model, and the hybrid mamba.left predicted before it existed: one attention layer for every seven Mamba layers, with mixture-of-experts layered on top, holding 256k of context in a fraction of the memory a pure-attention model would need.
Built on
- Mamba →It is not what runs the AI you use today — attention came back into every serious version that followed it — but it is why "constant memory no matter how long the conversation runs" stopped being science fiction. A fixed-size summary instead of the whole transcript, paid for with everything that summary cannot hold onto.
- Mixture of experts →This is why a model that behaves like it has hundreds of billions of parameters can still answer as fast and as cheap as one a tenth the size — most of it sits idle on any given word. Only a handful of experts fire per token, so it runs like a small model and has to be held in memory like a huge one.
only a note so far: the paper is worth more than this, and it will get it