Why does Kimi K2 matter?
Kimi K2: Moonshot AI’s open-weight model —
the paper →Kimi K2: Open Agentic IntelligenceKimi Team, Moonshot AI, 2025 ↗Moonshot AI’s open-weight model — the same mixture-of-experts-plus-latent-attention approach DeepSeek used, taken further: 1.04 trillion total parameters, 384 experts, only 32 billion active on any given token.
Built on
- Mixture of experts →This is why a model that behaves like it has hundreds of billions of parameters can still answer as fast and as cheap as one a tenth the size — most of it sits idle on any given word. Only a handful of experts fire per token, so it runs like a small model and has to be held in memory like a huge one.
- DeepSeek →This is why a lab locked out of the best chips could still ship a model that rivalled OpenAI’s and Google’s, and why it rattled the assumption that more capital always wins — by refusing to treat memory as somebody else’s problem. Its latent attention is the most aggressive attack yet on how much state a single token has to carry.
only a note so far: the paper is worth more than this, and it will get it