Back to glossary

Term

Kimi K2 Thinking

Kimi K2 Thinking (November 2025) is Moonshot AI's open reasoning-agent model — a trillion-parameter MoE with a 256K context that runs long thinking and tool-use chains on its own.

Kimi K2 Thinking — explained in more detail

Moonshot AI released Kimi K2 Thinking in early November 2025 (November 6) as an open model with freely available weights under a modified MIT license. It builds on Kimi K2 — a Mixture-of-Experts model with around one trillion parameters, of which roughly 32 billion are active per token. The Thinking variant is designed as a reasoning agent: it does not merely respond but plans, acts and self-corrects. Technically it uses INT4 quantization with QAT, cutting the model size from about 1 TB to roughly 594 GB and speeding up inference. The context window reaches up to 256K tokens.

A core feature is the interleaving of thinking and tool use: the model alternates between reasoning, calling tools, interpreting results and planning — per Moonshot, consistently across 200 to 300 tool calls.

Example / Practical use

For agentic applications, the long, autonomous tool-use chain is the real value — for instance research agents that run many steps without human intervention. A “Heavy Mode” runs eight independent reasoning paths in parallel for difficult problems. In benchmarks Moonshot reports 43% on Humanity’s Last Exam and positions the model ahead of GPT-5 and Claude Sonnet 4.5. Because the weights are open, Kimi K2 Thinking can be self-hosted; it is also accessible via the Kimi platform and API providers.

Distinction

Kimi K2 Thinking belongs to the class of open-weight models and thus competes with proprietary frontier models. Within the Kimi range, the Thinking variant is the reasoning- and agent-specialized offshoot of Kimi K2 — the focus is not fast single answers but the ability to work through complex, multi-step tasks autonomously.

See everything in one place:Kimi