Kimi (Moonshot AI) — the open model line at a glance
Where things stand
As of September 2026, Kimi K3 (July 2026) is the flagship of the line. This article covers the line as a whole — parameter counts, context windows and prices for individual versions live in the linked glossary entries.
What Kimi is and what the line stands for
Kimi is the open model line from the Chinese AI lab Moonshot AI. The weights of every generation are free to download, including on Hugging Face — so self-hosting is possible across the whole line. Two things characterise Kimi: a very large context window and a consistent design for agentic work with many tool calls.
The current flagship Kimi K3 pushes that to the limit. According to the maker it has 2.8 trillion parameters, making it the largest published open-weight model to date. That size is the point: Kimi positions itself as the open frontier option for tasks where smaller models run out of context or tool coordination.
The line’s strength profile
- Coding 4.5 / 5 · Claude Opus 5 u. a.
- Reasoning 4 / 5 · Claude Opus 5 u. a.
- Text 4 / 5 · Claude Fable 5.1
- Vision 3.5 / 5 · Gemini 3.1 Pro
- Speed 3.5 / 5 · Claude Haiku 4.5 u. a.
- Kosten-Eff. 4.5 / 5 · GPT Luna u. a.
Eignung 0–5 · redaktionelle Einordnung, kein Benchmark · gestrichelt = Feld-Bestwert je Achse
Kimi is strong in coding and cost efficiency, a solid all-rounder across the remaining axes. The axes are an editorial assessment, not a benchmark — they show the balance, but don’t replace a test on your own use case.
Where Kimi sits in the field
Redaktionelle Einordnung, kein Benchmark
Next to DeepSeek V, GLM and Qwen, Kimi sits in the strong all-rounder band with a good cost profile. What sets Kimi apart from the other open lines doesn’t show on this map: the sheer size of the context window and the scaling of its agent-swarm feature, which lets many model instances work on a task in a coordinated way.
Use profile — open, large, agentic
Kimi is at its best where context length and tool coordination matter:
- Agentic workflows with many tool calls and multi-step action chains.
- Multi-agent systems (agent swarm), where several instances collaborate.
- Tasks with a very long context — large document sets, long codebases.
- Self-hosting when data can’t leave the building. The weights sit on Hugging Face.
The weights are free, but self-hosting costs ride on your own infrastructure — a 2.8-trillion-parameter model needs serving. If you don’t want to host it yourself, you reach for hosted access. Basics on that are under Running local LLMs.
Which Kimi version fits
For new projects the current flagship Kimi K3 is the right pick. Older versions matter when you run smaller setups or compare behaviour across generations:
- Kimi K3 — current flagship (July 2026).
- Kimi K2 Thinking — reasoning-oriented K2 variant.
- Kimi K2.6, K2.5 and K2 — earlier generations.
Related terms
- AI model families at a glance — where Kimi sits next to the other open lines.
- Running local LLMs — self-hosting options.
- The Hugging Face ecosystem — where the weights live.
FAQ
- Kimi is the open model line of the Chinese AI lab Moonshot AI. The weights of every generation are free to download, including on Hugging Face.
- Kimi K3, released in July 2026. According to the maker it has 2.8 trillion parameters, making it the largest published open-weight model to date.
- Agentic workflows with many tool calls, self-hosted deployments, multi-agent systems (agent swarm) and tasks with a very long context.
- The weights are free to download; self-hosting costs depend on your own infrastructure. Hosted access via Kimi.com, the official API or third parties like Cloudflare Workers AI and OpenRouter each has its own pricing.
- All four are open Chinese frontier lines. Kimi stands out mainly through the size of its context window and the scaling of its agent-swarm feature, while DeepSeek and Qwen offer broader model families and GLM stronger coding benchmarks.