Moonshot Kimi LLM Series
by Moonshot AI
Kimi is Moonshot AI’s flagship model family, spanning K1 (Nov 2023) through K3 (July 2026). K3 is frontier-class with 2.8T parameters, 1M context, native multimodal, open weights, and Agent Swarm capabilities. Emphasizes agentic autonomy, long-context reasoning, and cost-efficiency vs Claude/GPT.
Overview
Kimi represents China’s major push into frontier AI. The model family progresses from baseline reasoning (K1) to a comprehensive agentic platform (K3), emphasizing open-source releases, long-horizon task execution, and competitive positioning against Claude and GPT on cost-per-task and agentic performance.
Company Context: Moonshot AI, founded 2023 by Tsinghua grad Yang Zhilin; backed by Alibaba and Chinese internet giants; achieved unicorn status ($1B valuation).
Complete Model Lineage
| Model | Release | Parameters | Key Feature | Status |
|---|---|---|---|---|
| K1 | Nov 16, 2023 | Undisclosed | Baseline reasoning | Archived |
| K1.5 | Jan 20, 2025 | Undisclosed | o1-level reasoning | Limited |
| K2 | July 11, 2025 | 1T / 32B active | 384-expert MoE, open-weight | Active |
| K2 Linear | Oct 2025 | 48B / 3B active | Lightweight KDA | Active |
| K2.5 | Jan 2026 | 1T / 32B active | Multimodal + 100 agents | Active |
| K2.6 | April 2026 | 1T / 32B active | 300 agents, 4K steps, #1 open-weight | Active |
| K2.7 Code | June 2026 | 1T / 32B active | +21.8% coding, -30% tokens | Active |
| K3 | July 16, 2026 | 2.8T / ~50B active | Frontier: 1M context, KDA, always-on reasoning, open weights | Active |
K3 (Current Flagship) — Detailed Specifications
Architecture
Parameters & Scaling:
- Total: 2.8 trillion parameters (896 experts)
- Active per token: ~50 billion (16 active experts via learned routing)
- Attention: Hybrid KDA (Kimi Delta Attention) — 3:1 ratio of linear to full attention
- Efficiency gain: 2.5× better compute-to-intelligence vs K2
- Quantization: MXFP4 weights, MXFP8 activations (quantization-aware training)
- Attention Residual (AttnRes): Improved information flow across layers
Capabilities
Context & Speed:
- Context window: 1 million tokens (~750,000 words)
- Max output tokens: ~64,000
- Decoding speed: 6.3× faster than K2 for long-context queries
- Reasoning mode: Always-on by default (can disable for speed)
Multimodal:
- Native vision understanding via MoonViT-3D encoder
- Image + text reasoning at 1M context (unique capability)
- Chart, diagram, and document analysis
Agent Swarm:
- 300+ parallel sub-agents
- 4,000+ coordinated reasoning steps
- 4.5× faster execution vs single-agent approach
- Multi-path exploration and consensus-building
Performance & Benchmarks
Agentic/Autonomous Tasks:
- Arena.ai Frontend Code: Rank #1 (beats Claude/GPT on web dev)
- Agentic Benchmarks: ~89.5% (5th most capable frontier model)
- Real-world task wins: 5 of 6 vs Claude Sonnet 5
Coding Performance:
- SWE-Bench Pro (via K2.6): 58.6%
- Terminal-Bench 2.0: 66.7%
- Internal coding benchmarks: +21.8% vs K2.6 (via K2.7 Code)
General Intelligence:
- Artificial Analysis Index: Rank #3 overall (behind Fable 5 ~60, GPT-5.6 Sol ~59)
- Math/Reasoning: Frontier-level performance
- Knowledge: Current through April 2026
Pricing & Availability
API Pricing (July 2026):
- Input: 0.30 with cache-hit)
- Output: $15 per million tokens
- Cost vs alternatives: 50-65% cheaper per completed task than Claude Sonnet 5
Subscription Status:
- Paused new subscriptions July 17, 2026 (overwhelming demand)
- Existing subscribers unaffected
- Capacity reopening batched as infrastructure scales
- Hosted API:
platform.kimi.ai - Web Chat:
kimi.com
Open Weights:
- Released July 27, 2026 (Modified MIT license)
- Available on Hugging Face
- vLLM support announced for self-hosting
- Enables local deployment without cloud dependency
K2.6 (Agentic Production Model) — Specifications
Key Features
Agent Swarm 2.0:
- 300 parallel sub-agents (vs K2.5’s 100)
- 4,000 coordinated reasoning steps
- 4.5× faster execution on large-scale search tasks
- 3-4.5× reduction in critical path steps
Performance:
- Artificial Analysis (April 2026): Top ranked open-weight model
- Benchmarks: 58.6% SWE-Bench Pro, 66.7% Terminal-Bench 2.0
- Multimodal: Native vision via MoonViT-3D
- Context: 256K tokens
Pricing:
- Input: $0.60 per million tokens
- Output: $2.50 per million tokens
- Cost advantage: 7-10× cheaper than K3 for standard tasks
Use Cases
- Production agentic workflows (parallel research, coordinated reasoning)
- Cost-optimized deployments requiring open weights
- Coding assistance and software engineering
- Multi-agent systems
K2.5 — Multimodal + Agent Introduction
Release: January 2026
Features:
- First Kimi with native multimodal (vision + text)
- Agent Swarm 1.0: up to 100 parallel sub-agents
- 256K context window
- Open-source (Modified MIT)
- Instant & thinking modes
Positioning: Entry point to Kimi’s agentic capabilities; more capable than K2 but less scaled than K2.6.
K2 (Open-Weight Baseline) — Specifications
Release: July 2025
Architecture:
- 384-expert MoE (1T total, 32B active)
- Pre-trained on 15.5 trillion tokens
- Text-only (no vision)
- Open weights (Modified MIT)
Purpose: Credibility reset — released open-source to restore developer trust after consumer chatbot popularity declined.
Availability: Open-weight; self-hostable; lower pricing tier.
K2 Linear — Edge Deployment Variant
Release: October 2025
Specifications:
- 48B total, 3B active parameters
- KDA (Kimi Delta Attention) for efficiency
- Reduced memory footprint
- Fast inference
Use Case: Memory-constrained edge devices, cost-optimized inference at scale.
Competitive Positioning
vs. Claude Sonnet 5
| Metric | Winner | K3 | Sonnet 5 |
|---|---|---|---|
| Intelligence | Claude | 57/100 | 53/100 |
| Agentic Performance | Kimi | 89.5 | 81.9 |
| Generation Speed | Claude | 32.1 tok/s | 76.1 tok/s |
| Context Window | Kimi | 1M | 256K |
| Price/Task | Kimi | -50-65% | baseline |
| Open Weights | Kimi | Yes | No |
| Real-world Wins | Kimi | 5 of 6 | 1 of 6 |
Summary: Kimi K3 wins on agentic autonomy, context, cost-per-task, and openness. Claude wins on raw speed.
vs. GPT-5.6 Sol
| Metric | Winner |
|---|---|
| Frontier reasoning | GPT Sol |
| Agentic coding | Kimi K3 |
| Cost-per-task | Kimi K3 |
| Open deployment | Kimi K3 |
| Ecosystem maturity | GPT |
| Speed | GPT |
vs. DeepSeek V4 Pro
| Metric | Winner |
|---|---|
| Price-per-token | DeepSeek |
| Agent Swarm | Kimi |
| Multimodal | Kimi |
| Reasoning depth | Comparable |
| 1M context | Comparable |
Unique Differentiators
1. Agent Swarm (Moonshot-Exclusive)
- Parallel sub-agent orchestration (300 agents, 4K steps)
- Enables multi-path reasoning and exploration
- 4.5× faster execution for large-scale searches
- Unique capability vs Claude and GPT
2. Open Weights Strategy
- K3, K2.6, K2.5, K2 all open or promised open
- Differentiates vs Claude/GPT (closed)
- Addresses Chinese chip supply constraints (enables broader adoption without vendor lock-in)
- Enables private/local deployment
3. Long Context + Multimodal Native
- 1M token context with vision reasoning (K3)
- Full-codebase analysis, long-document reasoning
- Vision understanding integrated, not bolted-on
4. Cost-Per-Task Advantage
- 50-65% cheaper than Claude for complete task execution
- K2.6 at 2.50 undercuts most competitors
- DeepSeek cheaper per-token, but Kimi wins per-task due to efficiency
Chinese Market Context
Market Trajectory
- Aug 2024: Kimi chatbot ranked #3 in China (monthly active users)
- June 2026: Dropped to #7 after DeepSeek R1 disruption (Jan 2026)
- Strategic pivot: Yang suspended chatbot marketing; pivoted to enterprise + developer focus
- K3 release: Aims to reclaim “Chinese AI innovation leadership” narrative
Strategic Differentiation
- Hardware Constraints: US export controls limit chip access → open models enable adoption without vendor dependency
- Agent-First: Targets enterprise workflow automation and developer productivity
- Cost Leadership: Competitive on effective cost-per-task despite DeepSeek’s per-token pricing
- Chinese-Language Optimization: Built-in multilingual capability for Chinese market
Competitive Dynamics
- DeepSeek disrupted pricing expectations (Jan 2026) with ultra-low per-token costs
- Moonshot responds with Agent Swarm differentiation + performance edge on agentic tasks
- Both companies emphasize open weights (vs. US closed-model strategies)
When to Use Each Kimi Model
Use K3 for:
- Frontier-class reasoning and autonomous task execution
- Long-context document/codebase analysis (1M tokens)
- Agentic workflows requiring parallel sub-agent coordination
- Open-weight deployment without cloud dependency
- Cost-per-task optimization on complex problems
Use K2.6 for:
- Production agentic systems with moderate budget
- Parallel research and multi-step reasoning
- Open-source applications
- Cost-optimized deployments
Use K2.5 for:
- Multimodal reasoning with agent support
- Entry to Kimi’s agentic capabilities
- Vision + text tasks
Use K2 for:
- Open-source, cost-sensitive deployments
- Text-only reasoning
Use K2 Linear for:
- Edge/embedded deployment
- Memory-constrained environments