Moonshot Kimi LLM Series

by Moonshot AI

Kimi is Moonshot AI’s flagship model family, spanning K1 (Nov 2023) through K3 (July 2026). K3 is frontier-class with 2.8T parameters, 1M context, native multimodal, open weights, and Agent Swarm capabilities. Emphasizes agentic autonomy, long-context reasoning, and cost-efficiency vs Claude/GPT.

Overview

Kimi represents China’s major push into frontier AI. The model family progresses from baseline reasoning (K1) to a comprehensive agentic platform (K3), emphasizing open-source releases, long-horizon task execution, and competitive positioning against Claude and GPT on cost-per-task and agentic performance.

Company Context: Moonshot AI, founded 2023 by Tsinghua grad Yang Zhilin; backed by Alibaba and Chinese internet giants; achieved unicorn status ($1B valuation).

Complete Model Lineage

ModelReleaseParametersKey FeatureStatus
K1Nov 16, 2023UndisclosedBaseline reasoningArchived
K1.5Jan 20, 2025Undisclosedo1-level reasoningLimited
K2July 11, 20251T / 32B active384-expert MoE, open-weightActive
K2 LinearOct 202548B / 3B activeLightweight KDAActive
K2.5Jan 20261T / 32B activeMultimodal + 100 agentsActive
K2.6April 20261T / 32B active300 agents, 4K steps, #1 open-weightActive
K2.7 CodeJune 20261T / 32B active+21.8% coding, -30% tokensActive
K3July 16, 20262.8T / ~50B activeFrontier: 1M context, KDA, always-on reasoning, open weightsActive

K3 (Current Flagship) — Detailed Specifications

Architecture

Parameters & Scaling:

  • Total: 2.8 trillion parameters (896 experts)
  • Active per token: ~50 billion (16 active experts via learned routing)
  • Attention: Hybrid KDA (Kimi Delta Attention) — 3:1 ratio of linear to full attention
  • Efficiency gain: 2.5× better compute-to-intelligence vs K2
  • Quantization: MXFP4 weights, MXFP8 activations (quantization-aware training)
  • Attention Residual (AttnRes): Improved information flow across layers

Capabilities

Context & Speed:

  • Context window: 1 million tokens (~750,000 words)
  • Max output tokens: ~64,000
  • Decoding speed: 6.3× faster than K2 for long-context queries
  • Reasoning mode: Always-on by default (can disable for speed)

Multimodal:

  • Native vision understanding via MoonViT-3D encoder
  • Image + text reasoning at 1M context (unique capability)
  • Chart, diagram, and document analysis

Agent Swarm:

  • 300+ parallel sub-agents
  • 4,000+ coordinated reasoning steps
  • 4.5× faster execution vs single-agent approach
  • Multi-path exploration and consensus-building

Performance & Benchmarks

Agentic/Autonomous Tasks:

  • Arena.ai Frontend Code: Rank #1 (beats Claude/GPT on web dev)
  • Agentic Benchmarks: ~89.5% (5th most capable frontier model)
  • Real-world task wins: 5 of 6 vs Claude Sonnet 5

Coding Performance:

  • SWE-Bench Pro (via K2.6): 58.6%
  • Terminal-Bench 2.0: 66.7%
  • Internal coding benchmarks: +21.8% vs K2.6 (via K2.7 Code)

General Intelligence:

  • Artificial Analysis Index: Rank #3 overall (behind Fable 5 ~60, GPT-5.6 Sol ~59)
  • Math/Reasoning: Frontier-level performance
  • Knowledge: Current through April 2026

Pricing & Availability

API Pricing (July 2026):

  • Input: 0.30 with cache-hit)
  • Output: $15 per million tokens
  • Cost vs alternatives: 50-65% cheaper per completed task than Claude Sonnet 5

Subscription Status:

  • Paused new subscriptions July 17, 2026 (overwhelming demand)
  • Existing subscribers unaffected
  • Capacity reopening batched as infrastructure scales
  • Hosted API: platform.kimi.ai
  • Web Chat: kimi.com

Open Weights:

  • Released July 27, 2026 (Modified MIT license)
  • Available on Hugging Face
  • vLLM support announced for self-hosting
  • Enables local deployment without cloud dependency

K2.6 (Agentic Production Model) — Specifications

Key Features

Agent Swarm 2.0:

  • 300 parallel sub-agents (vs K2.5’s 100)
  • 4,000 coordinated reasoning steps
  • 4.5× faster execution on large-scale search tasks
  • 3-4.5× reduction in critical path steps

Performance:

  • Artificial Analysis (April 2026): Top ranked open-weight model
  • Benchmarks: 58.6% SWE-Bench Pro, 66.7% Terminal-Bench 2.0
  • Multimodal: Native vision via MoonViT-3D
  • Context: 256K tokens

Pricing:

  • Input: $0.60 per million tokens
  • Output: $2.50 per million tokens
  • Cost advantage: 7-10× cheaper than K3 for standard tasks

Use Cases

  • Production agentic workflows (parallel research, coordinated reasoning)
  • Cost-optimized deployments requiring open weights
  • Coding assistance and software engineering
  • Multi-agent systems

K2.5 — Multimodal + Agent Introduction

Release: January 2026

Features:

  • First Kimi with native multimodal (vision + text)
  • Agent Swarm 1.0: up to 100 parallel sub-agents
  • 256K context window
  • Open-source (Modified MIT)
  • Instant & thinking modes

Positioning: Entry point to Kimi’s agentic capabilities; more capable than K2 but less scaled than K2.6.

K2 (Open-Weight Baseline) — Specifications

Release: July 2025

Architecture:

  • 384-expert MoE (1T total, 32B active)
  • Pre-trained on 15.5 trillion tokens
  • Text-only (no vision)
  • Open weights (Modified MIT)

Purpose: Credibility reset — released open-source to restore developer trust after consumer chatbot popularity declined.

Availability: Open-weight; self-hostable; lower pricing tier.

K2 Linear — Edge Deployment Variant

Release: October 2025

Specifications:

  • 48B total, 3B active parameters
  • KDA (Kimi Delta Attention) for efficiency
  • Reduced memory footprint
  • Fast inference

Use Case: Memory-constrained edge devices, cost-optimized inference at scale.

Competitive Positioning

vs. Claude Sonnet 5

MetricWinnerK3Sonnet 5
IntelligenceClaude57/10053/100
Agentic PerformanceKimi89.581.9
Generation SpeedClaude32.1 tok/s76.1 tok/s
Context WindowKimi1M256K
Price/TaskKimi-50-65%baseline
Open WeightsKimiYesNo
Real-world WinsKimi5 of 61 of 6

Summary: Kimi K3 wins on agentic autonomy, context, cost-per-task, and openness. Claude wins on raw speed.

vs. GPT-5.6 Sol

MetricWinner
Frontier reasoningGPT Sol
Agentic codingKimi K3
Cost-per-taskKimi K3
Open deploymentKimi K3
Ecosystem maturityGPT
SpeedGPT

vs. DeepSeek V4 Pro

MetricWinner
Price-per-tokenDeepSeek
Agent SwarmKimi
MultimodalKimi
Reasoning depthComparable
1M contextComparable

Unique Differentiators

1. Agent Swarm (Moonshot-Exclusive)

  • Parallel sub-agent orchestration (300 agents, 4K steps)
  • Enables multi-path reasoning and exploration
  • 4.5× faster execution for large-scale searches
  • Unique capability vs Claude and GPT

2. Open Weights Strategy

  • K3, K2.6, K2.5, K2 all open or promised open
  • Differentiates vs Claude/GPT (closed)
  • Addresses Chinese chip supply constraints (enables broader adoption without vendor lock-in)
  • Enables private/local deployment

3. Long Context + Multimodal Native

  • 1M token context with vision reasoning (K3)
  • Full-codebase analysis, long-document reasoning
  • Vision understanding integrated, not bolted-on

4. Cost-Per-Task Advantage

  • 50-65% cheaper than Claude for complete task execution
  • K2.6 at 2.50 undercuts most competitors
  • DeepSeek cheaper per-token, but Kimi wins per-task due to efficiency

Chinese Market Context

Market Trajectory

  • Aug 2024: Kimi chatbot ranked #3 in China (monthly active users)
  • June 2026: Dropped to #7 after DeepSeek R1 disruption (Jan 2026)
  • Strategic pivot: Yang suspended chatbot marketing; pivoted to enterprise + developer focus
  • K3 release: Aims to reclaim “Chinese AI innovation leadership” narrative

Strategic Differentiation

  1. Hardware Constraints: US export controls limit chip access → open models enable adoption without vendor dependency
  2. Agent-First: Targets enterprise workflow automation and developer productivity
  3. Cost Leadership: Competitive on effective cost-per-task despite DeepSeek’s per-token pricing
  4. Chinese-Language Optimization: Built-in multilingual capability for Chinese market

Competitive Dynamics

  • DeepSeek disrupted pricing expectations (Jan 2026) with ultra-low per-token costs
  • Moonshot responds with Agent Swarm differentiation + performance edge on agentic tasks
  • Both companies emphasize open weights (vs. US closed-model strategies)

When to Use Each Kimi Model

Use K3 for:

  • Frontier-class reasoning and autonomous task execution
  • Long-context document/codebase analysis (1M tokens)
  • Agentic workflows requiring parallel sub-agent coordination
  • Open-weight deployment without cloud dependency
  • Cost-per-task optimization on complex problems

Use K2.6 for:

  • Production agentic systems with moderate budget
  • Parallel research and multi-step reasoning
  • Open-source applications
  • Cost-optimized deployments

Use K2.5 for:

  • Multimodal reasoning with agent support
  • Entry to Kimi’s agentic capabilities
  • Vision + text tasks

Use K2 for:

  • Open-source, cost-sensitive deployments
  • Text-only reasoning

Use K2 Linear for:

  • Edge/embedded deployment
  • Memory-constrained environments

See Also