Superseded (verified 2026-09-30)
Kimi K2.5
by Moonshot AI
Open-weight multimodal MoE with “agent swarm” execution (January 2026). 1T total / 32B active parameters, 256K context, Modified MIT licence.
Specifications (Hugging Face model card)
- Release: 2026-01-29 (card)
- Parameters: 1T total / 32B activated; 384 experts, 8 selected per token plus 1 shared; 61 layers
- Context: 256K tokens
- Vision: MoonViT encoder (400M parameters); image and video input. (An earlier version of this note said “MoonViT-3D”: not what the card says.)
- Training: continual pretraining on ~15T mixed visual and text tokens (earlier note said 15.5T: corrected to the card’s figure)
- Licence: Modified MIT
- Modes: instant and thinking
Agent swarm
The card describes a self-directed, coordinated swarm-like execution scheme that splits tasks into parallel sub-tasks run by dynamically created agents. Moonshot’s K2.6 post gives K2.5’s scale as 100 sub-agents and 1,500 coordinated steps.
Benchmarks (thinking mode, vendor-reported)
AIME 2025 96.1; GPQA-Diamond 87.6; MMLU-Pro 87.1; SWE-Bench Verified 76.8; VideoMMU 86.6.
See also
Sources
- https://huggingface.co/moonshotai/Kimi-K2.5 (accessed 2026-09-30)
- https://www.kimi.ai/blog/kimi-k2-6 (K2.5 swarm figures, accessed 2026-09-30)