Kimi K2 Linear (Kimi-Linear-48B-A3B)
by Moonshot AI
Research-style efficient model: 48B total / 3B active parameters, 1M context, MIT licence, built around Kimi Delta Attention (KDA). It is a long-context efficiency design, not a member of the K2.5/K2.6/K3 flagship line.
Lineage (then → 2026 now)
Oct 2025: Kimi Linear (this model, paper arXiv 2510.26692) tested KDA at 48B/3B. 2026: the KDA idea was carried into the flagship Kimi K3 (3:1 hybrid linear-to-full attention per the earlier note). No newer standalone “Kimi Linear” checkpoint appears on the Moonshot Hugging Face org page (2026-10-01; its listed models are K3, K2.7 Code, K2.6, K2.5, K2 Instruct, VL and Moonlight), so this remains the reference research release.
Facts (Hugging Face card, arXiv)
- Parameters: 48B total / 3B activated
- Context: 1M tokens
- Licence: MIT
- Attention: KDA (a refined Gated DeltaNet), hybrid with global MLA at a 3:1 KDA-to-MLA ratio
- Claims (vendor/paper): up to 6x decoding throughput and up to 75% smaller KV cache at 1M context; RULER at 128K: 84.3 with 3.98x speedup
- Serving: Transformers, vLLM, SGLang
- Release: October 2025 (arXiv ID 2510.26692; v2 dated 2025-11-01). The Hugging Face card states no day.
Correction
Earlier versions positioned this as an edge/mobile/IoT model. A 48B-parameter model is not edge-class; the pitch is long-context decoding efficiency.
See also
Sources
- https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct (accessed 2026-10-01)
- https://arxiv.org/abs/2510.26692 (via search result; accessed 2026-10-01)
- https://huggingface.co/moonshotai (accessed 2026-10-01)