Kimi K2 Linear

by Moonshot AI

Lightweight edge variant (October 2025). 48B total / 3B active parameters. KDA (Kimi Delta Attention) for memory efficiency. Optimized for cost-sensitive and constrained deployments.

Release

  • Date: October 2025
  • Parameters: 48B total / 3B active
  • Focus: Edge and lightweight deployment

Architecture

  • Attention: KDA (Kimi Delta Attention)
  • Memory footprint: Significantly reduced vs K2
  • Inference: Fast, optimized for latency

Capabilities

  • Reasoning on memory-constrained devices
  • Fast inference for real-time applications
  • Cost-optimized inference at scale
  • Open-source deployment

Use Cases

  • Edge devices (mobile, embedded, IoT)
  • Cost-optimized inference clusters
  • Latency-critical applications
  • Memory-constrained environments

Positioning

  • Entry-level Kimi for resource-constrained scenarios
  • Trades some capability for efficiency
  • Complements K2/K2.5/K2.6 for broader deployment options

See Also