Qwen 3.8 Max

by Alibaba Cloud

Alibaba’s flagship Qwen 3.8 release, shipped open-weight — a 2.4T-parameter sparse MoE model with
~95B active parameters per token, 1M-token context, and native multimodal (text/image/video) input.

Part of the Qwen family, succeeding the Qwen 3.6 series (April 2026).

Specs

  • Architecture: sparse Mixture-of-Experts, 2.4T total parameters, ~95B active per token
  • Context window: up to 1M tokens
  • Modality: text, image, and video input; text output
  • Open weights: released alongside a smaller Qwen 3.8 27B for local use
  • API pricing: positioned well below Western frontier models, especially on output tokens

Timeline

DateEvent
2026-08-03Announced/tested — early coverage claims it beats Claude Opus 4.8 Max in side-by-side outputs
2026-08-04Open weights released; framed against Gemini 3.5 Pro’s ongoing delay
2026-08-14Benchmarked directly against [[./glm-5-dot-3-flash

Benchmarks & reception

Reception is mixed rather than a clean win: strong on coding, front-end/UI generation, 3D/interactive
scenes, multimodal and computer-use-style tasks, and cited on SWE Pro / Terminal-Bench 2.1 / Frontier
SWE as matching or beating some frontier competitors. At the same time, independent testing places it
below Opus and DeepSeek on some overall evaluations, and notes a tendency to “overthink” simple
tasks. The consistent throughline across sources is price: frontier-adjacent capability at a fraction
of closed-model API cost.

Notes

Qwen 3.8 Max lands in the same late-August 2026 wave as GLM-5.3 Flash and DeepSeek V4 — three
open-weight Chinese releases competing directly with each other on local-runnability and cost as much
as with Western closed labs. Coverage repeatedly frames Gemini 3.5 Pro’s delayed release as the
backdrop this and GLM-5.3 landed against.

  • Qwen — parent model line and Qwen 3.6-era coverage
  • GLM-5.3 Flash · DeepSeek (V4 Pro/Flash) — direct open-weight rivals
    from the same window