GLM-5.3 Flash
by Zhipu AI (Z.ai)
The first natively multimodal model in the GLM-5 line — a 320B MoE (18B active) released under an
MIT license, revealed August 26, 2026 as the identity behind the stealth-tested “Ox Alpha.”
Part of the GLM-5 series; the Flash variant of GLM-5.3.
Specs
- Architecture: MoE, 320B total parameters, 18B active per token
- Attention: hybrid attention with MHC-related scaling improvements for long context
- Context window: up to 1M tokens
- Modality: first natively multimodal GLM-5 model (vision input confirmed — image-to-dashboard demo)
- License: MIT, full weights on Hugging Face
- Local hardware: runs on CPU + RAM alone; recurring figure across coverage is ~$800 of hardware
for frontier-adjacent local capability (paired with DeepSeek V4 Flash in that comparison)
Timeline
| Date | Event |
|---|---|
| 2026-08-26 | Revealed as GLM-5.3 Flash, the stealth-tested “Ox Alpha” — Z.ai confirms identity, publishes specs |
| 2026-08-26 | Stealth preview reportedly served 44T tokens to 500,000+ users across 13M sessions before reveal |
| 2026-08-27 | Full weights released on Hugging Face (MIT); local hardware requirements tested |
| 2026-08-27 → 08-28 | Benchmarked head-to-head against [[./qwen-3-dot-8-max |
Reception
Reviewers describe it as one of the best models for local use released in 2026: strong vision and
coding performance at unusually low API cost, and in at least one direct comparison judged to handle a
prototyping task better than Claude Opus 5 with a more pixel-perfect result. Creative-coding testing
(browser-OS recreation, game demos, a Blender/Godot build) showed clear wins alongside a few weaknesses.
One retest after the stealth-to-named transition scored slightly lower than during the anonymous
preview period — attributed to preview-vs-GA variance rather than a regression.
Notes
GLM-5.3 Flash and Qwen 3.8 Max landed in the same late-August 2026 window as the
two headline open-weight Chinese releases, both explicitly framed as validating domestic inference
hardware and serving stacks, not just model quality.
Related
- GLM-5 series · GLM-5.3 · GLM-5.2
- Qwen 3.8 Max — direct rival, same release window
- DeepSeek (V4 Flash) — paired in local-hardware coverage