Zhipu releases GLM-5.3-Flash, first natively multimodal GLM-5 model

Zhipu AI / Z.ai

Models / LLM official 5 src. ~1 min

Zhipu AI (Z.ai) shipped GLM-5.3-Flash on 2026-08-25 as the first natively multimodal entry in the GLM-5 series: a 320B-total / 18B-active MoE trained on a 30T-token multimodal corpus with hybrid sparse-plus-linear attention and Manifold-Constrained Hyper-Connections. The team claims it outperforms GLM-5.2 across benchmarks and real workloads at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic evaluations (Terminal-Bench 2.1 84.3, DeepSWE 63.4, HLE 55.3). MIT-licensed weights in BF16 and FP8 are on Hugging Face, with serving recipes for SGLang, vLLM, TokenSpeed, and KTransformers.

Why it matters

GLM-5.3-Flash marks the GLM line's first multimodal-from-pretraining model and pairs an aggressive cost cut (1/10 of GLM-5.2) with frontier-tier agentic numbers, putting Zhipu back into direct competition with Claude Opus 4.8 / Qwen3.8 / DeepSeek-V4 in the open-weight MoE tier.

Importance: 3/5

5 sources

Sources