Qwen3.8-Flash-Next released as experimental Qwen4 architecture preview

Alibaba (Qwen Team)

Models / LLM official 5 src. ~1 min

Alibaba's Qwen Team open-sourced Qwen3.8-Flash-Next on 2026-08-24 (HF) / 2026-08-26 (GitHub) as an experimental preview of the architecture underpinning Qwen4: a 125B-total / 6B-active MoE with a 51B-parameter n-gram embedding table, 4B MTP head, hybrid Gated DeltaNet + Qwen Sparse Attention operating at micro-block granularity, and 262K native / 1M-token YaRN-extended context. Reported benchmarks include DeepSWE 1.1 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, LiveCodeBench v6 91.9, GPQA Diamond 91.7, and AndroidWorld 84.5, with weights shipped in BF16 and FP8 under qwen-community-1.0 and serving recipes for SGLang, vLLM, and TokenSpeed.

Why it matters

It is the first public signal that Qwen4 will lean on hybrid attention + n-gram embedding for long-context efficiency rather than pure dense or pure MoE scaling, and it lets the open-weight community benchmark the Qwen4 stack before the flagship ships.

Importance: 3/5

5 sources

Sources