DeepSeek open-sources V4-Flash-Vision-Exp, its first multimodal V4 model, under MIT

DeepSeek

Models / LLM official 2 src. ~1 min

On August 31, 2026 DeepSeek published weights for DeepSeek-V4-Flash-Vision-Exp on Hugging Face (repo created 2026-08-31T06:16Z), its first experimental multimodal model in the V4 family. The 305B-parameter MoE model (MIT license, FP8/BF16) extends DeepSeek-V4-Flash with a vision encoder and aligner, DFlash attention, and a DSpark speculative-decoding path for SGLang. Self-reported scores: Terminal Bench 2.1 at 83.9 and ApexBench Pass@1 at 36.5 (vs 26.2 for the text-only V4-Flash-0731). The model had been API-only since August 21; the weights drop makes it the largest open MIT-licensed multimodal model from a Chinese lab to date.

Why it matters

Moves a frontier-scale open multimodal model into the fully-open MIT tier just days after its API debut, and signals DeepSeek is extending the V4 line from text into vision agentic work.

Importance: 4/5

Flagship-scale open-weights multimodal release, official sources

Sources