DeepSeek releases experimental multimodal V4-Flash-Vision-Exp

DeepSeek

Models / LLM official + media 4 src. ~1 min

DeepSeek shipped V4-Flash-Vision-Exp on its API platform as its first multimodal entry, matching V4-Flash on text tasks (agents, reasoning, world knowledge) while adding image understanding billed at up to 384 tokens per image. The vision model lands close to Opus-4.8 on DeepSeek's multimodal-agent benchmark table and ships with a free Files API for image reuse and DeepSeek Harness 0.1.1 day-one support.

Why it matters

Marks DeepSeek's move beyond text into the multimodal race at Flash-tier pricing, with day-one coverage in Chinese financial press (Caixin) framing it as the start of multimodal competition alongside Moonshot and Qwen.

Importance: 3/5

Frontier Chinese lab first-multimodal release with 4 independent sources (2 official + 2 primary-media)

Sources