DeepSeek V4.1 Flash goes GA with open MIT weights and V4 Pro retirement
DeepSeek
DeepSeek moved V4.1 Flash from beta to full release on September 10, publishing MIT-licensed weights and a technical report on Hugging Face. The 552B-parameter multimodal MoE (8B active at prefill, 16B at decode) uses a Causal Encoder-Decoder with FP4 KV caching to cut persistent cache to roughly 1/8 of V4 Flash, supports 1M-token context, and cuts Flash API prices by roughly 11-57%. From September 14 deepseek-v4-pro requests will be routed to the new model.
Why it matters
The first production model on DeepSeek's new architecture family delivers frontier-adjacent open weights at drastically lower memory and price per token, directly undercutting closed API pricing for long-context agent workloads.
Importance: 4/5
Flagship-class open-weights release on a new architecture; official + Reuters and SiliconAngle coverage