vLLM v0.30.0 published: DeepSeek-V4.1-Flash with MXFP8 FlashMLA KV on Blackwell
vLLM project
Follow-up to yesterday's tag-only sighting: vLLM released v0.30.0 on Sep 22 with 762 commits from 315 contributors (104 new). Headline: DeepSeek-V4.1-Flash support with the entire KV cache kept in MXFP8 via FlashMLA V4.1 on SM100, plus more model additions.
Why it matters
Day-scale support for DeepSeek's new Flash model with a SM100-optimized attention path makes it immediately servable at scale on Blackwell GPUs.
Importance: 2/5
Continuation of yesterday's story: the release is now published