vLLM v0.27.0 Ships Kimi K3 Support and DeepSeek-V4 Optimizations

vLLM Project

Tools official 1 src. ~1 min

vLLM 0.27.0 (561 commits, 242 contributors) adds full-stack Kimi K3 support, new Qwen3.5 and K-EXAONE-2.0-750B-A37B model support, DeepSeek-V4 routing-kernel optimizations, a PyTorch 2.13.0 upgrade, FlashAttention 4 on SM100 with FP8 KV cache, Model Runner V2 for embedding/classification workloads, and early NVIDIA Rubin and ROCm gfx1250 hardware support.

Why it matters

A major release expanding vLLM's model coverage and inference performance for large-scale self-hosted serving.

Importance: 3/5

Notable release: major version bump with broad new model and hardware support, single official source.

Sources

official vLLM releases