#serving
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs Mind Lab research
- vLLM v0.25.0: Model Runner V2 Default, PagedAttention Retired, Transformers Backend Parity tools
- vLLM Adds Day-0 Support for MiniMax M3 Open Weights with 1M-Context Sparse Attention MiniMax tools
- vLLM v0.24.0: Model Runner V2 Default, Rust Frontend, SM90 FP8 Speedups vLLM tools
- vLLM v0.27.0 Ships Kimi K3 Support and DeepSeek-V4 Optimizations vLLM Project tools
- Modal Launches Auto Endpoints for Production-Grade Open-Model LLM Inference Modal tools
- ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14% Microsoft Research research
- Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler Dharma-AI tools
- Groq closes $350M Series A to build an AI inference cloud Groq industry
- SGLang v0.5.18 lands major model and perf updates tools
- SGLang v0.5.18 SGLang tools
- SGLang v0.5.19 lands beam search, DeepEP v2, and Qwen3.8 support across 786 PRs tools
- vLLM v0.29.0 makes Model Runner V2 the default and adds Mamba prefix caching vLLM tools
- vLLM v0.30.0 tagged on GitHub; release not yet published to PyPI vLLM project tools