vLLM 0.28.0 landed with 584 commits, bringing stack-wide Kimi-K3 support, end-to-end DeepSeek-V4 sparse MLA with DFlash2/DSpark speculative decoding, Model Runner V2 disaggregation, and tiered disk KV cache.

Key Takeaways

  • Adaptive speculative token budgets slash end-to-end TTFT by 55–65% with 1.5–3x faster sequence-parallel kernels;
  • End-to-end DeepSeek-V4 sparse MLA and Kimi-K3 shared-expert sharding save ~17 GiB VRAM per GPU;
  • Tiered KV cache introduces a disk tier, doubles max batched tokens to 16,384, and supports diverse hardware backends.