releasesvLLM releasesSep 22, 2026
vLLM releases: v0.30.0
Sentiment: neutral
TL;DR
vLLM released version 0.30.0 with significant contributions from many developers, introducing new models like DeepSeek-V4.1-Flash and DeepGEMM Mega-mHC, marking a substantial update in the model's capabilities.
Detailed Summary
vLLM has released version 0.30.0, which includes significant contributions from 315 contributors, including 104 new ones. This update introduces several new models such as DeepSeek-V4.1-Flash, DeepGEMM Mega-mHC, and async Engram prefetching, all designed to enhance performance through advanced memory management techniques. The broader impact is expected to be improved efficiency and speed in large language model processing across various applications.
Key Points
- • New models include DeepSeek-V4.1-Flash
- • Models utilize MXFP8 for KV storage via FlashMLA V4.1 on SM100
- • Async Engram prefetching is also introduced