← Back to News
releasesvLLM releasesSep 9, 2026

vLLM releases: v0.29.0

Read original ↗

Sentiment: neutral

TL;DR

vLLM released version 0.29.0, which includes 594 commits from 277 contributors and marks the full rollout of Model Runner V2 as the default for all models, enhancing performance with CUDA graph memory profiling for KV cache auto-sizing. This update is significant as it improves model efficiency and scalability in large language model deployments.

Detailed Summary

vLLM v0.29.0, a major update to the vLLM model runner, includes 594 commits from 277 contributors, with 91 being new participants. Notably, Model Runner V2 (MRV2) is now the default for all models, finishing a rollout that started with pooling models and adding CUDA graph memory profiling for KV cache auto-sizing. This release significantly enhances the model's performance and efficiency.

Key Points

  • • Model Runner V2 is now the default for all models.
  • • MRV2 gained CUDA graph memory profiling for KV cache auto-sizing.

Source: vLLM releases

Score: 42