Topic: kv
5 stories found
Today
vLLM releases: v0.30.0
vLLM released version 0.30.0 with significant contributions from many developers, introducing new models like DeepSeek-V4.1-Flash and DeepGEMM Mega-mHC, marking a substantial update in the model's capabilities.
Wednesday, September 16, 2026

Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
AIUC's CEO discusses the company's Series A funding round, focusing on making superintelligent agents more legally accountable by allowing them to be sued. This matters because it addresses potential liability issues in an increasingly complex AI landscape.
ggml/llama.cpp releases: b11009
The ggml/llama.cpp project released a new version addressing issues with split states and granularity for fused QKV gemm operations, crucial for models like Qwen35. This update ensures correct handling of attention layers when using specific configurations, enhancing the model's performance and accuracy.
Friday, September 11, 2026
ggml/llama.cpp releases: b10907
The ggml/llama.cpp project released updates to fix MTP context kv cache allocation issues for specific architectures like deepseek2, glm4moe, and cohere2moe. These changes also include adding inverse architecture gating and comprehensive testing for the MTP layer filter, enhancing model stability and performance.
Wednesday, September 9, 2026
vLLM releases: v0.29.0
vLLM released version 0.29.0, which includes 594 commits from 277 contributors and marks the full rollout of Model Runner V2 as the default for all models, enhancing performance with CUDA graph memory profiling for KV cache auto-sizing. This update is significant as it improves model efficiency and scalability in large language model deployments.
🌿 That's all for now. Come back tomorrow.
5 of 5 items shown. Sources: 123 days indexed.