Topic: async

3 stories found

Today

releases42

vLLM releases: v0.30.0

vLLM released version 0.30.0 with significant contributions from many developers, introducing new models like DeepSeek-V4.1-Flash and DeepGEMM Mega-mHC, marking a substantial update in the model's capabilities.

github.com

Tuesday, September 15, 2026

releases48

ggml/llama.cpp releases: b10985

The ggml/llama.cpp project updated to hash-cache only weights for transfers above a certain threshold, allowing the rpc-server to serve cached files from a weight-specific cache, enhancing efficiency in model loading and management. This update is significant as it optimizes the workflow by reducing redundant data transfers and improving the performance of the model deployment process.

github.com

Thursday, September 10, 2026

releases48

ggml/llama.cpp releases: b10901

The ggml/llama.cpp project released version b10901, which includes a Vulkan update allowing CPU writes in asynchronous tensor copies when the context is idle. This update aims to optimize performance by leveraging CPU resources more efficiently during idle GPU contexts.

github.com

🌿 That's all for now. Come back tomorrow.

3 of 3 items shown. Sources: 123 days indexed.