Topic: quantized

3 stories found

Sunday, July 12, 2026

releases48

ggml/llama.cpp releases: b9969

ggml/llama.cpp updates address issues with Vulkan support, particularly fixing problems related to longer prompt sizes and removing unused hardware support, enhancing compatibility and performance for quantized networks on mobile GPUs.

github.com

Saturday, July 11, 2026

releases42

vLLM releases: v0.25.0

vLLM released version 0.25.0 with 558 commits from 232 contributors, including 64 new ones. This update makes Model Runner V2 the default for all dense models, expanding on previous quantized-model support.

github.com

Sunday, July 5, 2026

releases48

ggml/llama.cpp releases: b9874

The ggml/llama.cpp project has released an update with commit b9874, introducing a CUDA implementation for concatenating quantized types, which enhances performance for GPU-based operations, particularly in machine learning models. Additionally, the update includes code optimization based on a user's suggestion, improving code efficiency and readability.

github.com

🌿 That's all for now. Come back tomorrow.

3 of 3 items shown. Sources: 60 days indexed.