Topic: quantized
3 stories found
Sunday, July 12, 2026
ggml/llama.cpp releases: b9969
ggml/llama.cpp updates address issues with Vulkan support, particularly fixing problems related to longer prompt sizes and removing unused hardware support, enhancing compatibility and performance for quantized networks on mobile GPUs.
Saturday, July 11, 2026
vLLM releases: v0.25.0
vLLM released version 0.25.0 with 558 commits from 232 contributors, including 64 new ones. This update makes Model Runner V2 the default for all dense models, expanding on previous quantized-model support.
Sunday, July 5, 2026
ggml/llama.cpp releases: b9874
The ggml/llama.cpp project has released an update with commit b9874, introducing a CUDA implementation for concatenating quantized types, which enhances performance for GPU-based operations, particularly in machine learning models. Additionally, the update includes code optimization based on a user's suggestion, improving code efficiency and readability.
🌿 That's all for now. Come back tomorrow.
3 of 3 items shown. Sources: 60 days indexed.