releasesggml/llama.cpp releases2026-07-12
ggml/llama.cpp releases: b9969
Sentiment: neutral
TL;DR
ggml/llama.cpp updates address issues with Vulkan support, particularly fixing problems related to longer prompt sizes and removing unused hardware support, enhancing compatibility and performance for quantized networks on mobile GPUs.
Detailed Summary
ggml/llama.cpp released a fix addressing issues with the llama-cli breaking for longer prompt sizes, particularly affecting q4_0 quantized networks due to insufficient shared memory on Adreno devices. The update also removed an unused Adreno device configuration. This patch improves compatibility and performance across various hardware platforms, enhancing user experience in handling larger prompts.
Key Points
- • [Vulkan] Fixes llama-cli breaking over longer prompt sizes
- • Removed unused Adreno device
Source: ggml/llama.cpp releases
Score: 48