Topic: kernel
8 stories found
Sunday, September 20, 2026
ggml/llama.cpp releases: b11064
The ggml/llama.cpp project updated its dsv4_hc_pre kernels to support arbitrary hardware contexts (hc), addressing a limitation that previously forced it to use CPU fallbacks. This change is significant as it enhances compatibility and performance across different hardware configurations, particularly for Kimi-K3 which uses varying hc values for banked checkpoints.
Friday, September 18, 2026
ggml/llama.cpp releases: b11046
The ggml/llama.cpp project has released an update that adds support for the `flash_attn_f32_f16_bin` kernel in OpenCL, enhancing computational efficiency for certain operations. This update is significant as it improves performance in processing tasks related to large language models.
ggml/llama.cpp releases: b11044
The ggml/llama.cpp project released an update that includes improvements to the hexagon backend, specifically enhancing IM2COL operations for both 1D and padded inputs. These changes are significant as they optimize memory access patterns, potentially improving performance in certain machine learning tasks.
Wednesday, September 16, 2026
ggml/llama.cpp releases: b11006
The ggml/llama.cpp project released a new version that includes support for K-Quants Q4_K and Q6_K, enhancing quantization techniques crucial for optimizing machine learning models' performance on resource-constrained devices. This update is significant as it improves model efficiency without compromising too much on accuracy, making it more viable for deployment in edge computing scenarios.
Friday, September 11, 2026
ggml/llama.cpp releases: b10908
The ggml/llama.cpp project released a fix for idle threads in specific multiplication kernels used for neural networks with fewer than 1024 elements. This update generalizes a previous row split to additional related kernels, improving thread utilization and potentially enhancing performance.
🌿 That's all for now. Come back tomorrow.
8 of 8 items shown. Sources: 123 days indexed.