Topic: ggml

70 stories found

Today

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b11094

The ggml/llama.cpp project updated the cpp-httplib library to version 0.57.1, a change signed off by Adrien Gallouët that enhances the project's functionality. This update is significant as it improves compatibility and performance for users of the llamacpp software suite.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11095

The ggml/llama.cpp project has released updates focusing on HMX-optimized GATED_DELTA_NET, marking progress towards faster and more efficient processing. These developments are crucial as they aim to enhance the performance of large language models by optimizing hardware support.

github.com

Yesterday

releases48

ggml/llama.cpp releases: b11090

The ggml/llama.cpp project fixed a compilation error related to CUDA sm_70 tiles by generalizing the tile shape in version b11090, addressing an issue where a recent update mismatched tile definitions. This update is crucial for ensuring compatibility across different GPU architectures.

github.com

Sunday, September 20, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b11065

The ggml/llama.cpp project released an update that tunes the FA parameter for use with the Gemma 4 on Ampere or newer GPUs, enhancing performance. This update is significant as it optimizes machine learning model processing for specific hardware, potentially improving speed and efficiency in applications like natural language processing.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11064

The ggml/llama.cpp project updated its dsv4_hc_pre kernels to support arbitrary hardware contexts (hc), addressing a limitation that previously forced it to use CPU fallbacks. This change is significant as it enhances compatibility and performance across different hardware configurations, particularly for Kimi-K3 which uses varying hc values for banked checkpoints.

github.com

Saturday, September 19, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b11056

The ggml/llama.cpp project released version b11056, which includes a change to enable I32 GET_ROWS (#29116). This update is significant for improving the handling of 32-bit integer operations in the library, potentially enhancing performance and functionality.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11057

The ggml/llama.cpp project added a dedicated Ling 3.0 (Bailing V3) parser to improve chat functionality by pre-opening think blocks in generation prompts, allowing for more seamless tool calls without premature

github.com

Friday, September 18, 2026

releases4 sources🛡️ Verified48

ggml/llama.cpp releases: b11046

The ggml/llama.cpp project has released an update that adds support for the `flash_attn_f32_f16_bin` kernel in OpenCL, enhancing computational efficiency for certain operations. This update is significant as it improves performance in processing tasks related to large language models.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11044

The ggml/llama.cpp project released an update that includes improvements to the hexagon backend, specifically enhancing IM2COL operations for both 1D and padded inputs. These changes are significant as they optimize memory access patterns, potentially improving performance in certain machine learning tasks.

github.com

Thursday, September 17, 2026

releases4 sources🛡️ Verified48

ggml/llama.cpp releases: b11028

The ggml/llama.cpp project released version b11028, which includes a fix for evicting old files. This update is important as it enhances the project's stability and efficiency, particularly for users on Apple Silicon platforms.

Covered by ggml/llama.cpp releases

70 of 70 items shown. Sources: 123 days indexed.