Topic: prefill

2 stories found

Friday, September 18, 2026

releases48

ggml/llama.cpp releases: b11046

The ggml/llama.cpp project has released an update that adds support for the `flash_attn_f32_f16_bin` kernel in OpenCL, enhancing computational efficiency for certain operations. This update is significant as it improves performance in processing tasks related to large language models.

github.com

Thursday, September 10, 2026

releases48

ggml/llama.cpp releases: b10900

The ggml/llama.cpp project released version b10900, which includes a Vulkan update enabling topk_moe fusion for prefill, enhancing performance. This update is significant as it improves the efficiency of the model's processing capabilities on compatible hardware.

github.com

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 123 days indexed.