Topic: llama.cpp
74 stories found
Yesterday
ggml/llama.cpp releases: b10819
The ggml/llama.cpp project released a fix for a memory leak in early return functionality (commit b10819). This update is crucial as it addresses a potential stability issue, enhancing the reliability of the software.
Nous Research Adds One-Click Local Model Setup to Hermes Desktop
Nou Research has simplified local model setup in Hermes Desktop to a single click, automatically optimizing and configuring the process for users based on their hardware. This update aims to make AI model deployment more accessible and efficient for end-users.
Friday, September 4, 2026
ggml/llama.cpp releases: b10816
The ggml/llama.cpp project released a new version including tuning updates for the M3 model and additional precision settings, addressing formatting issues. These changes are significant for developers working with the M3 model to optimize performance on metal GPUs.
ggml/llama.cpp releases: b10796
A new function, `n_expert_used_max`, was added to the ggml/llama.cpp project. This update allows each layer in the model to have a specific number of experts, enhancing flexibility and potentially improving performance in certain configurations.
ggml/llama.cpp releases: v0.4.0
Version 0.4.0 of llama.cpp was released, adding support for Qwen3.8-Flash-Next and Nemotron-3-Puzzle models, along with several new features like on-demand tensor reading and video input options, making it more versatile for AI language tasks. This update is significant as it enhances the model's capabilities and flexibility, catering to a broader range of applications in natural language processing.

Georgi Gerganov on llama.cpp/ggml future after Nvidia acquisition of HuggingFace
Georgi Gerganov discussed the potential impact of Nvidia's acquisition of HuggingFace on his project llama.cpp/ggml, emphasizing concerns about increased computational costs and reduced accessibility in the AI research community. This matters as it could alter the landscape for open-source AI models and affect how researchers and developers access powerful computing resources.
Thursday, September 3, 2026
ggml/llama.cpp releases: b10793
The ggml/llama.cpp project released a fix to ensure the entire source code is not rebuilt on each new commit, addressing an issue that could slow development. This update is crucial for improving workflow efficiency in contributing to the llama model's implementation.
Wednesday, September 2, 2026
Tuesday, September 1, 2026
74 of 74 items shown. Sources: 107 days indexed.