Topic: llama.cpp

74 stories found

Yesterday

releases48

ggml/llama.cpp releases: b10819

The ggml/llama.cpp project released a fix for a memory leak in early return functionality (commit b10819). This update is crucial as it addresses a potential stability issue, enhancing the reliability of the software.

github.com
industry24

Nous Research Adds One-Click Local Model Setup to Hermes Desktop

Nou Research has simplified local model setup in Hermes Desktop to a single click, automatically optimizing and configuring the process for users based on their hardware. This update aims to make AI model deployment more accessible and efficient for end-users.

marktechpost.com

Friday, September 4, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10816

The ggml/llama.cpp project released a new version including tuning updates for the M3 model and additional precision settings, addressing formatting issues. These changes are significant for developers working with the M3 model to optimize performance on metal GPUs.

Covered by ggml/llama.cpp releases
releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10796

A new function, `n_expert_used_max`, was added to the ggml/llama.cpp project. This update allows each layer in the model to have a specific number of experts, enhancing flexibility and potentially improving performance in certain configurations.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: v0.4.0

Version 0.4.0 of llama.cpp was released, adding support for Qwen3.8-Flash-Next and Nemotron-3-Puzzle models, along with several new features like on-demand tensor reading and video input options, making it more versatile for AI language tasks. This update is significant as it enhances the model's capabilities and flexibility, catering to a broader range of applications in natural language processing.

github.com
trending45

Georgi Gerganov on llama.cpp/ggml future after Nvidia acquisition of HuggingFace

Georgi Gerganov discussed the potential impact of Nvidia's acquisition of HuggingFace on his project llama.cpp/ggml, emphasizing concerns about increased computational costs and reduced accessibility in the AI research community. This matters as it could alter the landscape for open-source AI models and affect how researchers and developers access powerful computing resources.

twitter.com

Thursday, September 3, 2026

releases2 sources⚡ Corroborated48

ggml/llama.cpp releases: b10793

The ggml/llama.cpp project released a fix to ensure the entire source code is not rebuilt on each new commit, addressing an issue that could slow development. This update is crucial for improving workflow efficiency in contributing to the llama model's implementation.

Covered by ggml/llama.cpp releases

Wednesday, September 2, 2026

releases54

Ollama releases: v0.33.3

github.com

Tuesday, September 1, 2026

74 of 74 items shown. Sources: 107 days indexed.