Topic: releases
89 stories found
Yesterday
ggml/llama.cpp releases: b11118
The ggml/llama.cpp project released a new version that includes an update to introduce a direct-mapped DMA cache for better handling of HVX FA mask operations, enhancing performance. This update is significant as it optimizes memory management, particularly relevant for Apple Silicon (arm64) systems on macOS and iOS.
Ollama releases: v0.34.4
Ollama released version 0.34.4, addressing issues like intermittent "model not found" errors and improving structured outputs processing. These updates aim to enhance the stability and efficiency of the server and application functionalities.
Tuesday, September 22, 2026
ggml/llama.cpp releases: b11115
The ggml/llama.cpp project has released a new version including an OpenCL bin kernel for the `kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin`, enhancing binary kernel selection and support for non-MoE dp4a operations, which is crucial for optimizing large language model computations on specific hardware.
ggml/llama.cpp releases: b11094
The ggml/llama.cpp project updated the cpp-httplib library to version 0.57.1, a change signed off by Adrien GallouΓ«t that enhances the project's functionality. This update is significant as it improves compatibility and performance for users of the llamacpp software suite.
ggml/llama.cpp releases: b11114
The ggml/llama.cpp server was updated to fix router eviction race conditions by routing all model loads through a queue, ensuring that no model is evicted prematurely during loading. This update is crucial for improving the stability and reliability of model handling in the server.
Monday, September 21, 2026
NousResearch releases: Hermes Agent v0.21.4 (v2026.9.21)
Hermes Agent version v0.21.4 was released on September 21, 2026, consolidating approximately 1,800 merged pull requests to provide a stable release for various deployment environments, including Docker images and Hermes Cloud services.
ggml/llama.cpp releases: b11090
The ggml/llama.cpp project fixed a compilation error related to CUDA sm_70 tiles by generalizing the tile shape in version b11090, addressing an issue where a recent update mismatched tile definitions. This update is crucial for ensuring compatibility across different GPU architectures.
Sunday, September 20, 2026
ggml/llama.cpp releases: b11065
The ggml/llama.cpp project released an update that tunes the FA parameter for use with the Gemma 4 on Ampere or newer GPUs, enhancing performance. This update is significant as it optimizes machine learning model processing for specific hardware, potentially improving speed and efficiency in applications like natural language processing.
ggml/llama.cpp releases: b11064
The ggml/llama.cpp project updated its dsv4_hc_pre kernels to support arbitrary hardware contexts (hc), addressing a limitation that previously forced it to use CPU fallbacks. This change is significant as it enhances compatibility and performance across different hardware configurations, particularly for Kimi-K3 which uses varying hc values for banked checkpoints.
Saturday, September 19, 2026
Ollama releases: v0.34.3
Ollama released version v0.34.3, which includes updates to the `GET /api/show` endpoint to advertise each model's thinking controls and default settings, enhancing API transparency. This update is significant for users needing detailed control over model behavior.
89 of 89 items shown. Sources: 124 days indexed.