Topic: releases

89 stories found

Yesterday

releases2 sources⚑ Corroborated48

ggml/llama.cpp releases: b11118

The ggml/llama.cpp project released a new version that includes an update to introduce a direct-mapped DMA cache for better handling of HVX FA mask operations, enhancing performance. This update is significant as it optimizes memory management, particularly relevant for Apple Silicon (arm64) systems on macOS and iOS.

Covered by ggml/llama.cpp releases
releases54

Ollama releases: v0.34.4

Ollama released version 0.34.4, addressing issues like intermittent "model not found" errors and improving structured outputs processing. These updates aim to enhance the stability and efficiency of the server and application functionalities.

github.com↗

Tuesday, September 22, 2026

releases2 sources⚑ Corroborated48

ggml/llama.cpp releases: b11115

The ggml/llama.cpp project has released a new version including an OpenCL bin kernel for the `kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin`, enhancing binary kernel selection and support for non-MoE dp4a operations, which is crucial for optimizing large language model computations on specific hardware.

Covered by ggml/llama.cpp releases, vLLM releases
releases2 sources⚑ Corroborated48

ggml/llama.cpp releases: b11094

The ggml/llama.cpp project updated the cpp-httplib library to version 0.57.1, a change signed off by Adrien GallouΓ«t that enhances the project's functionality. This update is significant as it improves compatibility and performance for users of the llamacpp software suite.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11114

The ggml/llama.cpp server was updated to fix router eviction race conditions by routing all model loads through a queue, ensuring that no model is evicted prematurely during loading. This update is crucial for improving the stability and reliability of model handling in the server.

github.com↗

Monday, September 21, 2026

releases60

NousResearch releases: Hermes Agent v0.21.4 (v2026.9.21)

Hermes Agent version v0.21.4 was released on September 21, 2026, consolidating approximately 1,800 merged pull requests to provide a stable release for various deployment environments, including Docker images and Hermes Cloud services.

github.com↗
releases48

ggml/llama.cpp releases: b11090

The ggml/llama.cpp project fixed a compilation error related to CUDA sm_70 tiles by generalizing the tile shape in version b11090, addressing an issue where a recent update mismatched tile definitions. This update is crucial for ensuring compatibility across different GPU architectures.

github.com↗

Sunday, September 20, 2026

releases2 sources⚑ Corroborated48

ggml/llama.cpp releases: b11065

The ggml/llama.cpp project released an update that tunes the FA parameter for use with the Gemma 4 on Ampere or newer GPUs, enhancing performance. This update is significant as it optimizes machine learning model processing for specific hardware, potentially improving speed and efficiency in applications like natural language processing.

Covered by ggml/llama.cpp releases
releases48

ggml/llama.cpp releases: b11064

The ggml/llama.cpp project updated its dsv4_hc_pre kernels to support arbitrary hardware contexts (hc), addressing a limitation that previously forced it to use CPU fallbacks. This change is significant as it enhances compatibility and performance across different hardware configurations, particularly for Kimi-K3 which uses varying hc values for banked checkpoints.

github.com↗

Saturday, September 19, 2026

releases2 sources⚑ Corroborated54

Ollama releases: v0.34.3

Ollama released version v0.34.3, which includes updates to the `GET /api/show` endpoint to advertise each model's thinking controls and default settings, enhancing API transparency. This update is significant for users needing detailed control over model behavior.

Covered by Ollama releases, ggml/llama.cpp releases

89 of 89 items shown. Sources: 124 days indexed.