Topic: fix

21 stories found

Today

releases48

ggml/llama.cpp releases: b11093

The ggml/llama.cpp project released a new version addressing an issue with mask bounds in the flash attention block pre-pass, aiming to improve performance and stability. This update is significant for users relying on macOS Apple Silicon (arm64) as it resolves specific technical problems affecting their systems.

github.com↗

Yesterday

releases48

ggml/llama.cpp releases: b11090

The ggml/llama.cpp project fixed a compilation error related to CUDA sm_70 tiles by generalizing the tile shape in version b11090, addressing an issue where a recent update mismatched tile definitions. This update is crucial for ensuring compatibility across different GPU architectures.

github.com↗

Sunday, September 20, 2026

releases48

ggml/llama.cpp releases: b11060

The ggml/llama.cpp project released a new version (b11060) that includes fixes to make time-step projection input contiguous and skip unnecessary contiguous copies after normalization, aimed at improving efficiency. These updates are significant for enhancing the performance of the model without compromising accuracy.

github.com↗

Saturday, September 19, 2026

releases48

ggml/llama.cpp releases: b11054

The ggml/llama.cpp project updated to enable support for the TOP_K operation on the Hexagon processor, improving row partitioning and optimizing large-row selection. This update is crucial as it enhances the efficiency of operations involving top-k elements in machine learning models running on Hexagon hardware.

github.com↗

Friday, September 18, 2026

releases48

ggml/llama.cpp releases: b11042

The ggml/llama.cpp project released a new version including an OpenCL binary kernel for A8 Q6_K non-MoE computations and fixes to layout compatibility, aimed at improving performance and compatibility in GPU-accelerated machine learning tasks. These updates are significant for users looking to optimize their computational resources when running large language models.

github.com↗

Thursday, September 17, 2026

releases48

ggml/llama.cpp releases: b11025

The ggml/llama.cpp project released a new version that includes a fix for extending Nemotron MTP support, removing unnecessary declarations. This update is important as it enhances the model's compatibility and efficiency.

github.com↗

Wednesday, September 16, 2026

releases48

ggml/llama.cpp releases: b11009

The ggml/llama.cpp project released a new version addressing issues with split states and granularity for fused QKV gemm operations, crucial for models like Qwen35. This update ensures correct handling of attention layers when using specific configurations, enhancing the model's performance and accuracy.

github.com↗

Tuesday, September 15, 2026

releases54

Ollama releases: v0.34.2

Ollama released version 0.34.2, which introduces a first-run setup option for signing in or continuing locally, enhances direct access to apps via URL on macOS and Windows, and addresses memory issues. These updates improve user experience and stability across desktop platforms.

github.com↗

Monday, September 14, 2026

releases60

NousResearch releases: Hermes Agent v0.21.3 (v2026.9.14)

Hermes Agent v0.21.3 was released on September 14, 2026, consolidating over 338 merged pull requests to address sign-in issues and ensure stability for various deployment environments.

github.com↗

21 of 21 items shown. Sources: 123 days indexed.