Topic: fix
21 stories found
Today
ggml/llama.cpp releases: b11093
The ggml/llama.cpp project released a new version addressing an issue with mask bounds in the flash attention block pre-pass, aiming to improve performance and stability. This update is significant for users relying on macOS Apple Silicon (arm64) as it resolves specific technical problems affecting their systems.
Yesterday
ggml/llama.cpp releases: b11090
The ggml/llama.cpp project fixed a compilation error related to CUDA sm_70 tiles by generalizing the tile shape in version b11090, addressing an issue where a recent update mismatched tile definitions. This update is crucial for ensuring compatibility across different GPU architectures.
Sunday, September 20, 2026
ggml/llama.cpp releases: b11060
The ggml/llama.cpp project released a new version (b11060) that includes fixes to make time-step projection input contiguous and skip unnecessary contiguous copies after normalization, aimed at improving efficiency. These updates are significant for enhancing the performance of the model without compromising accuracy.
Saturday, September 19, 2026
ggml/llama.cpp releases: b11054
The ggml/llama.cpp project updated to enable support for the TOP_K operation on the Hexagon processor, improving row partitioning and optimizing large-row selection. This update is crucial as it enhances the efficiency of operations involving top-k elements in machine learning models running on Hexagon hardware.
Friday, September 18, 2026
ggml/llama.cpp releases: b11042
The ggml/llama.cpp project released a new version including an OpenCL binary kernel for A8 Q6_K non-MoE computations and fixes to layout compatibility, aimed at improving performance and compatibility in GPU-accelerated machine learning tasks. These updates are significant for users looking to optimize their computational resources when running large language models.
Thursday, September 17, 2026
ggml/llama.cpp releases: b11025
The ggml/llama.cpp project released a new version that includes a fix for extending Nemotron MTP support, removing unnecessary declarations. This update is important as it enhances the model's compatibility and efficiency.
Wednesday, September 16, 2026
ggml/llama.cpp releases: b11009
The ggml/llama.cpp project released a new version addressing issues with split states and granularity for fused QKV gemm operations, crucial for models like Qwen35. This update ensures correct handling of attention layers when using specific configurations, enhancing the model's performance and accuracy.
Tuesday, September 15, 2026
Ollama releases: v0.34.2
Ollama released version 0.34.2, which introduces a first-run setup option for signing in or continuing locally, enhances direct access to apps via URL on macOS and Windows, and addresses memory issues. These updates improve user experience and stability across desktop platforms.
Monday, September 14, 2026
EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development
NousResearch releases: Hermes Agent v0.21.3 (v2026.9.14)
Hermes Agent v0.21.3 was released on September 14, 2026, consolidating over 338 merged pull requests to address sign-in issues and ensure stability for various deployment environments.
21 of 21 items shown. Sources: 123 days indexed.