Topic: releases
83 stories found
Yesterday
Ollama releases: v0.34.0
Ollama released version 0.34.0, allowing users to run their own models in ChatGPT Desktop and improving structured output performance on Apple Silicon, enhancing flexibility and functionality for model users.
NVIDIA Releases Personal AI Router (PAIR): An Open Source Virtual Inference Router that Distributes Local AI Requests Across RTX, DGX Spark, and Mac Nodes
NVIDIA released Personal AI Router (PAIR), an open-source tool that distributes AI requests across various devices like RTX, DGX Spark, and Mac nodes within a home network. This allows for efficient use of existing hardware without requiring changes to agent harnesses, potentially lowering the barrier for local AI development and deployment.
Friday, September 4, 2026
ggml/llama.cpp releases: b10816
The ggml/llama.cpp project released a new version including tuning updates for the M3 model and additional precision settings, addressing formatting issues. These changes are significant for developers working with the M3 model to optimize performance on metal GPUs.
ggml/llama.cpp releases: b10796
A new function, `n_expert_used_max`, was added to the ggml/llama.cpp project. This update allows each layer in the model to have a specific number of experts, enhancing flexibility and potentially improving performance in certain configurations.
ggml/llama.cpp releases: v0.4.0
Version 0.4.0 of llama.cpp was released, adding support for Qwen3.8-Flash-Next and Nemotron-3-Puzzle models, along with several new features like on-demand tensor reading and video input options, making it more versatile for AI language tasks. This update is significant as it enhances the model's capabilities and flexibility, catering to a broader range of applications in natural language processing.
Thursday, September 3, 2026
ggml/llama.cpp releases: b10793
The ggml/llama.cpp project released a fix to ensure the entire source code is not rebuilt on each new commit, addressing an issue that could slow development. This update is crucial for improving workflow efficiency in contributing to the llama model's implementation.
Wednesday, September 2, 2026
Tuesday, September 1, 2026
Monday, August 31, 2026
83 of 83 items shown. Sources: 107 days indexed.