Topic: quantization

3 stories found

Monday, September 14, 2026

releases54

Ollama releases: v0.34.1

Ollama released version 0.34.1, which marks the full integration of MLX safetensors support in their `ollama create` command and introduces stricter conditions for detecting runaway repeat tokens. These updates enhance model handling on Apple Silicon and improve overall stability.

github.com

Wednesday, September 9, 2026

releases48

ggml/llama.cpp releases: b10876

The ggml/llama.cpp project updated its CUDA implementation to provide more control over quantization options, allowing users to configure specific combinations and enabling runtime fallbacks for unsupported configurations. These changes enhance flexibility and usability, making the library more adaptable to various hardware setups.

github.com

🌿 That's all for now. Come back tomorrow.

3 of 3 items shown. Sources: 123 days indexed.