Topic: cache

2 stories found

Yesterday

releases48

ggml/llama.cpp releases: b10255

The ggml/llama.cpp project updated its SYCL oneDNN SDPA to support Q4_0-Q8_0 and FP32 KV caches, extending the handling of non-FP16 key-value pairs. This update is significant as it enhances the flexibility and performance of the model in various computational scenarios.

github.com↗

Friday, July 31, 2026

releases48

ggml/llama.cpp releases: b10213

The ggml/llama.cpp project released version b10213, which includes support for rotated key-value cache quantization. This update enhances the efficiency and performance of the model, making it more suitable for resource-constrained environments like macOS Apple Silicon.

github.com↗

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 77 days indexed.