releasesggml/llama.cpp releasesSep 9, 2026
ggml/llama.cpp releases: b10877
Sentiment: neutral
Key Points
- • CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3
- • CUDA: pick MMQ tile size against ncols_opt set on the host side
- • Assisted-by: Claude Fable 5.1