Topic: cohere2moe

1 stories found

Friday, September 11, 2026

releases48

ggml/llama.cpp releases: b10907

The ggml/llama.cpp project released updates to fix MTP context kv cache allocation issues for specific architectures like deepseek2, glm4moe, and cohere2moe. These changes also include adding inverse architecture gating and comprehensive testing for the MTP layer filter, enhancing model stability and performance.

github.com

🌿 That's all for now. Come back tomorrow.

1 of 1 items shown. Sources: 123 days indexed.