Topic: mtp

6 stories found

Thursday, September 17, 2026

releases2 sourcesโšก Corroborated48

ggml/llama.cpp releases: b11026

The ggml/llama.cpp project released version b11026, which includes a model update to skip gate_up_exps if TENSOR_SKIP is set. This update is crucial for qwen35moe when MTP tensors are fused but not loaded, ensuring compatibility and functionality across different configurations.

Covered by ggml/llama.cpp releases

Wednesday, September 16, 2026

releases48

ggml/llama.cpp releases: b11007

The ggml/llama.cpp project released a new version that enables CUDA graph usage for MTP, improving performance. This update is significant as it enhances the efficiency of the software, particularly relevant for users requiring high computational power.

github.comโ†—

Tuesday, September 15, 2026

releases48

ggml/llama.cpp releases: b10988

The ggml/llama.cpp project has released updates to optimize speculative decoding and memory management for speculative tensor operations, focusing on OpenCL implementations. These changes are crucial for improving the efficiency and performance of large language models during inference.

github.comโ†—

Sunday, September 13, 2026

Friday, September 11, 2026

releases48

ggml/llama.cpp releases: b10907

The ggml/llama.cpp project released updates to fix MTP context kv cache allocation issues for specific architectures like deepseek2, glm4moe, and cohere2moe. These changes also include adding inverse architecture gating and comprehensive testing for the MTP layer filter, enhancing model stability and performance.

github.comโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

6 of 6 items shown. Sources: 124 days indexed.