Topic: qwen35

2 stories found

Thursday, September 17, 2026

releases48

ggml/llama.cpp releases: b11026

The ggml/llama.cpp project released version b11026, which includes a model update to skip gate_up_exps if TENSOR_SKIP is set. This update is crucial for qwen35moe when MTP tensors are fused but not loaded, ensuring compatibility and functionality across different configurations.

github.comโ†—

Wednesday, September 16, 2026

releases48

ggml/llama.cpp releases: b11009

The ggml/llama.cpp project released a new version addressing issues with split states and granularity for fused QKV gemm operations, crucial for models like Qwen35. This update ensures correct handling of attention layers when using specific configurations, enhancing the model's performance and accuracy.

github.comโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 123 days indexed.