Topic: tensor
6 stories found
Yesterday
ggml/llama.cpp releases: b11081
The ggml/llama.cpp project released a new version with updates to make tensor data standard deviation configurable and added more examples to the documentation, enhancing flexibility and usability for developers. These changes are significant as they improve the toolkit's adaptability and user-friendliness in handling large language model architectures.
Thursday, September 17, 2026
ggml/llama.cpp releases: b11026
The ggml/llama.cpp project released version b11026, which includes a model update to skip gate_up_exps if TENSOR_SKIP is set. This update is crucial for qwen35moe when MTP tensors are fused but not loaded, ensuring compatibility and functionality across different configurations.
Tuesday, September 15, 2026
ggml/llama.cpp releases: b10985
The ggml/llama.cpp project updated to hash-cache only weights for transfers above a certain threshold, allowing the rpc-server to serve cached files from a weight-specific cache, enhancing efficiency in model loading and management. This update is significant as it optimizes the workflow by reducing redundant data transfers and improving the performance of the model deployment process.
Monday, September 14, 2026
Ollama releases: v0.34.1
Ollama released version 0.34.1, which marks the full integration of MLX safetensors support in their `ollama create` command and introduces stricter conditions for detecting runaway repeat tokens. These updates enhance model handling on Apple Silicon and improve overall stability.
Thursday, September 10, 2026
ggml/llama.cpp releases: b10901
The ggml/llama.cpp project released version b10901, which includes a Vulkan update allowing CPU writes in asynchronous tensor copies when the context is idle. This update aims to optimize performance by leveraging CPU resources more efficiently during idle GPU contexts.
Tuesday, September 8, 2026
ggml/llama.cpp releases: b10867
The ggml/llama.cpp project updated its codebase to disable lazy tensor loading by default on integrated GPUs (iGPUs) to address performance issues, impacting users of these less powerful graphics cards. This change is crucial for ensuring smoother model operation across different hardware configurations without sacrificing functionality.
🌿 That's all for now. Come back tomorrow.
6 of 6 items shown. Sources: 123 days indexed.