Topic: tensor

6 stories found

Yesterday

releases48

ggml/llama.cpp releases: b11081

The ggml/llama.cpp project released a new version with updates to make tensor data standard deviation configurable and added more examples to the documentation, enhancing flexibility and usability for developers. These changes are significant as they improve the toolkit's adaptability and user-friendliness in handling large language model architectures.

github.com

Thursday, September 17, 2026

releases48

ggml/llama.cpp releases: b11026

The ggml/llama.cpp project released version b11026, which includes a model update to skip gate_up_exps if TENSOR_SKIP is set. This update is crucial for qwen35moe when MTP tensors are fused but not loaded, ensuring compatibility and functionality across different configurations.

github.com

Tuesday, September 15, 2026

releases48

ggml/llama.cpp releases: b10985

The ggml/llama.cpp project updated to hash-cache only weights for transfers above a certain threshold, allowing the rpc-server to serve cached files from a weight-specific cache, enhancing efficiency in model loading and management. This update is significant as it optimizes the workflow by reducing redundant data transfers and improving the performance of the model deployment process.

github.com

Monday, September 14, 2026

releases54

Ollama releases: v0.34.1

Ollama released version 0.34.1, which marks the full integration of MLX safetensors support in their `ollama create` command and introduces stricter conditions for detecting runaway repeat tokens. These updates enhance model handling on Apple Silicon and improve overall stability.

github.com

Thursday, September 10, 2026

releases48

ggml/llama.cpp releases: b10901

The ggml/llama.cpp project released version b10901, which includes a Vulkan update allowing CPU writes in asynchronous tensor copies when the context is idle. This update aims to optimize performance by leveraging CPU resources more efficiently during idle GPU contexts.

github.com

Tuesday, September 8, 2026

releases48

ggml/llama.cpp releases: b10867

The ggml/llama.cpp project updated its codebase to disable lazy tensor loading by default on integrated GPUs (iGPUs) to address performance issues, impacting users of these less powerful graphics cards. This change is crucial for ensuring smoother model operation across different hardware configurations without sacrificing functionality.

github.com

🌿 That's all for now. Come back tomorrow.

6 of 6 items shown. Sources: 123 days indexed.