Topic: cache

5 stories found

Friday, September 18, 2026

trending46

Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)

Researchers have developed a method called "Cache-to-Cache" that allows large language models to communicate directly, enhancing collaboration and potentially improving model performance. This breakthrough could significantly advance the field of artificial intelligence by enabling more efficient and effective information sharing among different language models.

arxiv.org

Tuesday, September 15, 2026

releases48

ggml/llama.cpp releases: b10985

The ggml/llama.cpp project updated to hash-cache only weights for transfers above a certain threshold, allowing the rpc-server to serve cached files from a weight-specific cache, enhancing efficiency in model loading and management. This update is significant as it optimizes the workflow by reducing redundant data transfers and improving the performance of the model deployment process.

github.com

Monday, September 14, 2026

releases54

Ollama releases: v0.34.1

Ollama released version v0.34.1, addressing issues with ChatGPT model selector spacing and improving memory management by evicting cache snapshots and checking system free memory before loading new models, all while raising the token repeat limit to 100. These updates enhance stability and performance, making the software more reliable for users.

github.com

Friday, September 11, 2026

releases48

ggml/llama.cpp releases: b10907

The ggml/llama.cpp project released updates to fix MTP context kv cache allocation issues for specific architectures like deepseek2, glm4moe, and cohere2moe. These changes also include adding inverse architecture gating and comprehensive testing for the MTP layer filter, enhancing model stability and performance.

github.com

Wednesday, September 9, 2026

releases42

vLLM releases: v0.29.0

vLLM released version 0.29.0, which includes 594 commits from 277 contributors and marks the full rollout of Model Runner V2 as the default for all models, enhancing performance with CUDA graph memory profiling for KV cache auto-sizing. This update is significant as it improves model efficiency and scalability in large language model deployments.

github.com

🌿 That's all for now. Come back tomorrow.

5 of 5 items shown. Sources: 123 days indexed.