Topic: allocation

2 stories found

Today

research40

AdaMem: Adaptive Memory Token Allocation for Soft Compression in Retrieval-Augmented Generation

AdaMem introduces an adaptive memory token allocation method for soft compression in retrieval-augmented generation, aiming to reduce the cost of processing long passages while minimizing distracting information. This innovation matters because it enhances the efficiency and effectiveness of language models using RAG techniques.

arxiv.org

Friday, September 11, 2026

releases48

ggml/llama.cpp releases: b10907

The ggml/llama.cpp project released updates to fix MTP context kv cache allocation issues for specific architectures like deepseek2, glm4moe, and cohere2moe. These changes also include adding inverse architecture gating and comprehensive testing for the MTP layer filter, enhancing model stability and performance.

github.com

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 123 days indexed.