Topic: rag

28 stories found

Yesterday

trending37

Show HN: We Beat MLPerf: Modern Storage for KV Offload and LLM Training

A tech company has developed modern storage solutions that outperformed existing methods in key value-offloading and large language model training tasks, as measured by MLPerf benchmarks. This breakthrough could significantly enhance the efficiency and speed of AI training processes, making advanced AI technologies more accessible and practical.

theopenlake.com

Friday, September 4, 2026

releases48

ggml/llama.cpp releases: b10814

The ggml/llama.cpp project has released updates that extend OpenCL support by adding nine new elementwise operations, which previously relied on the CPU, thereby improving performance on GPU-accelerated systems. This update is significant because it enhances the efficiency and versatility of the library for tasks requiring extensive mathematical computations.

github.com
research40

R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG

A new adapter called R$^{2}$Adapter addresses the limitations of traditional Retrieval-Augmented Generation (RAG) by improving its ability to handle complex, multi-hop reasoning tasks, making Large Language Models more versatile and efficient. This advancement is crucial as it enhances the capability of LLMs to process more intricate queries, thereby broadening their practical applications.

arxiv.org
industry32

Architecting memory and storage in the AI era

The AI inference era enables real-time analysis of vast datasets, crucial for advancements like accelerated medical research and efficient customer service. This technology is pivotal as it demonstrates the potential of AI to transform various industries through rapid data processing and decision-making.

technologyreview.com

Wednesday, September 2, 2026

28 of 28 items shown. Sources: 107 days indexed.