Topic: rag
28 stories found
Yesterday
Show HN: We Beat MLPerf: Modern Storage for KV Offload and LLM Training
A tech company has developed modern storage solutions that outperformed existing methods in key value-offloading and large language model training tasks, as measured by MLPerf benchmarks. This breakthrough could significantly enhance the efficiency and speed of AI training processes, making advanced AI technologies more accessible and practical.
Friday, September 4, 2026
ggml/llama.cpp releases: b10814
The ggml/llama.cpp project has released updates that extend OpenCL support by adding nine new elementwise operations, which previously relied on the CPU, thereby improving performance on GPU-accelerated systems. This update is significant because it enhances the efficiency and versatility of the library for tasks requiring extensive mathematical computations.
R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG
A new adapter called R$^{2}$Adapter addresses the limitations of traditional Retrieval-Augmented Generation (RAG) by improving its ability to handle complex, multi-hop reasoning tasks, making Large Language Models more versatile and efficient. This advancement is crucial as it enhances the capability of LLMs to process more intricate queries, thereby broadening their practical applications.

Architecting memory and storage in the AI era
The AI inference era enables real-time analysis of vast datasets, crucial for advancements like accelerated medical research and efficient customer service. This technology is pivotal as it demonstrates the potential of AI to transform various industries through rapid data processing and decision-making.
Thursday, September 3, 2026
Wednesday, September 2, 2026
Tuesday, September 1, 2026
Monday, August 31, 2026
28 of 28 items shown. Sources: 107 days indexed.