Topic: inference

8 stories found

Friday, September 18, 2026

research40

VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

A new method called VisKG-LM compiles knowledge graphs into visual memory to improve multiple-choice question answering, addressing the issue of redundant encoding by storing graph information beforehand. This approach aims to enhance efficiency and accuracy in processing complex queries.

arxiv.orgโ†—

Wednesday, September 16, 2026

research40

Optimal Model Activation Policies for Inference Networks of Large Language Models

A new study explores optimal model activation policies for large language models to optimize performance while managing high inference costs, crucial for efficient use in NLP tasks.

arxiv.orgโ†—

Saturday, September 12, 2026

research35

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

Researchers have developed an enhanced training method for generating natural-language proofs in advanced mathematics, focusing on Olympiad problems, which could revolutionize how such complex mathematical challenges are approached and solved. This work matters because it potentially provides a robust framework for automating proof generation, aiding both education and research in mathematics.

arxiv.orgโ†—

Friday, September 11, 2026

research40

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The study challenges assumptions about membership evidence in language models by showing that detectable duplications are rare and hard to verify, suggesting caution when inferring training data content from prediction ease. This matters because it impacts how reliably we can use language model behavior to deduce their training data composition.

arxiv.orgโ†—

Thursday, September 10, 2026

research40

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

A new paper proposes X-CoSD, a communication-efficient method for collaborative speculative decoding that involves an on-device small language model generating candidates while a server large language model verifies them, aiming to improve distributed inference processes in large language models. This approach is crucial as it could enhance the efficiency and scalability of AI applications across devices.

arxiv.orgโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

8 of 8 items shown. Sources: 123 days indexed.