Topic: benchmark
20 stories found
Today
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
A new benchmark called LayerRAG-Bench has been introduced to evaluate the reliability of agentic retrieval-augmented generation systems across multiple layers, highlighting their potential failures in grounding answers. This benchmark is crucial for improving the overall trustworthiness and practical utility of these systems by identifying specific areas where they may fall short.
Yesterday
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
The study examines the validity of using large language models (LLMs) as substitutes for human respondents in surveys across different domains, highlighting scenarios where such simulations may fail due to limitations in the models' understanding. This research is crucial as LLMs are increasingly relied upon for making significant decisions in product development, policy-making, and market analysis.
Wednesday, July 29, 2026
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
Enabling two specific API settings significantly boosted GPT-5.6's performance on the ARC-AGI-3 benchmark, tripling scores while improving efficiency by enhancing reasoning capabilities and compacting data retention.
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
MyoCardBench is a new benchmark for evaluating large language models in realistic cardiovascular care scenarios, addressing limitations of existing benchmarks which often focus on isolated tasks or knowledge rather than longitudinal, multimodal, and safety-critical clinical workflows.
Tuesday, July 28, 2026
Monday, July 27, 2026
Saturday, July 25, 2026
Why the OpenAI Agent Broke Into Hugging Face: Reward Hacking, Not Malice, Explained for Engineers
Friday, July 24, 2026
Wednesday, July 22, 2026
Tuesday, July 21, 2026
20 of 20 items shown. Sources: 72 days indexed.