Topic: benchmark

20 stories found

Today

research40

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

A new benchmark called LayerRAG-Bench has been introduced to evaluate the reliability of agentic retrieval-augmented generation systems across multiple layers, highlighting their potential failures in grounding answers. This benchmark is crucial for improving the overall trustworthiness and practical utility of these systems by identifying specific areas where they may fall short.

arxiv.org↗

Yesterday

research40

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

The study examines the validity of using large language models (LLMs) as substitutes for human respondents in surveys across different domains, highlighting scenarios where such simulations may fail due to limitations in the models' understanding. This research is crucial as LLMs are increasingly relied upon for making significant decisions in product development, policy-making, and market analysis.

arxiv.org↗

Wednesday, July 29, 2026

ai_labs75

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Enabling two specific API settings significantly boosted GPT-5.6's performance on the ARC-AGI-3 benchmark, tripling scores while improving efficiency by enhancing reasoning capabilities and compacting data retention.

openai.com↗
research40

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

MyoCardBench is a new benchmark for evaluating large language models in realistic cardiovascular care scenarios, addressing limitations of existing benchmarks which often focus on isolated tasks or knowledge rather than longitudinal, multimodal, and safety-critical clinical workflows.

arxiv.org↗

Tuesday, July 28, 2026

20 of 20 items shown. Sources: 72 days indexed.