Topic: judges

2 stories found

Yesterday

research40

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

A new method uses cheaper open-source language models to grade natural-language mathematical proofs, reducing costs associated with evaluating math-reasoning systems. This approach addresses the high expense of using advanced language models like LLMs for such tasks.

arxiv.org↗

Monday, August 3, 2026

research40

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

A new method called Chain-of-Models aims to audit large language models (LLMs) used as judges to reduce bias, addressing limitations of current mitigation techniques that are either ineffective against various biases or impractical at scale due to reliance on human evaluation.

arxiv.org↗

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 77 days indexed.