Topic: evaluation

10 stories found

Today

research40

Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models

A new study evaluates large language models (LLMs) based on their coding behaviors rather than just performance metrics like pass@k, highlighting that as models improve, traditional evaluation methods become less effective in distinguishing between them.

arxiv.orgโ†—

Yesterday

ai_labs75

Building standards for the next phase of AI

OpenAI proposes establishing global standards for AI to enhance safety through coordinated evaluation and governance. This initiative aims to address AI risks by fostering international cooperation in the tech sector.

openai.comโ†—
research40

From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators

A new study proposes using large language models (LLMs) in a persona-grounded, open-ended simulation as discharge educators to better adapt to patients' literacy, recall, and personality needs, addressing limitations of current LLM evaluations that focus on static or artifact-generation tasks. This approach aims to improve patient understanding and adherence to discharge plans.

arxiv.orgโ†—

Thursday, September 17, 2026

research40

Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation

Large language models were evaluated against physicians in diagnosing and treating cases within traditional Chinese medicine, with a clinical case library of 349 de-identified patients used to assess their effectiveness. This study highlights the potential of LLMs in TCM but also underscores the need for further validation in real-world applications.

arxiv.orgโ†—

Wednesday, September 16, 2026

research40

Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions

Chen, Zhao, and Cohan's study evaluates LLM-generated research ideas but faces criticism for survivorship bias in its human baseline, which includes only published papers, while the LLM baseline considers one-shot responses. This discrepancy highlights a potential flaw in how the LLM's performance is being compared to real-world scenarios.

arxiv.orgโ†—

Saturday, September 12, 2026

research35

When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

The study argues that while independent evaluation is crucial to rejecting harmful updates, it should not hinder the continual learning process for embodied agents. It suggests assessing update admission based on both error control and preserving learning opportunities within a set interaction limit.

arxiv.orgโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

10 of 10 items shown. Sources: 123 days indexed.