Topic: scores

3 stories found

Yesterday

ai_labs75

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Enabling two specific API settings significantly boosted GPT-5.6's performance on the ARC-AGI-3 benchmark, tripling scores while improving efficiency by enhancing reasoning capabilities and compacting data retention.

openai.com↗
research40

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

Researchers have developed CogArena, a tool for evaluating cognitive abilities in large language models (LLMs) across multiple methods, aiming to assess whether LLM scores reflect true cognitive structures that consistently emerge and generalize. This matters because it could improve the understanding of LLM capabilities and their alignment with human cognitive functions.

arxiv.org↗

🌿 That's all for now. Come back tomorrow.

3 of 3 items shown. Sources: 71 days indexed.