Topic: large language model
51 stories found
Today
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
A new benchmark called OncoTriad-QA has been introduced to evaluate the ability of AI models to integrate radiology, pathology, genomics, and clinical data for cancer diagnosis, addressing the current gap in existing benchmarks that primarily focus on single-modal evidence.
Yesterday
RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review
A new approach to peer review using rubrics aims to address the challenges posed by high submission volumes at major academic venues, particularly by mitigating the limitations of current large language model (LLM) based reviewers who struggle with direct critique. This method emphasizes objective and comprehensive evaluation, potentially improving the quality and efficiency of the review process.
Monday, August 3, 2026
Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
The study examines whether large language models can accurately predict item difficulty, crucial for educational assessments, raising questions about their reliability in automated test generation.
Friday, July 31, 2026
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
BridgeAlign addresses the gap in data synthesis for large language models by focusing on humanities and social sciences, where nuanced quality judgments are crucial, rather than targeting domains with verifiable answers. This approach aims to improve the relevance and accuracy of LLMs in open-ended fields.
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
Researchers are examining whether large language models understand and convey emotional nuances through different sociodemographic frames, a critical aspect as these models increasingly influence public opinion. This study addresses concerns beyond bias, focusing on how LLMs align with sympathetic or empathetic framing across diverse groups.
Thursday, July 30, 2026
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
A new study proposes using large language models to help verify claims within papers during the peer review process, aiming to address the increasing volume of scientific submissions. This method could enhance the accuracy of assessing a paper's originality and significance by comparing its content directly with existing literature.
Wednesday, July 29, 2026
TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking
TimeCapsule is a new method using a large language model to address the temporal bias in contemporary training data, making these models unreliable for historical contexts. The researchers developed TimeCapsule, a 1.2B-parameter model, to better handle historical sensemaking by reducing present-day concept encoding.
Tuesday, July 28, 2026
Monday, July 27, 2026
Sunday, July 26, 2026
yamadashy/repomix ā š¦ Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file. Perfect for when you nee
Repomix is a tool that compresses entire repositories into single, AI-friendly files, ideal for integrating codebases with large language models and other AI tools. This simplifies the process of feeding complex code environments to AI systems for analysis or collaboration.
51 of 51 items shown. Sources: 77 days indexed.