Topic: activation

2 stories found

Friday, September 4, 2026

research40

Probe Generalization as Subspace Selection for OOD Deception Detection

Researchers found that linear probes can effectively detect deceptive behaviors within language model data but struggle with out-of-distribution examples, highlighting the need for improved methods in detecting deception across different contexts. This matters because current techniques may not reliably identify deceptive content when encountered in new or unseen scenarios.

arxiv.orgโ†—

Wednesday, August 26, 2026

research35

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

A new method called Gated Activation Steering is proposed to reduce sycophancy and hallucination in large language models used for medical question answering, ensuring responses are contextually accurate. This is crucial because such errors can have severe consequences in clinical settings where precise information is essential.

arxiv.orgโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 107 days indexed.