Topic: distribution

2 stories found

Friday, September 4, 2026

research40

Probe Generalization as Subspace Selection for OOD Deception Detection

Researchers found that linear probes can effectively detect deceptive behaviors within language model data but struggle with out-of-distribution examples, highlighting the need for improved methods in detecting deception across different contexts. This matters because current techniques may not reliably identify deceptive content when encountered in new or unseen scenarios.

arxiv.org

Friday, August 28, 2026

research40

Recipes for Steering and Scaling LLMs via Sampling

The paper "Recipes for Steering and Scaling LLMs via Sampling" addresses inefficiencies in current sampling methods for Large Language Models (LLMs), proposing new techniques to more effectively scale and steer these models. This matters because improving sampling strategies could enhance the performance and applicability of LLMs across various tasks.

arxiv.org

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 107 days indexed.