Topic: probe

5 stories found

Friday, September 4, 2026

research40

Probe Generalization as Subspace Selection for OOD Deception Detection

Researchers found that linear probes can effectively detect deceptive behaviors within language model data but struggle with out-of-distribution examples, highlighting the need for improved methods in detecting deception across different contexts. This matters because current techniques may not reliably identify deceptive content when encountered in new or unseen scenarios.

arxiv.org

Monday, August 24, 2026

research40

Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit

The study investigates if stereotype-loaded queries reveal more personally identifiable information from a RAG system compared to neutral ones, focusing on culturally marked topics across four cultures. This matters as it highlights potential biases and privacy risks in AI systems when handling sensitive cultural queries.

arxiv.org

🌿 That's all for now. Come back tomorrow.

5 of 5 items shown. Sources: 107 days indexed.