Topic: hate speech

2 stories found

Friday, August 28, 2026

research40

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

A new approach aims to make AI models used for detecting hate speech more transparent by aligning their reasoning with human rationales, addressing the challenge of culturally coded multilingual hate speech online that conventional systems can miss but lack explainability. This matters because it could help reduce bias and improve moderation accuracy without over-censorship or under-moderation, especially concerning Muslim communities.

arxiv.org

🌿 That's all for now. Come back tomorrow.

2 of 2 items shown. Sources: 107 days indexed.