ArXiv

CATCH-ME if you RAG presents the first large-scale, expert-curated multilingual…

Authors
Helena Bonaldi, Genoveffa Martone, Marco Guerini
Categories
cs.CL
arXiv
https://arxiv.org/abs/2606.20369v1
PDF
https://arxiv.org/pdf/2606.20369v1

Brief

CATCH-ME if you RAG presents the first large-scale, expert-curated multilingual, multi-turn dataset for counterspeech tackling the intersection of hate and misinformation. Dialogues are anchored to verified sources (fact-checks and NGO reports) and include document- and chunk-level span annotations to support RAG; it spans five languages and targets seven marginalized groups to enable training and evaluation of more persuasive, factually grounded counterspeech models.

Source evidence

Abstract

Online hate speech and misinformation frequently overlap, yet NLP research has mainly treated them in isolation. While LLMs represent a scalable solution for assisting humans in the generation of counterspeech for both threats, zero-shot models frequently generate repetitive and vague responses, underscoring the need for high-quality examples to steer model generation. However, existing counterspeech datasets against the overlap of hate and misinformation are scarce and limited to single-turn English dialogues, while real-life interactions span across multiple turns and languages. To bridge this gap, we introduce the first large-scale, expert-curated, multilingual dataset of dialogues tackling the intersection of hate and misinformation. To ensure factual grounding, the dialogues are also anchored in verified external knowledge (i.e., fact-checking articles and NGO reports) and include document- and chunk-level span annotations, making it directly applicable for RAG systems. Covering five languages and targeting hate directed at seven marginalized groups, this novel resource enables the training and evaluation of more persuasive, factually grounded counterspeech models.