ArXiv

Hamid Reza Firoozfar et al.

Authors
Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2606.27314v1
PDF
https://arxiv.org/pdf/2606.27314v1

Brief

A mechanism-oriented taxonomy for indirect linguistic expressions (ILE) categorizes the operations that encode and recover meaning (e.g., algospeak, euphemisms, adversarial obfuscation) rather than speakers' intents. The authors integrate this taxonomy into LLM prompts and, on 2,000 annotated TikTok and Bluesky posts evaluated with three LLMs, report the best document- and span-level results—improving accuracy by 4.7% and F1 by 5.4% over prior taxonomies. Summary based on the abstract (full text not reviewed).

Why it matters

Hamid Reza Firoozfar et al. (published 2026-06-25; arXiv:2606.27314v1) introduce a comprehensive, mechanism-oriented taxonomy of indirect linguistic expressions (ILE) that categorizes the underlying encoding and recovery operations rather than communicative goals.

Key details

  • They embed the taxonomy into LLM prompts and evaluate it against four existing taxonomies plus a no-taxonomy baseline on 2,000 manually annotated TikTok and Bluesky posts using three LLMs, achieving the strongest document- and span-level detection performance.
  • The taxonomy yields an absolute improvement of 4.7% accuracy and 5.4% F1 over the best-performing benchmark; the work was submitted for ARR review for EMNLP 2026 and the abstract warns of profane/offensive content.
Source evidence

Abstract

To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive meanings. Such expressions surface as algospeak, euphemisms, and adversarial obfuscation, depending on intent and context, and they involve recurring encoding mechanisms. We propose a comprehensive, mechanism-oriented taxonomy of ILE that abstracts away from communicative goals and instead categorizes the underlying operations through which meaning is encoded and recovered. We evaluate the taxonomy by incorporating it into LLM prompts and comparing it with four existing taxonomies and a no-taxonomy baseline, using 2,000 manually annotated TikTok and Bluesky posts. The proposed taxonomy attains the strongest document- and span-level performance across the three LLMs, achieving an improvement of 4.7% in accuracy and 5.4% in F1 over the best-performing benchmark. The empirical results reveal the importance of a comprehensive, mechanism-oriented taxonomy as a stable scaffold for detecting emerging coded language and a useful input to content moderation. Disclaimer: This paper contains content that may be profane, vulgar, or offensive.

Comment: Submitted for review in ARR for EMNLP 2026