ArXiv

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

Authors
Jay L. Cunningham, Mark Atta Mensah, Richard Martinez...
Categories
cs.CL, cs.CY, cs.HC
arXiv
https://arxiv.org/abs/2608.06141v1
PDF
https://arxiv.org/pdf/2608.06141v1

Brief

Decolonizing linguistic policies in automated speech recognition, Cunningham et al. argue that ASR data, metrics, and model priors enact colonial language hierarchies and limit access to public services, healthcare, and education. They introduce the Three Harms (3M) taxonomy (Misrecognition, Misalignment, Mistrust), a seven-layer situatedness model for linguistic diversity, and a participatory framework with a minimum audit protocol that centers affected communities as co-designers and governance partners.

Why it matters

Introduces the Three Harms (3M) taxonomy—Misrecognition, Misalignment, and Mistrust—to frame ASR failures as implicit linguistic policies that reproduce colonial language hierarchies.

Key details

  • Proposes a seven-layer situatedness model for linguistic diversity in ASR and an accompanying participatory framework plus a minimum audit protocol that positions affected communities as co-designers, evaluators, and governance partners.
  • Argues data, metrics, and model priors determine whose voices become machine-legible; paper is on arXiv (2608.06141v1), authored by Cunningham et al., submitted 2026-08-06 and listed as an Interspeech 2026 contribution (10 pages, 2 figures, 2 tables).
Source evidence

Abstract

This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.

Comment: 10 Pages, 2 Figures, 2 Tables, Interspeech 2026 - Sydney, Australia