ArXiv

Linguistic Monoculture in LLM-Assisted Language Use

Authors
Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen
Categories
cs.AI, cs.CL, cs.GT
arXiv
https://arxiv.org/abs/2607.27134v1
PDF
https://arxiv.org/pdf/2607.27134v1

Brief

Linguistic monoculture in LLM-assisted language use: Thejaswi et al. (2026) develop a mathematical framework treating authors and LLMs as distributions over linguistic features and study three interaction regimes (fixed shared, recursively updated shared, personalized). They characterize equilibria and convergence, show over-conformity creates a negative externality with a potentially unbounded price-of-monoculture in extreme preferences, and support findings with simulations.

Why it matters

Authors and LLMs are modeled as distributions over linguistic features; the paper analyzes three interaction mechanisms: (1) a fixed shared model, (2) a shared model recursively updated from author outputs, and (3) personalized models updated via author-specific and population-level feedback.

Key details

  • Equilibrium results: fixed shared models drive authors toward a common linguistic norm; recursive feedback can relocate the shared norm without changing pairwise spread under common conformity; personalization can sustain a family of distinct author–model equilibria with nonzero linguistic diversity.
  • A game-theoretic utility model endogenizes conformity and finds individually rational authors over-conform, producing a negative externality. The authors define a 'price of monoculture' that is finite for each fixed instance but can grow unbounded when distinctiveness dominates authenticity; synthetic simulations illustrate differing long-run diversity outcomes.
Source evidence

Abstract

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.