ArXiv

Literary Non-Style in LLM-Generated Text

Authors
Cory Massaro
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.17228v1
PDF
https://arxiv.org/pdf/2607.17228v1

Brief

Literary Non-Style in LLM-Generated Text (Cory Massaro, arXiv 2026-07-19) analyzes n-gram statistics to expose systematic departures of LLM outputs from human writing. The paper finds simple, consistent distributional patterns—especially in higher-order n-grams—whose qualitative properties reveal stylistic deficiencies and support the claim that style and semantics are entangled. Full text was not provided here (abstract-only).

Why it matters

Identifies simple, consistent patterns in the statistical distribution of n-grams within LLM-generated text, with particular emphasis on higher-order n-grams as measurable markers that depart from human writing (Massaro, 2026-07-19).

Key details

  • Presents qualitative analysis showing stylistic deficiencies in LLM output and argues higher-order n-grams correlate with semantic content, implying style and semantics are not cleanly separable.
  • ArXiv preprint by Cory Massaro (cs.CL), published 2026-07-19 as arXiv:2607.17228v1 (PDF: https://arxiv.org/pdf/2607.17228v1).
Source evidence

Abstract

Prior work on LLM-generated text has demonstrated quantitative and qualitative departures from text produced by humans. LLM-generated texts differ from human writing in style, resulting in a characteristic textual "feel," while the semantic range of LLMs is much restricted compared to that of humans. In this contribution, I note simple but consistent patterns in the statistical distribution of n-grams within LLM-generated text. Via qualitative analysis of these n-grams, I reveal deficiencies in LLM style. Because higher-order n-grams correlate to semantic content, I conclude that questions of style and semantics are not cleanly separable.