ArXiv

How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation

Authors
Maria Levchenko
Categories
cs.CL, cs.DL
arXiv
https://arxiv.org/abs/2606.27275v1
PDF
https://arxiv.org/pdf/2606.27275v1

Brief

The paper evaluates LLMs' handling of historical language by proposing a four-part diagnostic (tokenization cost, predictive surprisal, semantic robustness, context sensitivity) and testing on 17th‑century Italian (1610–1689), 19th‑century I Promessi Sposi (control), and 18th‑century Russian. It finds 25–30% tokenization inflation, 17th‑century Italian 2.4× surprisal (3.2× in academic prose), embeddings >0.85, and a temporal prompt cutting surprisal ~60%.

Why it matters

The paper introduces a four-dimension diagnostic (tokenization cost, predictive uncertainty/surprisal, semantic robustness, context sensitivity) and evaluates it on three corpora: 17th‑century Italian (1610–1689), 19th‑century I Promessi Sposi (high-exposure control), and 18th‑century Russian civil print books.

Key details

  • Tokenization inflation is comparable for Russian and early modern Italian (25–30%), but 17th‑century Italian is on average 2.4× more surprising than modern Italian (academic prose up to 3.2×); Russian shows only a modest surprisal increase.
  • Semantic representations remain robust (embedding similarity >0.85 across datasets), and a minimal temporal context prompt reduces historical surprisal by ~60%, offering a simple, model-agnostic mitigation for generative tasks.
Source evidence

Abstract

Large language models (LLMs) are increasingly critical to digital library workflows, yet their ability to process historical language remains poorly understood. Historical difficulty is typically treated as a monolithic barrier, conflating orthographic variation, linguistic distance, and pretraining exposure. In this paper, we propose a diagnostic framework that decomposes this difficulty into four distinct dimensions: tokenization cost, predictive uncertainty (surprisal), semantic robustness, and context sensitivity. We evaluate this framework on three datasets spanning three centuries: (1) a newly curated corpus of 17th-century Italian texts (1610-1689) digitized from original page images; (2) canonical 19th-century Italian "I Promessi Sposi" serving as a high-exposure control; and (3) 18th-century Russian civil print books as a contrastive orthographic stress test. Our results reveal a distinct dissociation between encoding cost and comprehension. While Russian and early modern Italian incur comparable tokenization penalties (25-30% inflation), their predictive difficulty diverges sharply. 17th-century Italian is on average 2.4 times more surprising than its modern equivalent - with academic prose reaching 3.2 times - whereas Russian shows only a modest increase. But predictive uncertainty does not imply representational degradation: embedding similarity remains robust (> 0.85) across all datasets, confirming that models can represent historical meaning even when generation is unstable. Finally, we demonstrate that a minimal temporal context prompt reduces historical surprisal by approximately 60%, offering a simple, model-agnostic mitigation. These findings suggest that while historical text imposes a consistent encoding tax, digital libraries can safely deploy LLMs for semantic retrieval tasks, provided that generative applications are carefully adapted.

Comment: The 22nd Conference on Information and Research Science Connecting to Digital and Library Science