ArXiv

Linear representations of grammaticality in neural language models

Authors
Jane Li, Najoung Kim
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.15175v1
PDF
https://arxiv.org/pdf/2607.15175v1

Brief

Li and Kim investigate whether grammaticality is encoded in internal sentence representations of pretrained neural language models. Using mass-mean probing on multiple models, they report robust linear separation between grammatical and ungrammatical strings that cannot be fully attributed to confounds like frequency or plausibility, and that generalizes across phenomena and partly across languages. Summary is based on the paper's abstract.

Why it matters

Li & Kim (2026-07-16; arXiv:2607.15175v1) apply mass-mean probing and find grammatical vs. ungrammatical sentences are linearly separable in sentence representations of a wide range of pretrained neural language models; this representational separation is not fully explained by correlated sentence-level factors (e.g., lexical frequency, plausibility, world knowledge).

Key details

  • The grammaticality signal generalizes across a broad set of grammatical phenomena and, to some degree, across languages, providing a complementary, non-probability-based framework for evaluating syntactic competence in NLMs.
Source evidence

Abstract

Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that grammaticality is inherently entangled with likelihood. Model-assigned probability is a function of many related sentence properties, such as lexical frequency, plausibility, and world knowledge. In this work, we move beyond probability-based evaluations and investigate whether grammaticality is encoded in the internal representations of NLMs. Using mass-mean probing, we test whether grammatical and ungrammatical sentences are systematically separated in representational space. We further examine the extent to which these representations are independent of sentence properties that are correlated with grammaticality, as well as their generalization across grammatical phenomena and languages. Our results provide evidence that grammaticality is robustly encoded in sentence representations of a wide range of pretrained NLMs, yielding clear representational separation on the dimension of grammaticality that cannot be fully explained by alternative sentence-level factors. Moreover, this encoding generalizes across a broad range of grammatical phenomena and to some degree, across languages, suggesting that grammaticality constitutes a coherent representational dimension in contemporary NLMs. These findings contribute new evidence to debates about the nature of syntactic knowledge in language models and offer a complementary framework for evaluating grammatical competence that is not dependent on string probabilities alone.