ArXiv

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

Authors
Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino
Categories
cs.CL, cs.AI, cs.LG
arXiv
https://arxiv.org/abs/2608.03930v1
PDF
https://arxiv.org/pdf/2608.03930v1

Brief

Logic pre-pretraining (Logic-PPT) initializes language models by training on formal derivations to impart structural and linguistic biases lacking in Dyck/procedural primitives. At a 100B-token scale, Logic‑PPT accelerates skill acquisition—reaching 80% accuracy on linguistic tasks with 36B fewer tokens than standard initialization—and reorganizes representations into a lower‑rank, spectrally concentrated space that permits pruning to ≈33% sparsity without loss.

Why it matters

Logic-PPT pre-pretraining on formal derivations at a 100B-token scale accelerates skill acquisition, reaching 80% accuracy on linguistic tasks with 36B fewer tokens than standard initialization and outperforming alternative pre-pretraining baselines.

Key details

  • Formal-derivation pre-pretraining reorganizes internal representations into a lower-rank, spectrally concentrated space that improves compressibility: pruning to ≈33% sparsity matches dense baseline performance.
Source evidence

Abstract

Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.