ArXiv

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

Authors
Zhihua Liang
Categories
cond-mat.dis-nn, cs.CL, cs.LG
arXiv
https://arxiv.org/abs/2607.17146v1
PDF
https://arxiv.org/pdf/2607.17146v1

Brief

A continuous geometric framework models Transformer operations as an integro-differential equation on a semantic fiber bundle, translating RMSNorm, RoPE, attention, FFNs, residual streams, and optimization into differential-geometry and stochastic-calculus terms. It predicts Attention as an entropic optimal-transport Schrödinger bridge and SGD as an Itô diffusion; six experiments on five models (124M–8B) confirm ε^{-1/2} scaling and O(1/√k) recurrence suppression.

Why it matters

Proposes a continuous geometric framework that models Transformers as an integro-differential equation on a semantic fiber bundle M × R^d, starting from one axiom (token sequences as a discrete 1-manifold with a canonical measure lattice) and mapping RMSNorm, RoPE, Softmax Attention, FFN, residual stream, SGD, and weight decay into differential-geometric/stochastic vocabulary.

Key details

  • Predicts specific quantitative phenomena: Attention ≃ entropic optimal transport / Schrödinger bridge; SGD ≃ Itô diffusion violating detailed balance; reports ε^{-1/2} Lipschitz scaling calibrated at machine precision (R^2 = 1.000), O(1/√k) suppression of Poincaré recurrence on the RoPE torus, Lie–Trotter operator-splitting torsion, a thermodynamic context-limit phase transition, and a non-equilibrium steady-state parameter vortex.
  • Empirically validates these predictions in a six-part experimental campaign over five architectures (Qwen3, LLaMA-3.1, Gemma-3, GPT-2, Mistral) spanning 124M–8B parameters; key observables hold under both AdamW and pure SGD (excluding momentum artifacts).
Source evidence

Abstract

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token sequence forms a discrete $1$-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning $124$M to $8$B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the $ε^{-1/2}$ Lipschitz scaling calibration at machine precision ($R^2 = 1.000$), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the $\calO(1/\sqrt{k})$ thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.