ArXiv

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

Authors
Iosif Lytras, Nikolaos Makras, Sotirios Sabanis
Categories
cs.LG, math.OC, math.PR, stat.ML
arXiv
https://arxiv.org/abs/2608.06283v1
PDF
https://arxiv.org/pdf/2608.06283v1

Brief

The Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA) targets sampling from non-smooth, non-convex potentials with superlinear gradient growth by directly using subgradients and taming to obtain a stable explicit discretisation. The authors prove non-asymptotic Wasserstein-2 convergence bounds with constants explicit in dimension and inverse temperature, give excess-risk guarantees, and validate assumptions (with constants) on a regularized GPT-2 pretraining potential, reporting competitive pretraining performance versus AdamW and Muon.

Why it matters

Introduces SG-TULA (Subgradient Tamed Unadjusted Langevin Algorithm), an explicit Langevin discretisation that operates directly on subgradients (no smoothing) and uses taming to handle superlinear gradient growth in non-convex, non-smooth potentials.

Key details

  • Derives non-asymptotic convergence bounds in Wasserstein-2 distance with all constants tracked explicitly in terms of dimension and inverse temperature, improving known rates for subgradient-based Langevin methods; also provides excess risk estimates for the associated optimization problem.
  • Verifies assumptions (with explicit constants) for a regularized GPT-2–lineage pretraining potential and shows a boosted coordinate-wise SG-TULA variant pretrains competitively against finetuned AdamW and Muon; paper posted on arXiv 2026-08-06 (53 pages).
Source evidence

Abstract

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.

Comment: 53 pages