ArXiv

Autoregressive Boltzmann Generators

Authors
Danyal Rehman, Charlie B. Tan, Yoshua Bengio...
Categories
cs.LG, cs.AI
arXiv
https://arxiv.org/abs/2606.27361v1
PDF
https://arxiv.org/pdf/2606.27361v1

Brief

Autoregressive Boltzmann Generators (ArBG) introduce an autoregressive generative framework for sampling molecular Boltzmann distributions that avoids normalizing-flow topology constraints and supports sequential inference-time interventions. By leveraging LLM-style scalable architectures, ArBG outperforms flow-based Boltzmann Generators on all benchmarks—most notably on the 10-residue Chignolin—and yields Robin, a 132M-parameter transferable model that cuts zero-shot E-W2 error on 8-residue systems by over 60%.

Why it matters

ArBG introduces an autoregressive Boltzmann Generator framework that removes normalizing-flow topological constraints, enables sequential inference-time interventions, and scales via LLM-style architectures.

Key details

  • ArBG outperforms flow-based Boltzmann Generators across benchmarks—especially on the 10-residue peptide Chignolin—and Robin, a 132M-parameter ArBG model, reduces zero-shot energy error E-W2 on 8-residue systems by over 60%.
  • Code and models are available at https://github.com/danyalrehman/autobg; the paper appeared as an ICML 2026 (Spotlight) submission and is on arXiv: https://arxiv.org/abs/2606.27361v1.
Source evidence

Abstract

Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has driven the development of Boltzmann Generators (BGs), which allow rapid generation of uncorrelated equilibrium samples by combining a generative model with exact likelihoods and an importance sampling correction. However, modern BGs predominantly rely on normalizing flows (NFs), which either suffer from limited expressivity due to strict invertibility constraints (discrete time) or computationally expensive likelihoods (continuous time). In this paper, we propose Autoregressive Boltzmann Generators (ArBG) -- a novel autoregressive modelling framework -- that overcomes these limitations by departing from the flow-based BG paradigm. ArBG circumvents the topological constraints of flows and enables sequential inference-time interventions, while offering enhanced scalability by leveraging architectures effective in Large Language Models. We empirically demonstrate that ArBG leads to significant improvements over flow-based models across all benchmarks, but particularly in larger peptide systems such as the 10-residue Chignolin. Furthermore, we introduce Robin, a 132 million parameter transferable model trained with the ArBG framework which improves over the previous state-of-the-art, reducing the zero-shot energy error, E-W$_2$, on 8-residue systems by over 60$\%$. The code can be found at the following link: https://github.com/danyalrehman/autobg.

Comment: ICML 2026 (Spotlight)