ArXiv

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Authors
Shentong Mo, Yatao Bian
Categories
cs.LG, cs.AI, cs.MA
arXiv
https://arxiv.org/abs/2607.28553v1
PDF
https://arxiv.org/pdf/2607.28553v1

Brief

Atomic Policy Optimization (APO) proposes a fully unsupervised alignment method for 3D atomic-structure prediction to remove the expensive ground-truth coordinate supervision required by recent flow-matching models. APO combines group-relative policy optimization with a dual-reward mechanism—an eigen-decomposition reward to reinforce dominant latent structural modes and a thermodynamic-stability reward—to self-correct sampled groups. According to the abstract, APO achieves state-of-the-art match rates and structural fidelity on crystal and antibody benchmarks and improves inference efficiency; full text was not available for this summary.

Why it matters

Atomic Policy Optimization (APO), introduced by Shentong Mo and Yatao Bian (arXiv 2026-07-30), is a fully unsupervised alignment framework for 3D atomic-structure prediction that removes reliance on ground-truth coordinate supervision used by flow-matching models (e.g., FlowDPO).

Key details

  • APO adapts group-relative policy optimization with a dual-reward scheme: (i) an eigen-decomposition-based reward that reinforces dominant latent structural modes across sampled groups, and (ii) a thermodynamic-stability reward; the method reportedly outperforms fully supervised baselines on crystal and antibody benchmarks, achieving new state-of-the-art match rates and structural fidelity while straightening probability paths to improve inference efficiency.
Source evidence

Abstract

Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct'' by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.