APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the data bottleneck in three-dimensional atomic structure prediction caused by reliance on supervised alignment by proposing the first fully unsupervised atomic policy optimization framework. The approach integrates group-relative policy optimization with a dual-reward mechanism—comprising a dominant-mode reward derived from similarity-based feature decomposition and a thermodynamic stability reward—to guide the generation of physically plausible and thermodynamically stable configurations without requiring ground-truth structural labels. By incorporating flow-matching models and thermodynamic constraints, the framework surpasses fully supervised baselines in both crystal and antibody structure prediction tasks, achieving new state-of-the-art performance in structural matching accuracy and fidelity while significantly enhancing inference efficiency.
📝 Abstract
Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct'' by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.
Problem

Research questions and friction points this paper is trying to address.

3D structure prediction
unsupervised learning
atomic systems
data-scarce regimes
structural modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

unsupervised alignment
atomic structure prediction
policy optimization
dual-reward mechanism
thermodynamic stability