🤖 AI Summary
This study addresses the inefficiency and instability arising from the coupling of search strategies and structure generation in de novo crystal design by proposing the ANCHOR framework. This method decouples combinatorial space search from structure generation, leveraging a frozen crystal structure prediction (CSP) prior as an evaluation benchmark. It employs Group Relative Policy Optimization (GRPO) reinforcement learning with a multi-objective reward mechanism for adaptive exploration, alongside a continuous adaptive novelty score to guide the search process. Experimental results demonstrate that optimizing the search strategy under fixed physical priors significantly outperforms directly fine-tuning generative models. Specifically, ANCHOR achieves an MSUN of 47.6% and a SUN of 22.1%, while elevating the state-of-the-art MSUN to 41.3% within the MatterGen pipeline.
📝 Abstract
De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likelihood training, while rewards, including novelty measured against the search's own history, should act on a search over compositions. We introduce ANCHOR, a GRPO composition policy trained with multi-objective rewards around a frozen CSP model, and continuous adaptive novelty (CAN), a graded novelty score against known structures and a growing discovery history. Using the frozen CSP model as a fixed ruler under one evaluator, we test where adaptation should act. Replacing DNG compositions with ANCHOR's policy on the same CSP backbone raises MSUN from 11.4% to 47.6% and SUN from 1.1% to 22.1% at 99.9% formula uniqueness. Fine-tuning DNG models directly on the same rewards instead moves their composition marginal without raising their on-hull fraction. We show that KL-regularized fine-tuning of a DNG model can only reweight chemistry the pretrained model already supports by a bounded factor, while unregularized DNG fine-tunes move toward known or less stable chemistry. Even a stability-only reward routed into ANCHOR's CSP backbone roughly halves SUN relative to the frozen backbone, whereas likelihood training on structures found during search can improve a CSP backbone. Under MatterGen's evaluation pipeline, ANCHOR raises state-of-the-art MSUN from 29.2% to 41.3%, transfers without retraining to two further CSP backbones, and reaches 47.1% after distillation into Crystalite-CSP. As with any model optimised against a potential, its on-hull rate depends on that potential.