Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inverse design for targeted properties and reward hacking in reinforcement learning within crystal generation models. We propose a framework that aligns stochastic interpolation generative models using Group Relative Policy Optimization (GRPO). This approach applies policy gradients to discrete composition channels for the first time, directly optimizing atomic type transition probabilities via discrete flow matching to enable goal-directed crystal structure generation. Furthermore, it exposes and mitigates reward hacking by introducing a novel evaluation criterion that stratifies benchmarks according to the number of reference phases. Experimental results demonstrate a substantial increase in the yield of metastable, unique, and novel (mSUN) structures from 13.4% to 45.5%, validating the framework’s effectiveness in enhancing diffusion models while revealing critical limitations in existing evaluation metrics.
📝 Abstract
Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimization (GRPO) to align a generative model based on stochastic interpolants and discrete flow matching with general black-box reward functions through reinforcement learning. Atom types are generated by a discrete flow and the policy gradient of our generalization of GRPO directly acts on the likelihoods of the atom-type transitions, which differentiates our work from previous reinforcement-learning approaches for diffusion and flow-based generative models of crystalline materials. We introduce a reward function that raises the yield of metastable, unique and novel structures (mSUN) from 13.4% for the pretrained model to 45.5% for the reinforced model, as evaluated by a community benchmark. Our reward also improves the performance of a reinforcement learning framework for crystalline materials based on latent denoising diffusion models. At the same time, we find that directly reinforcing atom-type transition likelihoods enables reward exploitation that has to be prevented with explicit guards. The same analysis also exposes a gap in the community metric. Single-element structures in distinct packings are counted as metastable, unique and novel materials and inflate mSUN without yielding any new compounds. A stability claim is only as good as its reference hull. We report every result split by the number of reference phases behind it and argue that benchmarks should do the same.
Problem

Research questions and friction points this paper is trying to address.

inverse materials design
crystal generation
reinforcement learning
reward hacking
evaluation metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group-Relative Policy Optimization (GRPO)
Discrete Flow Matching
Reinforcement Learning
Inverse Materials Design
Reward Hacking
🔎 Similar Papers
No similar papers found.