🤖 AI Summary
This study addresses the limited generative capacity of language models in constrained, counterintuitive creative tasks by proposing a masked diffusion model-based framework for the conditional generation of chess puzzles. The core innovation lies in a non-directional diffusion method that enables flexible conditioning on tactical themes and local board positions. Furthermore, the model is jointly optimized through an auxiliary best-move prediction task and Denoising Diffusion Policy Optimization (DDPO) reinforcement learning. Experimental results demonstrate that this approach improves puzzle solution uniqueness by 11.6% and theme-matching accuracy by 2.5%. Following reinforcement learning fine-tuning, the yield of valid positions increases by 89.1%, substantially enhancing structured creative generation under complex constraints.
📝 Abstract
While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.