🤖 AI Summary
This work addresses the trade-off between performance and output diversity commonly observed in post-trained large language models. To reconcile this, the authors propose CreativeInstruct, an instruction-tuning approach that introduces a [StartCreativity] control token, enabling a single model to dynamically balance generation quality and creativity during inference. This method achieves flexible, on-the-fly control without requiring multiple specialized models. Additionally, the study introduces a novel structural diversity metric based on graph edit distance. Experimental results demonstrate that CreativeInstruct significantly enhances diversity in narrative generation without compromising quality, with 70.3% of human evaluators rating its outputs as more creative. When used as a base model for reinforcement learning, it also yields performance gains of approximately 4% and 5% on the AMC and MATH benchmarks, respectively.
📝 Abstract
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.