🤖 AI Summary
This study addresses the persistent challenge of hallucinations in large language models during knowledge-intensive tasks, where maintaining factual fidelity often compromises creative generation. To reconcile this trade-off, we propose HARPO, a reinforcement learning framework that jointly optimizes faithfulness and creativity. Methodologically, we introduce a novel hallucination-aware generative reward model to provide verifiable feedback signals, design a selective activation mechanism to dynamically balance these dual objectives, and incorporate a data curriculum strategy for progressive training. Extensive experiments demonstrate that HARPO significantly reduces hallucination rates while enhancing creative writing quality across multi-scale models. Ultimately, this work achieves synergistic optimization of knowledge faithfulness and generative creativity, offering an effective solution for aligning large language models with both accuracy and expressiveness requirements.
📝 Abstract
Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.