OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of maintaining persona consistency in multi-turn role-playing, distribution shift in external distillation, and reward ambiguity in reinforcement learning by proposing an on-policy self-distillation framework. The method leverages information asymmetry to enable a single model to serve simultaneously as teacher and student. By identifying a bimodal structure in teacher confidence, it introduces a role-aware divergence switching mechanism and designs a progressive feature masking curriculum for effective knowledge internalization, eliminating the need for external teachers or reward models throughout training. Experimental results demonstrate that this approach significantly outperforms supervised fine-tuning and reinforcement learning baselines on benchmarks such as CharacterBench, effectively enhancing persona consistency in multi-turn dialogues.
📝 Abstract
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
Problem

Research questions and friction points this paper is trying to address.

persona consistency
role-playing language models
multi-turn dialogues
distribution mismatch
reward ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Persona Consistency
Role-Aware Divergence Switching
Progressive Trait Masking
Asymmetric Information
🔎 Similar Papers
No similar papers found.