🤖 AI Summary
This study addresses the granularity mismatch between user rating criteria and model generation guidance in existing LLM personalization alignment by proposing the GRASP framework. Through online policy self-distillation, GRASP converts coarse-grained user ratings into fine-grained token-level supervision signals. Furthermore, it introduces a novel criterion-based teacher verification mechanism to filter high-quality instances, combined with token distribution alignment techniques to enhance training efficiency. Experimental results demonstrate that GRASP achieves state-of-the-art performance on the LaMP-QA benchmark, validating the effectiveness of fine-grained supervision for LLM personalization.
📝 Abstract
LLM personalization aims to generate responses aligned with individual users'preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.