Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of overlooked linguistic feedback and uneven credit assignment in training LLM-as-a-Judge models by proposing a position-selective self-distillation framework. Specifically, the method analyzes token-level entropy variations between teacher and student models to precisely identify and mask high-entropy positions, thereby distinguishing memorization from semantic understanding patterns. This mechanism optimizes supervision signal density and fully leverages natural language feedback to enhance model generalization. Experimental results demonstrate that the proposed approach outperforms outcome-supervised reinforcement learning baselines by 2–9 percentage points on subjective subcategories. Furthermore, it maintains competitive performance on objective tasks while significantly improving out-of-distribution generalization capabilities.
📝 Abstract
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
Problem

Research questions and friction points this paper is trying to address.

LLM judges
language feedback
subjective tasks
outcome-supervised RL
self-distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Position-Selective Self-Distillation
Entropy Shift
LLM Judges
Natural Language Feedback
Position Masking
🔎 Similar Papers
2024-09-23arXiv.orgCitations: 15
2024-01-18International Conference on Machine LearningCitations: 264
💼 Related Jobs
No related jobs found.