Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing personalized reward models rely on in-context learning, struggling to capture deep correlations among user historical preference pairs. This work proposes P-TTT, a method that encodes preference relations into user-specific fast weights via test-time training. By introducing sequence-level updates and a preference alignment objective, P-TTT achieves weight adaptation through a single forward pass, eliminating the need for backpropagation during inference. Experimental results demonstrate that P-TTT significantly enhances the modeling of preference relations, substantially outperforming state-of-the-art methods. Ultimately, this approach establishes an efficient personalization paradigm for reinforcement learning based on large language models.
📝 Abstract
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.
Problem

Research questions and friction points this paper is trying to address.

Personalized Reward Modeling
In-Context Learning
Preference Relations
Reinforcement Learning from Human Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Training
Personalized Reward Modeling
Fast Weights
RLHF
Preference Alignment
🔎 Similar Papers
No similar papers found.