🤖 AI Summary
This study addresses the performance bottlenecks of standard alignment paradigms in personalized generation for large language models and the prohibitive scoring costs of large reward models. We propose a parameter-efficient candidate matching framework that reformulates personalized generation as a ranking task. Specifically, the method employs a million-parameter MLP to reuse internal embeddings from a base generator for precise matching, integrated with Best-of-N sampling to enable test-time scaling. Experiments demonstrate that this framework consistently outperforms billion-parameter general-purpose reward models across nine datasets while utilizing less than 0.4% of their parameters. Furthermore, it reduces scoring latency by four orders of magnitude, unlocking substantial potential for personalized generation at minimal computational overhead.
📝 Abstract
Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.