🤖 AI Summary
This study addresses the challenge of accommodating multi-dimensional personalization demands during test-time scaling of large language models by proposing the PersonTTS framework. Its core innovation lies in introducing a cross-user experience reuse mechanism to achieve amortized agent policy discovery. Specifically, the framework initializes controllers via demand matching and employs source distillation for process guidance, efficiently transferring previously learned search experiences to novel requirements and substantially reducing policy discovery overhead. Experimental results demonstrate that PersonTTS significantly outperforms existing methods on benchmarks such as AIME. It effectively enhances joint satisfaction for unseen user profiles while markedly decreasing both the temporal and computational costs associated with policy discovery.
📝 Abstract
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.