Agentic-TTT: Training test-time policy for test-time training

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited applicability of existing test-time training (TTT) methods, whose indiscriminate deployment often incurs unnecessary computational overhead or performance degradation due to the absence of adaptive selection mechanisms. To overcome this limitation, this work introduces an agent-based paradigm that governs TTT by encapsulating it as a callable tool within a dynamic environment constructed from accumulated skills. A policy network is trained via reinforcement learning to autonomously determine when and how to execute parameter updates and self-improvement, guided by actual utility gains. The proposed approach effectively balances performance enhancement against computational cost, nearly doubling the utility relative to the base model on standard benchmarks while demonstrating strong generalization to unseen domains.
📝 Abstract
Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.
Problem

Research questions and friction points this paper is trying to address.

Test-time training
Large language models
Self-improvement
Agency
Compute efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-time training
Agentic-TTT
Test-time policy
Autonomous self-improvement
Tool use
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jiahao Lu
Jiahao Lu
National University of Singapore
AGI risksAI controlAI safetyAI alignmentAI security
M
Mohan Kankanhalli
NUS AI Institute, National University of Singapore