Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational costs of test-time training for large language models and the credit assignment challenges arising from coupled policy and implementation. We propose Guidance-TTT, a framework that introduces a novel guidance-execution decoupled architecture. Specifically, it freezes a large model to handle code implementation while exclusively training a small model to optimize high-level decision-making policies, thereby restricting test-time learning to short-horizon decision sequences. Furthermore, we design an adaptive group-relative reinforcement learning objective that leverages verifier feedback to enable online policy updates. Experimental results demonstrate that, in offline settings, this framework outperforms state-of-the-art methods across four major domains, including combinatorial optimization and machine learning, while significantly reducing computational overhead and enhancing exploration efficiency.
📝 Abstract
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce Guidance-TTT, which separates these roles. A compact guidance model is trained at test time to propose high-level strategic changes, while a frozen execution model implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Guidance-TTT produces strong solutions across four distinct domains: combinatorial optimization (Polyomino Packing), heuristic programming (AHC058), machine learning (Lasso), and GPU kernel optimization (TriMul). Across these tasks, it outperforms the best solutions reported in prior work while remaining competitive with state-of-the-art results on public online leaderboards. Code is available at https://github.com/Human-Agent-Society/reef/tree/guidance-ttt-support.
Problem

Research questions and friction points this paper is trying to address.

Open-ended scientific discovery
Test-time training
Credit assignment
Large language models
Computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-time training
Guidance-execution decoupling
Group-relative reinforcement learning
Open-ended scientific discovery
Small model adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chonghe Jiang
Massachusetts Institute of Technology
Ao Qu
Ao Qu
Massachusetts Institute of Technology
Language AgentMultisensory AIComputational Social Science
S
Siyuan Liu
Hong Kong Polytechnic University
R
Ruoyun Ma
ByteDance Inc.
Zijian Zhou
Zijian Zhou
National University of Singapore
statistical learninglarge language modelsmulti-agent machine learning
D
Dingyi Zhuang
Massachusetts Institute of Technology
B
Bo Liu
Stanford University
H
Han Zheng
Massachusetts Institute of Technology
Hanfei Yu
Hanfei Yu
Stevens Institute of Technology
Serverless ComputingLarge-Scale AI SystemsDistributed ML SystemsLLM Systems
Baichuan Mo
Baichuan Mo
PhD @ MIT, Research Scientist @ TikTok, Lyft
TransportationOptimizationMachine LearningDemand Modeling
J
Jinhua Zhao
Singapore-MIT Alliance for Research and Technology
P
Paul Pu Liang
Singapore-MIT Alliance for Research and Technology