EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in existing agent optimization where tools, decision-making, and models are tightly coupled, hindering the independent improvement of decision logic. To overcome this, it proposes an "evolution-internalization" bridging mechanism that validates novel decisions through program evolution and generates revised reasoning trajectories. Supervised fine-tuning (SFT) is then employed to transform explicit external instructions into the model's implicit reasoning capabilities, achieving decoupled decision optimization without reliance on external frameworks. Empirical results demonstrate that task success rates improve by 10.9 and 9.2 percentage points on in-domain and out-of-domain tasks, respectively, validating the method's cross-task generalizability and applicability across multiple models.
📝 Abstract
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a''potpourri''approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
Problem

Research questions and friction points this paper is trying to address.

Agent fine-tuning
Decision-making procedures
Internalization
Evolution
Reasoning traces
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Fine-Tuning
Decision-Making Internalization
Evolutionary Harness
Trace Rewriting
Generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shihan Dou
Shihan Dou
Fudan University
LLMsCode LMsRLAlignment
S
Shaofan Liu
Fudan University
Z
Zhonghang Lu
Fudan University
J
Jiahang Lin
Renmin University of China
Shichun Liu
Shichun Liu
Fudan University
NLP
B
Binghai Wang
Renmin University of China
Jiajie Jin
Jiajie Jin
Renmin University of China
Information RetrievalLarge Language Models
Guanting Dong
Guanting Dong
Remin University of China
LLM Reasoning & AlignmentDeep Search AgentAgentic RL
T
Tao Gui
Fudan University
Q
Qi Zhang
Renmin University of China
X
Xuanjing Huang
Fudan University