Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of supervision signals caused by the unavailability of specialized testing frameworks during deployment. To overcome this limitation, we propose a recursive self-rewriting framework built upon the Qwen-3.8-27B model. By orchestrating the collaboration of a planner, a critic, and an executor, our approach reconstructs problem-solving experiences from multiple frameworks into generalized training trajectories. Furthermore, reliability is ensured through runbook extraction, leakage screening, and sandboxed execution techniques. Experimental results demonstrate that the proposed method improves the task resolution rate by 34.3% and significantly outperforms baselines on the Pass@3 metric across benchmarks such as Terminal-Bench. These findings confirm that our framework effectively enables the reuse and generalization of cross-framework experiences in practical deployment scenarios.
📝 Abstract
Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.
Problem

Research questions and friction points this paper is trying to address.

complex tasks
trajectory scaling
deployment gap
supervised finetuning
harness generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Rewrite
Trajectory Reconstruction
Supervised Fine-Tuning
Multi-Harness Collaboration
Runbook Extraction
🔎 Similar Papers
2024-07-16Neural Information Processing SystemsCitations: 16