Self-Evaluating Recursive Agents

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge in recursive agent training where intermediate subtasks lack ground-truth annotations and external evaluation is prohibitively expensive. To overcome this, we propose the SERA framework, which reformulates evaluation as a learnable skill. Specifically, the policy endogenously generates weighted scoring criteria and integrates them with recursive decomposition, ranking-based objective functions, leaf coverage signals, and tree search to jointly optimize task decomposition, solving, and evaluation. This approach substantially reduces reliance on external supervision. Empirical results demonstrate that SERA achieves performance gains of 5.38 and 13.14 points on the TextCraft and TextWorld benchmarks, respectively, with an additional 2.43-point improvement during the inference phase.
πŸ“ Abstract
Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and carry no ground truth. Existing methods score each node with a verifier or judge, which is costly at scale and blind to decomposition quality. We argue that a recursive agent must learn three coupled capabilities within one set of weights: decomposing problems into subtasks, solving them, and evaluating the outcomes, each requiring its own training signal. SERA (Self-Evaluating Recursive Agents) turns evaluation into a learned capability of the policy itself. Before delegating, the parent writes a rubric of weighted success criteria for each child subtask; a ranking objective against verified outcomes then trains rubric generation so that the criteria track genuine subtask success. In addition, a complementary leaf-coverage signal provides direct credit for task decomposition. Our central finding is that \emph{training} the policy to generate aligned rubrics is what drives the gains: because the same weights both evaluate and execute, learning to judge subtasks sharpens the agent's ability to solve them. Notably, external supervision is also reduced: the judge is consulted only to train the rubric generator, while solving is trained against the agent's own rubric scores, which outperform direct use of the judge. Beyond training, the learned rubric doubles as an inference-time selector for tree search. On TextCraft-Synth and TextWorld-Sync, SERA improves over strong recursive-agent baselines by 5.38 and 13.14 points on average, and rubric-guided tree search at inference adds a further 2.43 points on TextWorld-Sync.
Problem

Research questions and friction points this paper is trying to address.

recursive agents
task decomposition
self-evaluation
training signal
intermediate subtasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evaluating Recursive Agents
Rubric Generation
Task Decomposition
Tree Search
Reduced External Supervision
πŸ”Ž Similar Papers