🤖 AI Summary
This study addresses the challenges faced by long-horizon agents in determining success or failure, safely retrying, and annotating experiences without reward feedback. To this end, we propose SelfSuite, a framework that leverages large language model bootstrapping to enable agents to autonomously construct evaluation suites from publicly available materials, generating weighted judges and task briefings. By integrating a gated retry mechanism with a typed memory bank, our approach innovatively achieves self-evaluation and experience management under zero-label conditions. Extensive experiments on the tau2-bench and AppWorld benchmarks demonstrate that SelfSuite significantly outperforms unlabeled baselines, with certain metrics matching few-shot expert-annotated performance.
📝 Abstract
A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.