🤖 AI Summary
This study addresses the challenge in agent self-improvement where variations in initial states hinder the accurate evaluation of meta-skill revisions. To overcome this, we propose Hindsight Meta-Experience Distillation (HMED), a mechanism that isolates variables through state restoration and re-execution to construct reusable meta-experience records. Crucially, this work pioneers a paradigm shift by redirecting the learning focus from branching outcomes toward the causal effects of the improvement process itself, leveraging contrastive analysis and structured distillation to optimize skill discovery. Extensive evaluations across multiple interactive benchmarks demonstrate that HMED significantly enhances both the meta-skill learning capabilities and overall task performance of open-source and closed-source models.
📝 Abstract
As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.