Mining Agent Skills from Production Traces

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively mining skills from agent trajectories in production environments where reliable execution feedback is unavailable. It systematically investigates how sampling strategies, feedback signals, and skill representations—specifically workflow planning versus declarative ontologies—affect downstream task performance through comparative experiments on two enterprise-level benchmarks. The findings reveal that distinct domains require customized meta-skills, thereby undermining the viability of universal approaches. Experimental results demonstrate that ThinkingBox favors labeled workflow skills, whereas APEX prefers ontological representations without exhibiting significant evidence-based preferences. These observations confirm that skill-mining strategies are highly contingent upon task-specific structural constraints.
📝 Abstract
Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the sampling of execution traces, access to success or failure information, and the form of the mined skills affect downstream task performance. Holding the mining pipeline fixed, we compare six combinations of mining evidence and skill forms. Mining evidence has three levels: successful trajectories only, successes and failures with their outcome labels, or the same mix with labels withheld. Skill form has two types: an ordered workflow plan, or a declarative ontology of entities, states, and policies. We evaluate the mined skills on two enterprise benchmarks, ThinkingBox-Bench and APEX-Agents. Analysis of task-level paired differences shows that the benefits of different configurations of mining evidence and skill forms depend on the enterprise domain. On ThinkingBox-Bench, paired differences show that workflows score better than ontology by 1.7 pp, Goldilocks beats success-only evidence type by 2.4 pp and Goldilocks blind simulating skills learnt without outcomes is worse by 3.1 pp. APEX-Agents shows a moderate preference for ontologies and no clear preference between evidence regimes. Within each domain, task structure related constraints drive uneven performance with mined skills. These findings motivate tailoring meta-skills to the demands of the target tasks rather than adopting a one-size-fits-all approach.
Problem

Research questions and friction points this paper is trying to address.

Agent skill mining
Execution traces
Production environments
Task performance
Skill representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Skill Mining
Execution Traces
Workflow Plan
Declarative Ontology
Production Environments
🔎 Similar Papers
Y
Yue Ran Kang
Massachusetts Institute of Technology
C
Colton Mikolajczyk
Massachusetts Institute of Technology
C
Chhaya Methani
Microsoft Corporation
H
Hazel Mak
Microsoft Corporation
S
Sahil Bhatnagar
Microsoft Corporation
S
Susheel Suresh
Microsoft Corporation
A
Alejandro Gutierrez Munoz
Microsoft Corporation