🤖 AI Summary
This study addresses the performance degradation caused by indiscriminate skill invocation in AI agents, highlighting the need to predict actual skill gains for specific tasks. We propose SkillDelta, a framework that predicts task-conditioned skill benefits from pairwise execution histories to enable selective invocation. The core innovation lies in establishing an error analysis theory under explicit transfer assumptions, where local predictors model support coverage and representation mismatch to transfer historical gains to novel tasks without retraining. Evaluated across 15 experimental settings, our approach significantly outperforms random activation strategies, improving average success rates by 4.3% while effectively optimizing computational resource allocation.
📝 Abstract
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-conditional skill gains from paired executions of the same agent with and without the skill. A local predictor transfers these historical gains to new tasks without retraining the agent. Under explicit transfer assumptions, our analysis links support coverage, representation mismatch, and execution noise to prediction error and decision regret. Across five benchmarks and three target agents, paired history improves observed-gain ranking over skill-assisted outcomes alone in 12 of 15 settings. At matched expected skill-use rates, SkillDelta improves success over random activation in all 15 settings, with an average absolute gain of 4.3%. Most of this advantage comes from allocation across task groups. Evidence for additional within-group selection value is strongest on ToolQA and weaker elsewhere. Code is available at https://github.com/TankTechnology/skilldelta.