COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

📅 2026-09-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出COBRA-Skills框架,通过结合上下文引导的优先级策略和基于证据的技能进化方法,有效降低了大型语言模型技能优化的成本和所需数据量。
📝 Abstract
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce \textbf{COBRA-Skills}, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
Problem

Research questions and friction points this paper is trying to address.

Large language model
skill optimization
execution-based evaluation
task data
Innovation

Methods, ideas, or system contributions that make the work stand out.

contextual bandit
skill evolution
budgeted sequential optimization
execution feedback
🔎 Similar Papers
No similar papers found.