Evaluating Agent Skills for Version-Specific Plugin Migration: A Retrospective Study

📅 2026-09-24
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This study addresses the limitation that skill evaluations of coding agents during plugin migration frequently rely on diagnostic scores, making it difficult to verify whether agents genuinely satisfy target-version contracts. To overcome this, we propose a traceable evaluation framework that conducts retrospective studies using static migration archives to link aggregated rewards with contract-level evidence. By integrating task-guided statistics, execution probes, and multi-model independent rescoring with weighted Kappa analysis, the approach systematically reveals scoring biases. Our findings demonstrate that although skill enhancements increase average rewards, such gains are concentrated and accompanied by scoring deficiencies. Nevertheless, corrected estimates confirm positive benefits, and multi-judge rescoring exhibits high agreement, thereby providing reliable checkpoints for auditing agent capabilities.
📝 Abstract
Agent skills package version-specific maintenance knowledge for coding agents, but a higher diagnostic score does not by itself show that the resulting migration advice satisfies the target version's contract. We study a shipped plugin-upgrade skill through an archive of 64 reports on 16 static migration tasks, with two attempts per condition and 328 criterion decisions. With the skill, mean recorded reward rises from 93.83 to 98.75, a gain of 4.92 points (95% task-bootstrap interval [0.31, 10.86]); the gain is concentrated in one task, and eight task pairs are at the ceiling. Tracing every decision to its contract domain and reviewing ten reports in depth exposes grading errors that favor either arm; in one, a containment predicate that accepts the parent directory still receives full credit. Executable probes confirm this defect and show that a working teardown repair is excluded only by a narrower lifecycle rubric. Replacing the reviewed decisions keeps the estimate positive (4.61 to 5.39 points) but moves its interval to or across zero. Re-grading all 64 reports with judges from two other model families, without arm labels or prior scores, agrees with the original judge on 91.8% and 95.7% of decisions (weighted $Îș=0.64$ and $0.72$) and gives gains of 10.63 and 6.09 points. The study contributes a traceable evaluation that connects aggregate reward to contract-level evidence and judge sensitivity, together with concrete review checks for migration advice. Executable end-to-end repairs, independent human annotation, and other frameworks are left to future work.
Problem

Research questions and friction points this paper is trying to address.

Agent skills evaluation
Plugin migration
Version-specific maintenance
Contract compliance
Reward grading reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Skills
Plugin Migration
Contract-level Evaluation
Executable Probes
Traceable Evaluation
🔎 Similar Papers
No similar papers found.
đŸ’Œ Related Jobs
No related jobs found.
B
Beiming Liu
Tsinghua University
H
Haihao Li
Fudan University
Minjie Chen
Minjie Chen
PetroChina Southwest Oil & Gasfield Company
N
Ning Chen
Jilin University
Y
Yiran Wang
The Frederick Gunn School
J
Jiming Ye
Shenzhen University
Puzhao Zhang
Puzhao Zhang
Dalian Neusoft University of Information
T
Tongtao Wang
Independent Researcher
Sheng Gao
Sheng Gao
Beijing University of Posts and Telecommunications
Machine Learning and Data mining
W
William Jin
Independent Researcher
W
Weihao Mu
Dalian University of Technology, Jiyin Zhiyuan (Shanghai) Technology Co., Ltd.
Chengzhi Liu
Chengzhi Liu
PhD, UC Santa Barbara
Vison Language ModelTruthworthy AIReasoning
Y
Yucheng Xia
Harbin Engineering University
G
Guangren Wang
Alibaba Cloud
C
Chaoyang Fan
Jianghan University, Jiuxiangxian (Beijing) Technology Co., Ltd.
C
Changfeng Huang
Sun Yat-sen University
X
Xunming Lin
Great Bay University
Y
Yuanjie Shen
Beijing Information Science and Technology University