Evaluating Agent Skills for Version-Specific Plugin Migration: A Retrospective Study
This study addresses the limitation that skill evaluations of coding agents during plugin migration frequently rely on diagnostic scores, making it difficult to verify whether agents genuinely satisfy target-version contracts. To overcome this, we propose a traceable evaluation framework that conducts retrospective studies using static migration archives to link aggregated rewards with contract-level evidence. By integrating task-guided statistics, execution probes, and multi-model independent rescoring with weighted Kappa analysis, the approach systematically reveals scoring biases. Our findings demonstrate that although skill enhancements increase average rewards, such gains are concentrated and accompanied by scoring deficiencies. Nevertheless, corrected estimates confirm positive benefits, and multi-judge rescoring exhibits high agreement, thereby providing reliable checkpoints for auditing agent capabilities.