MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a systematic evaluation framework for agent skills and the challenge of identifying high-potential skills prior to execution. To this end, we propose MCRI, an information-gain-based four-dimensional theoretical framework, coupled with a large language model-driven pre-evaluation method. By quantifying the information gain and behavioral constraints of skills, this approach establishes a novel paradigm for prioritizing skills without requiring actual execution. Experimental validation on the OpenClaw skill repository and multiple benchmarks, including BigCodeBench, demonstrates that the proposed method significantly improves skill selection accuracy, elevating downstream performance ranking percentiles by nearly 20%. These results confirm its effectiveness as a systematic solution for the efficient screening and selection of agent skills.
📝 Abstract
As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
Problem

Research questions and friction points this paper is trying to address.

agent skills
skill evaluation
analysis framework
modular agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

MCRI Framework
Skill Evaluation
Information Gain
Behavioral Constraint
LLM-based Assessment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.