Score
Designs and builds assessment frameworks and metrics to measure the quality and performance of skill nodes, aggregating independent model evaluations and producing robust skill-evaluation scores. Analyzes cross-model transferability and constructs transferability-scoring methods to verify and select effective skill nodes.
This study addresses the **reliable assessment of knowledge transferability** in transfer learning—a longstanding challenge hindered by inconsistent evaluation criteria, poor interpretability, and ill-defined applicability scopes. We propose the first **two-dimensional classification framework**, systematically organizing over 60 mainstream transferability metrics along axes of *transferable knowledge type* (e.g., features, relations, semantics) and *measurement granularity* (sample-, task-, or domain-level), while rigorously reconstructing their mathematical foundations, underlying assumptions, and failure boundaries. Through cross-modal and cross-task empirical analysis, we characterize the efficacy gradients and root limitations of metrics across paradigms (e.g., pretraining-finetuning). Our work establishes a standardized assessment pathway and principled metric selection guidelines for transferability evaluation, advancing trustworthy AI evaluation infrastructure, and identifying key future directions—including dynamic transferability modeling and causally grounded metrics.
This study addresses two core challenges in transfer learning: quantitative assessment of knowledge transferability and assurance of trustworthiness. First, it systematically formalizes transfer learning from a trustworthiness perspective, proposing a novel “transferability–trustworthiness” co-evaluation framework; theoretically characterizes transferability bounds under non-IID settings; and develops a new transfer paradigm incorporating multi-dimensional trust constraints—privacy, robustness, and fairness. Methodologically, it integrates statistical learning theory, adversarial robustness analysis, fairness metrics, differential privacy mechanisms, and mainstream transfer approaches (e.g., domain adaptation, meta-transfer learning, federated transfer learning). Key contributions include: (1) establishing the first holistic framework spanning theoretical modeling, quantitative trust attribute measurement, and empirical validation; and (2) identifying three open research directions—trustworthy non-IID transfer, standardized trustworthiness benchmarks, and cross-domain causal generalization.
This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.
Current transferability estimation benchmarks suffer from fundamental flaws—namely, unrealistic fixed model spaces and static performance hierarchies—which severely distort evaluation outcomes; simple, dataset-agnostic heuristics frequently outperform sophisticated metrics, exposing a critical mismatch between benchmark protocols and real-world model selection scenarios. Method: The authors conduct a systematic empirical re-evaluation of mainstream transferability metrics across diverse, realistic model spaces and dynamically varying performance rankings. Contribution/Results: They quantitatively identify and characterize the primary sources of benchmark bias for the first time. Crucially, they demonstrate that the prevailing evaluation paradigm is unreliable, propose a novel benchmarking framework that is realistic, dynamic, and task-aware, and provide both theoretical foundations and practical guidelines for designing robust transferability assessment systems.
This work addresses the lack of systematic evaluation of agent skills' practical utility in cross-domain, multi-model settings. We propose the first scalable evaluation framework specifically designed for agent skills, enabling skill developers to define custom tasks and evaluation dimensions grounded in real-world scenarios. The framework quantifies the enhancement provided by skills to large language model (LLM) agents through metrics of instruction following and task completion. It integrates automated task generation, scoring rules, and comparative experiments across 19 commercial and open-source LLMs, yielding a benchmark comprising 1,000 diverse tasks. Experimental results reveal significant disparities among models in adhering to skill instructions, thereby validating the effectiveness of skills in guiding agent behavior. The benchmark dataset is publicly released to foster further research in this area.
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
This study addresses the uncertainty and evaluation challenges associated with skill transfer in multi-agent systems by proposing Evo2Team, a framework that optimizes target team deployment through the selection, adaptation, and validation of source skills. It innovatively introduces a multi-objective function encompassing joint quality, cost, and model hierarchy, while emphasizing an execution-based transfer evaluation mechanism grounded in actual agent behaviors. The proposed method leverages GPT/Qwen model cascades, Count-Frequency and AgentsNet environments, and evolutionary algorithms for optimization. Experimental results demonstrate that Evo2Team effectively reduces exploration costs across most scenarios, significantly enhances task performance, and lowers deployment overhead.
研究探讨了手术技能模型在不同评估标准下的可迁移性问题,通过多种方法如端到端训练、ASAM及自监督学习等进行分析。
This study addresses the challenge of filtering erroneous and non-transferable knowledge from skill libraries of self-evolving agents by proposing a Proposer-Builder-Verifier architecture to validate skill reusability in unseen tasks. The method dynamically synthesizes test scenarios through conditional constraint generation, overcoming the limitations of traditional static evaluation. Furthermore, it employs an automated execution comparison mechanism to quantitatively assess skill utility and efficiency, enabling reliable retention or rejection decisions. Experimental results demonstrate that the proposed framework significantly improves downstream task performance and execution efficiency on the ALFWorld and WebShop benchmarks, while also confirming the accuracy of its skill reusability assessment.
研究通过对比任务级与子任务级技能诱导及文本与代码格式,解决LLM代理技能转移不可靠问题,提出技能效用评分以预测任务成功。
This study addresses the challenge of skill acquisition for LLM-based agents in industrial planning, which is hindered by heterogeneous and incomplete multi-source knowledge. To this end, we propose a cross-source skill induction and execution verification framework. The method extracts coordination patterns from discrepancies between documentation and practice, compiles them into standardized skill packages anchored by COM specifications, and achieves iterative refinement through unlabeled compliance screening and signal-conditioned trajectory attribution. Evaluated on the PIMS-Bench benchmark across four LLM backbones, the proposed approach yields absolute improvements of 14%–30% in component matching F1 scores, with particularly significant gains observed on complex tasks.
This study addresses the disconnect between agent skill relevance and task utility, along with the lack of mechanistic explanations, through an empirical investigation across 87 tasks. Methodologically, we define a downstream utility metric based on pass-rate differentials and employ both large language models and human review to analyze skill content, execution trajectories, and artifacts. Furthermore, we propose a re-ranking strategy supporting essential operations alongside a DAG-based dependency organization method. Results reveal that in 36.78% of tasks, identical skills exhibit opposite utility under different configurations. The proposed re-ranking strategy improves preferred pass rates by 4.35 to 5.80 percentage points. Additionally, this work distills 17 practical guidelines for effective skill authoring.