Score
Designs and implements a metric and computation pipeline that quantifies transfer effectiveness by adjusting observed target performance for source accuracy and example/task difficulty; builds the hardness‑adjusted transfer (HAT) score and related per‑task and aggregate computations to calibrate target results, reveal true transfer, and enable fair model benchmarking.
This study addresses the **reliable assessment of knowledge transferability** in transfer learning—a longstanding challenge hindered by inconsistent evaluation criteria, poor interpretability, and ill-defined applicability scopes. We propose the first **two-dimensional classification framework**, systematically organizing over 60 mainstream transferability metrics along axes of *transferable knowledge type* (e.g., features, relations, semantics) and *measurement granularity* (sample-, task-, or domain-level), while rigorously reconstructing their mathematical foundations, underlying assumptions, and failure boundaries. Through cross-modal and cross-task empirical analysis, we characterize the efficacy gradients and root limitations of metrics across paradigms (e.g., pretraining-finetuning). Our work establishes a standardized assessment pathway and principled metric selection guidelines for transferability evaluation, advancing trustworthy AI evaluation infrastructure, and identifying key future directions—including dynamic transferability modeling and causally grounded metrics.
This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.
Existing evaluation methods struggle to disentangle overall performance gains in source languages from genuine cross-lingual transfer capabilities in multilingual models. To address this limitation, this work proposes the Hardness-Adjusted Transfer (HAT) score, which isolates source-language performance to more accurately quantify transfer effectiveness from high-resource to low-resource languages. Leveraging HAT, we conduct a large-scale empirical analysis across 20 language models and three major multilingual benchmarks, revealing—for the first time—that small models retain meaningful transfer capacity, that scaling model size yields diminishing returns in transfer gains, and that overall cross-lingual transfer capability has steadily improved over time.
Current transferability estimation benchmarks suffer from fundamental flaws—namely, unrealistic fixed model spaces and static performance hierarchies—which severely distort evaluation outcomes; simple, dataset-agnostic heuristics frequently outperform sophisticated metrics, exposing a critical mismatch between benchmark protocols and real-world model selection scenarios. Method: The authors conduct a systematic empirical re-evaluation of mainstream transferability metrics across diverse, realistic model spaces and dynamically varying performance rankings. Contribution/Results: They quantitatively identify and characterize the primary sources of benchmark bias for the first time. Crucially, they demonstrate that the prevailing evaluation paradigm is unreliable, propose a novel benchmarking framework that is realistic, dynamic, and task-aware, and provide both theoretical foundations and practical guidelines for designing robust transferability assessment systems.
Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.
This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.
研究通过多种校准审计方法,解决了早期结果预测器在不同代理上的校准转移问题,发现特定目标组合存在持续的校准转移错误。
研究通过将任务参数化应用中的特定计算重新表述为适合并行执行的形式,解决了进化迁移优化(ETO)在扩展到更大任务集合时的评估效率问题。
This study addresses the limitations of existing transferability estimation metrics in medical image transfer learning, which are predominantly designed for natural images and struggle with class imbalance and instability across different experimental settings. For the first time, this work systematically evaluates the robustness of these metrics in medical imaging contexts by constructing miniature target datasets with varying sample sizes and multiple random seeds to isolate perturbations in target data. Through comprehensive comparisons involving diverse transferability estimation methods and classification evaluation metrics that account for class imbalance, the study reveals that minor variations in target data or choices of evaluation metrics can substantially alter the ranking of source models. Consequently, current transferability estimators exhibit consistently low agreement with actual performance rankings, highlighting their inadequacy in medical applications.
本文提出两种技术,通过增加推理时间和模拟真实部署环境来提高对齐评估的真实性,解决模型在测试与实际部署中表现差异的问题。
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.