cross-task transfer analysis

Designs and implements evaluation protocols and quantitative analyses that measure how capabilities learned on one task transfer to other tasks or task families; constructs experiments and metrics to assess cross-task transferability and cross-family evaluation while analyzing factors that affect transfer such as model architecture, backbone sharing, and other design choices.

cross-tasktransferanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Benchmarking Transferability: A Framework for Fair and Robust Evaluation

Apr 28, 2025
AK
Alireza Kazemi
🏛️ The University of Queensland

This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.

Addressing inconsistencies in transferability measurement methodsEvaluating reliability of transferability scores across domainsProposing standardized framework for robust transferability assessment

When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

Jun 05, 2025
QS
Quan Shi
🏛️ Princeton Language and Intelligence | OpenAI | Stanford University

This study investigates whether improvements in AI models’ reasoning capabilities naturally enhance their effectiveness as teachers—specifically, their ability to convey understandable and transferable knowledge to humans in human-AI collaboration. Method: We propose KITE, a novel evaluation framework that quantifies the causal impact of model explanations on humans’ subsequent independent problem-solving performance via a two-stage behavioral experiment (N=118). Contribution/Results: We provide the first systematic definition and empirical measurement of human-AI knowledge transfer efficacy, revealing only weak correlation between standard model benchmark performance and actual knowledge transfer—evidencing a “knowledge-rich but explanation-poor” phenomenon. We identify key behavioral strategy factors governing successful knowledge transfer and demonstrate that explainability and pedagogical alignment must be explicitly optimized. To support reproducibility and further research, we open-source the KITE toolkit—including implementation code, annotated datasets, and standardized evaluation protocols.

Assessing model explanation impact on human understandingIdentifying factors for successful knowledge transfer optimizationMeasuring AI-human knowledge transfer effectiveness

Crossover Designs in Software Engineering Experiments: Review of the State of Analysis

Aug 14, 2024
JF
Julian Frattini
🏛️ Blekinge Institute of Technology | Universidad Politécnica de Madrid

This study identifies the long-standing neglect of validity threats—particularly carryover effects—in crossover-design experiments within software engineering (SE). Method: Following Vegas et al.’s guidelines, we conducted the first quantitative assessment of 67 crossover-design experiments reported in 136 SE papers (2015–2024), employing forward snowball sampling and systematic content coding. Contribution/Results: Only 29.5% of validity threats were adequately addressed, and carryover effects were explicitly modeled in a mere 3% of studies. Although overall analytical validity has improved compared to earlier periods, practical adherence remains critically insufficient. Crucially, this work provides the first empirical evidence that low guideline adoption is a primary bottleneck. We propose a novel, taxonomy-based framework for classifying and assessing validity threats, offering actionable pathways and empirical grounding to enhance methodological rigor in SE experimentation.

Accuracy of Experimental ResultsCross-Design ExperimentData Analysis Quality

This work addresses the limited goal-directed execution capability of large language models in long-horizon tasks by introducing a Goal-Directed Execution (GDE) behavioral framework. The authors conduct post-training on the Qwen3.5-122B-A10B model using 363 long-horizon, multi-tool agent tasks from office scenarios, without relying on software engineering data. This approach yields a notable improvement on SWE-Bench Pro, increasing pass@1 by 5.8 percentage points. Experimental results demonstrate significant enhancements across four core GDE capabilities: goal selection, state construction, goal consistency maintenance, and environment validation. Furthermore, the model exhibits effective cross-domain transfer between office and software engineering tasks, confirming that long-horizon post-training can successfully drive the transfer of behavioral mechanisms.

behavioral generalizationcross-domain transfergoal-directed execution

Inferring Capabilities from Task Performance with Bayesian Triangulation

Sep 21, 2023
JB
John Burden
🏛️ University of Cambridge | The Alan Turing Institute | Universitat Politècnica de València

Addressing the challenge of reliably inferring AI systems’ cognitive capabilities from heterogeneous, few-shot task performance, this paper proposes a Bayesian triangulation framework for cognitive profiling. The method introduces a “measurement layout” generative model (implemented in PyMC) that jointly models task-instance features, latent capability dimensions, and system responses—thereby overcoming traditional psychometric reliance on large-scale, homogeneous datasets. Its key innovation lies in the first integration of Bayesian latent-variable modeling with multi-task cross-validation, enabling individualized, architecture-agnostic cognitive capability inversion. Evaluated on the AnimalAI Olympics benchmark (68 competing agents) and the O-PIAAGETS benchmark (30 synthetic agents), the framework successfully reconstructs fine-grained cognitive profiles, significantly enhancing discriminability and interpretability of inferred capabilities. Results empirically validate the feasibility and effectiveness of capability-oriented evaluation as a principled alternative to conventional behavioral benchmarks.

Enable capability inference from non-populational data using Bayesian methodsInfer cognitive profiles from diverse experimental task performance dataModel task-feature and capability interactions affecting system performance

Latest Papers

What's happening recently
View more

This study addresses the inference bias that arises when organizations evaluate expert competence solely based on project success or failure, a distortion attributable to differences in task bundling architectures. Drawing on Bayesian inference and Blackwell’s partial order theory, this work compares the informational value of bundled, outsourced, and unbundled projects, derives reliability thresholds, and quantifies the statistical costs associated with coarse-grained aggregation. The primary contribution lies in establishing, for the first time, a precise threshold relationship between task architecture and learning efficiency, revealing that under fixed workloads, the advantage of bundling strengthens as task scope expands. Furthermore, it demonstrates that bundling dominates when external technologies are unreliable, and that interim auditing can significantly broaden its region of superiority.

Blackwell orderbundlingcoarse performance

This work addresses systematic limitations in existing creative quality alignment (CQA) datasets, particularly their inadequate modeling of audience preferences and insufficient coverage of real-world logical constraints. To overcome these issues under stringent engineering and data scarcity conditions, the authors propose a low-resource CQA approach that leverages only around one hundred expert-annotated chain-of-thought (CoT) examples. By uncovering a dual mechanism between appreciation and generation tasks within conditional generative architectures, the method enables automatic transfer of calibrated knowledge from the appreciation module to the generation module. Experimental results demonstrate that the proposed framework substantially mitigates the shortcomings of current datasets and validates the practical feasibility of aligning generative models with nuanced creative quality metrics in real-world engineering settings.

Alignment Dataset BiasCalibrated SurpriseChain-of-Thought Fine-Tuning

This study addresses the current lack of human-centered, interpretable, and responsible evaluation criteria for AI in modeling and simulation. The authors propose the first multidimensional benchmark framework specifically designed to assess large language models (LLMs) through a human-centric lens, leveraging an open-source system dynamics AI platform to systematically evaluate performance across qualitative modeling, quantitative modeling, and model discussion tasks—emphasizing human-AI collaboration rather than replacement. The framework incorporates critical capabilities such as causal reasoning, iterative model refinement, and behavioral explanation, while embedding ethical and accountability considerations. Empirical results indicate that existing AI tools perform relatively well in qualitative tasks and model discussions but remain limited in causal reasoning and quantitative error correction; furthermore, different LLMs exhibit distinct strengths, with no single model emerging as universally superior.

AI for Modeling and SimulationBenchmarkingBias in AI

Hot Scholars

CX

Chenghao Xiao

Durham University
Natural Language ProcessingInformation RetrievalRepresentation Learning
KE

Kenneth Enevoldsen

Post-doc, Aarhus University
representation learningnatural language processingcognitive scienceelectronic health records
NM

Niklas Muennighoff

Stanford University
large language modelsartificial intelligencemachine learning
IC

Isaac Chung

Zendesk
Machine LearningComputer VisionNatural Language Processing
PM

Prasenjit Mitra

Research Professor, CMU-Africa and Department of ECE, CMU, Guest Professor, Leibniz Univ. Hannover
Machine LearningMedical InformaticsHuman Computer InteractionNatural Lang. Process.