Score
Design, build, and analyze models and training procedures that produce a single predictive system to solve multiple related tasks at once, including architectures with shared encoders and separate task heads, factorized multi-head or graph-structured task components, and outputs for multi-task regression and classification. This includes engineering shared versus task-specific representations, crafting and weighting multi-task loss functions, developing optimization and fine-tuning strategies to prevent interference or collapse, and applying structure- or sparsity-aware approaches to improve joint performance.
In heterogeneous multi-task linear regression, tasks exhibit distinct yet partially overlapping sparsity patterns (support sets) and heterogeneous nonzero coefficient values, posing challenges for joint modeling. Method: We propose the first heterogeneous sparse multi-task learning (MTL) framework enabling independent sharing of support sets and nonzero coefficient values across tasks. Our approach introduces a novel mixed-integer programming (MIP)-based formulation and develops a customized algorithm integrating block coordinate descent, combinatorial local search, and exact optimization—ensuring both global optimality and scalability. Contribution/Results: We establish theoretical guarantees on enhanced variable selection consistency. Extensive simulations and two biomedical case studies demonstrate substantial performance gains over existing sparse MTL methods. The corresponding open-source R package, sMTL, is publicly available on CRAN.
Existing multivariate time series forecasting methods often neglect dynamic inter-variable dependencies, leading to modeling bias. To address this, we formulate multivariate forecasting as a systematic multi-task learning problem for the first time. Our approach introduces a gradient-geometry-based task partitioning and balancing mechanism: (i) task correlation is quantified via gradient angle analysis; (ii) correlation-driven hierarchical clustering groups variables into coherent tasks; and (iii) an error-adaptive gradient reweighting strategy ensures balanced optimization. We further propose MTLinear—a lightweight linear architecture that achieves strong expressiveness without sacrificing computational efficiency. Extensive experiments on multiple benchmark datasets demonstrate that our method consistently outperforms strong baselines—including Informer and Autoformer—in both accuracy and inference speed, achieving superior forecasting performance with significantly lower computational overhead. The implementation is publicly available.
To address the low training efficiency, poor generalization, and high computational overhead of multi-task neural solvers for combinatorial optimization problems (COPs), this paper proposes an efficient unified training paradigm. Methodologically, we design an encoder-decoder architecture that jointly integrates theory-driven loss decomposition with task-influence matrix modeling, and introduce a bandit-based bounded multi-task sampling algorithm to enable resource-constrained collaborative optimization and approximate gradient propagation. Experiments on TSPLib and CVRPLib benchmarks demonstrate that our approach significantly outperforms both single-task and state-of-the-art multi-task baselines under identical training budgets—achieving faster convergence and superior generalization across diverse COPs. The implementation is open-sourced and has been widely adopted by the community. This work establishes a scalable, theoretically grounded framework for multi-task learning in combinatorial optimization.
The scientific and engineering communities lack unified, reproducible benchmarks for evaluating AI/ML methods in dynamical systems modeling. Method: This paper introduces the Common Task Framework (CTF), a general-purpose framework targeting multiple scientific objectives—including prediction, state reconstruction, generalization, and control—under realistic constraints of limited data and noisy measurements. CTF establishes standardized datasets, objective evaluation metrics, and an open benchmarking platform. Contribution/Results: CTF enables the first cross-disciplinary, physics-constrained comparison of system identification and machine learning algorithms, facilitating rapid iterative development and integration. Experimental results demonstrate that CTF significantly improves model development efficiency and deployment reliability, thereby addressing a critical gap in AI evaluation frameworks oriented toward scientific discovery.
This work addresses the performance degradation commonly observed in merged multi-task models due to parameter interference, which often results in inferior performance compared to single-task experts. Existing dynamic routing approaches typically require additional training or prior knowledge of task identities, limiting their practicality. To overcome these limitations, the authors propose a training-free, task-ID-agnostic dynamic routing mechanism that leverages a few task-specific support samples to construct low-rank task manifolds via singular value decomposition (SVD). Routing decisions are made by evaluating the projection residuals of test samples onto these manifolds. The method seamlessly integrates with lightweight subspace- or mask-based merging strategies and demonstrates consistent performance gains across multiple computer vision and natural language processing benchmarks, effectively narrowing the gap with single-task expert models even when task identities are unknown at inference time.
This work addresses the challenge of constructing fully conformal prediction regions in multi-task regression, which are typically intractable due to their reliance on an infinite ensemble of predictors. The authors propose a computationally feasible approximation within a vector-valued reproducing kernel Hilbert space, providing theoretical guarantees of achieving the desired coverage level under both known and estimated covariance matrices. In the case of known covariance, they derive an upper bound on the volume of the prediction region and establish its tightness. Empirical evaluations on synthetic data demonstrate that the proposed method substantially outperforms split conformal prediction, yielding significantly smaller prediction regions while maintaining accurate coverage.
This study addresses the challenge in traditional multi-task learning where heterogeneous output types—such as continuous and binary responses—lead to incomparable loss functions, hindering effective information sharing. To overcome this, the authors propose a multi-task transformation framework that unifies diverse response variables through unknown monotonic transformations. The approach integrates a shared first-layer deep neural network with group Lasso regularization, enabling joint modeling in high-dimensional settings. Notably, it is the first to combine monotonic transformations with a shared sparse structure, establishing a unified multi-task learning framework suitable for mixed output types. Theoretical guarantees for consistent variable selection are provided. Empirical results demonstrate that the method significantly outperforms existing approaches on both simulated data and real-world gene expression analysis, successfully identifying biologically meaningful shared predictors.
This work addresses the challenge in multi-task linear regression where some tasks may be contaminated and the smallest eigenvalue of the task covariance matrix can approach zero. To tackle this, the authors propose an estimator based on matrix-weighted norm regularization, which introduces a relative balance condition that does not rely on a lower bound for the eigenvalues of individual task covariance matrices. This approach adaptively leverages task similarity, is robust to arbitrary outlier tasks, and avoids negative transfer. Theoretical analysis shows that under moderate balance, the prediction mean squared error achieves the best-known rate, and the overall MSE is minimax optimal up to logarithmic factors. Moreover, when tasks are unrelated or exhibit weak balance, the method performs no worse than learning each task independently.
This work addresses negative transfer in multi-task learning, attributing it to limited shared representation capacity and insufficient inter-task redundancy. The authors propose a Capacity–Redundancy (CR) identity that decomposes task prediction information into a label-redundant component and a residual coupling term, thereby revealing the mechanism of task interference. They further establish a theoretical link between gradient similarity and total correlation (TC), providing a principled basis for quantifying redundancy, and derive necessary and sufficient conditions for a cluster-wise global sharing structure. Building on these insights, they design a clustered LoRA architecture that effectively reduces residual coupling under a Gaussian multi-task model, significantly outperforming random partitioning and achieving statistically significant performance gains across multiple sub-experiments.