Capacity and Redundancy Trade-offs in Multi-Task Learning

📅 2026-07-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses negative transfer in multi-task learning, attributing it to limited shared representation capacity and insufficient inter-task redundancy. The authors propose a Capacity–Redundancy (CR) identity that decomposes task prediction information into a label-redundant component and a residual coupling term, thereby revealing the mechanism of task interference. They further establish a theoretical link between gradient similarity and total correlation (TC), providing a principled basis for quantifying redundancy, and derive necessary and sufficient conditions for a cluster-wise global sharing structure. Building on these insights, they design a clustered LoRA architecture that effectively reduces residual coupling under a Gaussian multi-task model, significantly outperforming random partitioning and achieving statistically significant performance gains across multiple sub-experiments.
📝 Abstract
In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling $Δ$ from validation residual correlations, showing that clustered LoRA substantially reduces $\widehatΔ$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.
Problem

Research questions and friction points this paper is trying to address.

multi-task learning
negative transfer
capacity
redundancy
shared representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Capacity–Redundancy identity
total correlation
residual coupling
clustered sharing
gradient cosine similarity
🔎 Similar Papers
A
Asif Khan
Harvard Medical School, Boston, USA