🤖 AI Summary
This work addresses negative transfer in multi-task learning, attributing it to limited shared representation capacity and insufficient inter-task redundancy. The authors propose a Capacity–Redundancy (CR) identity that decomposes task prediction information into a label-redundant component and a residual coupling term, thereby revealing the mechanism of task interference. They further establish a theoretical link between gradient similarity and total correlation (TC), providing a principled basis for quantifying redundancy, and derive necessary and sufficient conditions for a cluster-wise global sharing structure. Building on these insights, they design a clustered LoRA architecture that effectively reduces residual coupling under a Gaussian multi-task model, significantly outperforming random partitioning and achieving statistically significant performance gains across multiple sub-experiments.
📝 Abstract
In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling $Δ$ from validation residual correlations, showing that clustered LoRA substantially reduces $\widehatΔ$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.