JIVEAdapter: A Multi-Task Additive Low-Rank Adapter via Joint and Individual Variation Explained (JIVE)

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of shared and task-specific signal entanglement and low parameter efficiency in multi-task fine-tuning by proposing a multi-task additive low-rank adapter grounded in the statistical principles of Joint and Individual Variation Explained (JIVE). The method introduces a novel additive decomposition mechanism that leverages orthogonality constraints to decouple weight updates into a shared joint structure and independent individual structures, combined with adaptive rank allocation for efficient knowledge reuse. Evaluated on the GLUE benchmark, the proposed approach achieves strong baseline performance at an equivalent effective rank. Furthermore, it enables incremental adaptation to new tasks without requiring additional modules, thereby significantly reducing computational costs.
📝 Abstract
Parameter-efficient fine-tuning adapts pretrained models at a fraction of the cost of full fine-tuning, yet most low-rank adapters are single-task and represent each weight update multiplicatively, leaving no explicit account of what is shared across tasks and what is task-specific. We introduce JIVEAdapter, a multi-task "additive" low-rank adapter inspired by statistical Joint and Individual Variation Explained (JIVE). JIVEAdapter decomposes every weight update into a Joint structure shared across all tasks plus a per-task Individual structure, penalizes the Individual structures to be near-orthogonal to the Joint so shared and task-specific signal stay "interpretable" and separated, and allocates rank adaptively across a shared Joint pool and a per-task Individual pool. The Joint is learned once, jointly over a task group or incrementally, one task at a time, then frozen and reused as a prior for new tasks without retraining the shared part. On GLUE and SuperGLUE with DeBERTaV3-base, JIVEAdapter is competitive with strong single-task and multi-task low-rank baselines at a matched per-task effective rank, without extra modules such as MoE, and when a related held-in task exists its frozen Joint serves a held-out task by reusing that task's Individual with only a cheap per-direction scale, otherwise training a small new one.
Problem

Research questions and friction points this paper is trying to address.

parameter-efficient fine-tuning
low-rank adapter
multi-task learning
joint and individual variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameter-efficient fine-tuning
Low-rank adapter
Multi-task learning
JIVE
Additive decomposition
🔎 Similar Papers
No similar papers found.