Where Should LoRA Go? Component-Type Placement in Hybrid Language Models

📅 2026-04-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency and performance degradation caused by standard LoRA’s uniform adaptation across heterogeneous components in mixture-of-experts language models, which overlooks the functional distinctions between attention and recurrent modules. The authors propose a component-aware LoRA adaptation strategy that differentially deploys adapters in sequential (Qwen3.5-0.8B) and parallel (Falcon-H1-0.5B) hybrid architectures. Experimental results demonstrate that adapting only the attention pathways achieves superior performance over full fine-tuning with 5–10× fewer parameters. Moreover, adapter application to recurrent components proves detrimental in sequential architectures—causing a 14.8% performance drop—but beneficial in parallel ones, yielding an 8.6% gain. The parallel architecture further exhibits positive cross-task transferability, whereas the sequential variant suffers from catastrophic forgetting.

Technology Category

Machine Learning: Mixture of Experts (MoE)Cognitive Modeling & Cognitive Systems: Adaptive BehaviorNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practice applies adapters uniformly without considering the distinct functional roles of each component type. We systematically study component-type LoRA placement across two hybrid architectures -- Qwen3.5-0.8B (sequential, GatedDeltaNet + softmax attention) and Falcon-H1-0.5B (parallel, Mamba-2 SSM + attention) -- fine-tuned on three domains and evaluated on five benchmarks. We find that the attention pathway -- despite being the minority component -- consistently outperforms full-model adaptation with 5-10x fewer trainable parameters. Crucially, adapting the recurrent backbone is destructive in sequential hybrids (-14.8 pp on GSM8K) but constructive in parallel ones (+8.6 pp). We further document a transfer asymmetry: parallel hybrids exhibit positive cross-task transfer while sequential hybrids suffer catastrophic forgetting. These results establish that hybrid topology fundamentally determines adaptation response, and that component-aware LoRA placement is a necessary design dimension for hybrid architectures.
Problem

Research questions and friction points this paper is trying to address.

LoRA placement
hybrid language models
component-type adaptation
attention mechanism
recurrent components
Innovation

Methods, ideas, or system contributions that make the work stand out.

LoRA placement
hybrid language models
component-aware adaptation
Mamba-2
parameter-efficient fine-tuning
🔎 Similar Papers
No similar papers found.
H
Héctor Borobia
VRAIN – Valencian Research Institute for Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain
E
Elies Seguí-Mas
Department of Economics and Social Sciences, Universitat Politècnica de València, Valencia, Spain
G
Guillermina Tormo-Carbó
Department of Business Organisation, Universitat Politècnica de València, Valencia, Spain