Loop Dropout: Regularizing Shared Updates in Looped Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the late-stage bias problem in LoRA fine-tuning for recurrent language models, wherein shared parameter updates disproportionately concentrate at later recurrence positions, resulting in inadequate early adaptation. To mitigate this, we propose Loop Dropout, a method that innovatively couples stochastic dropout with an inverse survival scaling mechanism. By randomly masking adapter layers and compensating for expected magnitude, our approach ensures balanced update amplitudes across all recurrence positions without introducing additional inference overhead. Extensive experiments demonstrate that Loop Dropout significantly outperforms existing LoRA variants on mathematical reasoning, instruction following, and code generation tasks. This work effectively resolves the challenge of uneven parameter adaptation inherent in shared-weight recurrent deep computation.
📝 Abstract
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update is more effective at later loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation.
Problem

Research questions and friction points this paper is trying to address.

Looped Language Models
Low-Rank Adaptation (LoRA)
Late-Loop Bias
Shared Update Regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Loop Dropout
Looped Language Models
Low-Rank Adaptation (LoRA)
Inverse-Survival Rescaling
Late-Loop Bias
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.