LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference

πŸ“… 2026-01-05
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the inefficiency of token-by-token decoding in autoregressive large language models, where existing acceleration methods often sacrifice accuracy due to uncompensated layer skipping or rely on additional routing mechanisms. The authors propose LoRA-Drop, a routing-free approach that reuses the hidden state from the previous token for most decoding steps through temporally scheduled computation, while applying low-rank LoRA corrections to maintain fidelity. To mitigate representation drift, the method periodically performs full forward passes. Compatible with standard KV caching, LoRA-Drop integrates dynamic layer skipping with periodic refreshes, achieving up to 2.6Γ— speedup and 45–55% KV cache reduction on LLaMA2/3 and Qwen2.5 models, with accuracy degradation limited to less than 0.5 percentage points.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
πŸ“ Abstract
Autoregressive large language models (LLMs) are bottlenecked by sequential decoding, where each new token typically requires executing all transformer layers. Existing dynamic-depth and layer-skipping methods reduce this cost, but often rely on auxiliary routing mechanisms or incur accuracy degradation when bypassed layers are left uncompensated. We present \textbf{LoRA-Drop}, a plug-and-play inference framework that accelerates decoding by applying a \emph{temporal compute schedule} to a fixed subset of intermediate layers: on most decoding steps, selected layers reuse the previous-token hidden state and apply a low-rank LoRA correction, while periodic \emph{refresh} steps execute the full model to prevent drift. LoRA-Drop requires no routing network, is compatible with standard KV caching, and can reduce KV-cache footprint by skipping KV updates in droppable layers during LoRA steps and refreshing periodically. Across \textbf{LLaMA2-7B}, \textbf{LLaMA3-8B}, \textbf{Qwen2.5-7B}, and \textbf{Qwen2.5-14B}, LoRA-Drop achieves up to \textbf{2.6$\times$ faster decoding} and \textbf{45--55\% KV-cache reduction} while staying within \textbf{0.5 percentage points (pp)} of baseline accuracy. Evaluations on reasoning (GSM8K, MATH, BBH), code generation (HumanEval, MBPP), and long-context/multilingual benchmarks (LongBench, XNLI, XCOPA) identify a consistent \emph{safe zone} of scheduling configurations that preserves quality while delivering substantial efficiency gains, providing a simple path toward adaptive-capacity inference in LLMs. Codes are available at https://github.com/hosseinbv/LoRA-Drop.git.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
sequential decoding
layer skipping
KV-cache
efficiency bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

LoRA-Drop
temporal decoding
efficient LLM inference
KV-cache reduction
low-rank adaptation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
H
Hossein Rajabzadeh
Department of Mechanical and Mechatronics Engineering, University of Waterloo
M
Maryam Dialameh
Department of Mechanical and Mechatronics Engineering, University of Waterloo
Chul B. Park
Chul B. Park
Distinguished Professor of Microcellular Engineered Plastics, University of Toronto
Polymer ProcessingMicrocellular Foaming
I
Il-Min Kim
Department of Electrical and Computer Engineering, Queen’s University
H
Hyock Ju Kwon
Department of Mechanical and Mechatronics Engineering, University of Waterloo