Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of fixed-expert MoE-LoRA architectures, which inefficiently allocate computation by employing a constant number of experts regardless of input token difficulty—wasting resources on easy tokens while under-provisioning for challenging ones. To overcome this, the authors propose CARE, a novel dynamic routing method that leverages both the confidence of the router’s output distribution and inter-expert disagreement as signals to guide expert selection. Experts are activated via nucleus sampling until their cumulative weights reach an adaptive threshold, while a budget thermostat regulates the average number of active experts. Requiring no additional parameters and only a single forward pass, CARE achieves comparable or superior performance to top-k=4 MoE-LoRA on LLaMA-3.1-8B and Qwen2.5-7B using fewer experts, while also significantly enhancing out-of-distribution detection capability.
📝 Abstract
Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncertainty signal: peaked mass indicates confidence, while a flat distribution indicates ambiguity. We introduce CARE (Confidence-Adaptive Routing of Experts), which admits experts in a nucleus fashion. Experts are activated in decreasing router weight until their cumulative mass reaches a threshold, with a small extension when the admitted experts disagree. A budget thermostat calibrates the threshold so that the average number of active experts matches any target. CARE is a drop-in, single-forward-pass rule with no extra parameters. Across eight commonsense benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, as well as math, code, and knowledge tasks, CARE improves over fixed top-k MoE-LoRA at matched compute and matches the fixed-k=4 baseline while activating fewer experts. The same confidence and disagreement signals also improve out-of-distribution detection over MSP, entropy, and multi-pass proxies. We support the design with nucleus fidelity, budget optimality, and an epistemic reading of disagreement, and we release code.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Low-Rank Adaptation
Confidence-Adaptive Routing
Token Uncertainty
Expert Allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Confidence-Adaptive Routing
Mixture-of-Experts LoRA
Nucleus Sampling
Dynamic Expert Activation
Uncertainty-aware Routing
💼 Related Jobs
No related jobs found.
T
Tom Saliencro
University of California, Irvine
R
Rohan Desai
University of Washington
P
Priya Nair
University of California, Irvine
M
Maya Lindqvist
University of California, Irvine
D
Daniel Whitmore
University of Washington