When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models

📅 2026-01-22
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy caused by the uniform application of attention mechanisms in hybrid models by proposing the AMOR architecture. Employing Mamba2 or Gated DeltaNet as the recurrent backbone, this method introduces a parameter-free dynamic entropy gating mechanism that leverages running batch statistics to construct a binary gate, adaptively scheduling attention modules based on predictive uncertainty. Experimental results demonstrate that activating attention at only approximately 40% of positions achieves state-of-the-art performance on commonsense reasoning tasks while significantly enhancing long-context robustness. Ultimately, AMOR realizes an effective balance between computational efficiency and model accuracy.
📝 Abstract
Recurrent-attention hybrids aim to combine the efficiency of recurrence with the expressivity of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state alone is sufficient for accurate prediction. We introduce AMOR (Adaptive Metacognitive Output Router), a post-hoc hybrid architecture that selectively invokes attention based on predictive uncertainty. A recurrent backbone is augmented with entropy-gated attention blocks that activate only when the model's output entropy exceeds a dynamic threshold derived from a running batch median and scaled standard deviation. This yields a simple, gradient-free routing mechanism inspired by uncertainty-driven computation and the System 1 / System 2 distinction. Across Mamba2 and Gated DeltaNet backbones (180M-1.5B), AMOR consistently matches or outperforms both pure recurrent models and fixed-schedule hybrid baselines while invoking attention on only ~22% of tokens. It achieves strong performance on common-sense reasoning benchmarks and maintains stable long-context performance on LongBench, where prior hybrid models degrade under distribution shift. These results suggest that when attention is applied matters as much as how much: selectively allocating attention based on predictive uncertainty improves both efficiency and robustness, offering a simple alternative to uniform or fixed routing strategies and pointing toward adaptive hybrid architectures that dynamically match computation to input difficulty.
Problem

Research questions and friction points this paper is trying to address.

hybrid models
recurrent-attention
predictive uncertainty
computational efficiency
selective attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Entropy Gate
Hybrid Architecture
Predictive Uncertainty
Parameter-free Routing
Recurrent-Attention
🔎 Similar Papers
No similar papers found.