🤖 AI Summary
Large reasoning language models excel at complex reasoning tasks but suffer from “overthinking,” leading to substantial computational overhead. To address this, we propose a training-free adaptive reasoning suppression method. Our approach introduces a multi-checkpoint confidence evaluation mechanism coupled with a progressive dynamic suppression threshold, enabling real-time monitoring and step-wise pruning of reasoning paths. Unlike static suppression strategies—which force a trade-off between accuracy and efficiency—our method operates entirely through forward inference without requiring retraining or fine-tuning. We evaluate it across multiple mathematical reasoning benchmarks (GSM8K, MATH) and diverse model architectures (Llama-3, Qwen2). Results show an average reduction of 53% in token consumption, 46.1% in latency, and 57.9% in energy consumption, while maintaining or even improving accuracy. This demonstrates that adaptive, confidence-guided suppression can significantly enhance inference efficiency without compromising reasoning fidelity.
📝 Abstract
Large Reasoning Language Models (LRLMs or LRMs) demonstrate remarkable capabilities in complex reasoning tasks, but suffer from significant computational inefficiencies due to overthinking phenomena. Existing efficient reasoning methods face the challenge of balancing reasoning quality with inference cost reduction. We propose extbf{Adaptive Reasoning Suppression (ARS)}, a novel training-free approach that dynamically suppresses redundant reasoning steps while preserving accuracy through adaptive certainty monitoring. ARS introduces a multi-checkpoint certainty estimation mechanism with progressive suppression thresholds, achieving superior efficiency compared to static suppression methods. Our extensive evaluation across mathematical reasoning benchmarks using multiple model architectures demonstrates that ARS achieves up to 53%, 46.1%, and 57.9% in token, latency and energy reduction, while maintaining or improving accuracy.