ARS: Adaptive Reasoning Suppression for Efficient Large Reasoning Language Models

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large reasoning language models excel at complex reasoning tasks but suffer from “overthinking,” leading to substantial computational overhead. To address this, we propose a training-free adaptive reasoning suppression method. Our approach introduces a multi-checkpoint confidence evaluation mechanism coupled with a progressive dynamic suppression threshold, enabling real-time monitoring and step-wise pruning of reasoning paths. Unlike static suppression strategies—which force a trade-off between accuracy and efficiency—our method operates entirely through forward inference without requiring retraining or fine-tuning. We evaluate it across multiple mathematical reasoning benchmarks (GSM8K, MATH) and diverse model architectures (Llama-3, Qwen2). Results show an average reduction of 53% in token consumption, 46.1% in latency, and 57.9% in energy consumption, while maintaining or even improving accuracy. This demonstrates that adaptive, confidence-guided suppression can significantly enhance inference efficiency without compromising reasoning fidelity.

Technology Category

Knowledge Representation and Reasoning: Computational Complexity of ReasoningCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningMachine Learning: Large Multimodal Models (LMMs)

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Large Reasoning Language Models (LRLMs or LRMs) demonstrate remarkable capabilities in complex reasoning tasks, but suffer from significant computational inefficiencies due to overthinking phenomena. Existing efficient reasoning methods face the challenge of balancing reasoning quality with inference cost reduction. We propose extbf{Adaptive Reasoning Suppression (ARS)}, a novel training-free approach that dynamically suppresses redundant reasoning steps while preserving accuracy through adaptive certainty monitoring. ARS introduces a multi-checkpoint certainty estimation mechanism with progressive suppression thresholds, achieving superior efficiency compared to static suppression methods. Our extensive evaluation across mathematical reasoning benchmarks using multiple model architectures demonstrates that ARS achieves up to 53%, 46.1%, and 57.9% in token, latency and energy reduction, while maintaining or improving accuracy.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational inefficiency in large reasoning language models
Balancing reasoning quality with inference cost reduction
Dynamically suppressing redundant reasoning steps while preserving accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic suppression of redundant reasoning steps
Training-free adaptive certainty monitoring mechanism
Multi-checkpoint certainty estimation with progressive thresholds