RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation and exacerbated overthinking in quantized reasoning models by proposing a training-free, token-level adaptive penalty mechanism. The method leverages quantization-aware behavioral analysis to precisely identify redundant thinking tokens specific to each model, and applies customized penalties guided by the full-precision counterpart, thereby overcoming the limitations of traditional predefined markers. Experimental results demonstrate that this mechanism improves accuracy by 9.8 percentage points while reducing chain-of-thought length by 51.3%. By preserving reasoning fidelity and substantially enhancing computational efficiency, the proposed approach achieves a superior accuracy-efficiency trade-off for quantized large language models.
📝 Abstract
Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.
Problem

Research questions and friction points this paper is trying to address.

post-training quantization
reasoning models
overthinking
inference efficiency
chain-of-thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training quantization
Reasoning models
Overthinking tokens
Token-level penalty
Chain-of-thought
💼 Related Jobs
No related jobs found.