On the Efficiency-Safety Dilemma in Large Reasoning Models

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了大型推理模型中效率与安全性之间的关系,通过量化和剪枝方法分析其对抗鲁棒性的影响,并提出最佳策略平衡效率与鲁棒性。
📝 Abstract
Large reasoning models (LRMs) incur high inference costs, often mitigated by efficiency techniques like quantization and pruning. However, the impact of these techniques on model adversarial robustness remains largely unexplored. This study provides the first comprehensive analysis of the interplay between efficiency, jailbreak vulnerability, and reasoning in LRMs. We find that while efficiency methods seemingly reduce the success rate of jailbreak attacks, this improvement is often superficial. It largely arises from degraded reasoning capabilities leading to "attempted but failed" malicious responses, rather than an increase in genuine alignment. Mechanistic analysis of representational drift confirms this, revealing a strict coupling between reasoning capability loss and the model's inability to maintain malicious semantic trajectories. Additionally, we identify quantization with pruning as the optimal strategy to balance efficiency and robustness. These findings clarify the distinction between true safety alignment and capability-induced failure, providing an empirical foundation for LRM deployment.
Problem

Research questions and friction points this paper is trying to address.

Large Reasoning Models
Adversarial Robustness
Efficiency Techniques
Jailbreak Vulnerability
Reasoning Capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

efficiency
jailbreak vulnerability
reasoning
quantization and pruning
representational drift