🤖 AI Summary
This study addresses the mischaracterization in existing research regarding large language models’ persistent explicit reasoning inertia under “no-thinking” instructions. To systematically quantify behavioral compliance, this work proposes a three-tier evaluation framework incorporating an “empty thinking rate,” which leverages structured response decomposition and LLM-as-a-judge automated assessment to precisely distinguish valid answers from filler text. The findings reveal how task difficulty modulates reasoning cessation, demonstrating that explicit control cannot entirely suppress the reasoning process. Furthermore, the analysis identifies a trade-off between instruction compliance and accuracy in open-ended tasks, establishing “no-thinking” as an independent capability dimension. These insights clarify the boundaries of instruction-following behavior and offer a rigorous methodology for evaluating cognitive control in large language models.
📝 Abstract
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.