MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy and inflexibility of large reasoning models on simple problems by proposing a lightweight metacognitive controller. Formulating reasoning regulation as a sequential decision-making problem, the method trains the controller via reinforcement learning to directly observe reasoning trajectories and intervene dynamically, with rewards balancing correctness and conciseness. Operating in a plug-and-play manner on frozen models, the controller requires no retraining, supervised data, or predefined budgets, while supporting zero-shot cross-model transfer. Experiments demonstrate that this approach improves accuracy by approximately 4.7% and reduces generation length by 53.3% across multiple benchmarks, significantly enhancing both reasoning efficiency and generalization capability.
📝 Abstract
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
Problem

Research questions and friction points this paper is trying to address.

Large Reasoning Models
Redundant Reasoning
Metacognitive Control
Reasoning Regulation
Inference Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Metacognitive Control
Reinforcement Learning
Reasoning Regulation
Plug-and-Play Controller
Large Reasoning Models
🔎 Similar Papers