🤖 AI Summary
This study addresses the computational inefficiency of long chain-of-thought reasoning and the challenge of automating routing decisions in reasoning models by proposing a joint optimization framework based on online reinforcement learning. Methodologically, it employs the GRPO algorithm with a paired-to-self-routing curriculum learning strategy, enabling end-to-end training without supervised fine-tuning warm-up. The core innovation lies in introducing the first hierarchical counterfactual credit assignment mechanism, which effectively decouples the credit signals for routing decisions from those for content generation. Experiments demonstrate that this approach improves accuracy on the Qwen3 model series while reducing token consumption by over 40%. Furthermore, it exhibits strong cross-domain generalization capabilities, establishing a new paradigm for efficient, adaptive inference.
📝 Abstract
Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Many dual-mode models leave this choice to users. Automating it is challenging because routing targets evolve with the policy, initial mode preferences destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and modeconditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts assign cross-mode credit only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced rollouts from both modes, then increases self-routed updates to improve autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods. Relative to always-thinking checkpoints, it improves macro-average accuracy while reducing mean generated tokens by 51% for Qwen3-8B and 41% for Qwen3-14B. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and instruction following, its routing behavior and response quality generalize to held-out coding, science, knowledge, and commonsense benchmarks.