Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inefficiency of small reasoning models constrained by parametric knowledge, where blindly increasing inference compute yields diminishing returns. We propose FlyBy, a framework that trains models to β€œreason first, diagnose later” by distinguishing execution from knowledge bottlenecks, enabling on-demand queries to stronger models when encountering knowledge deficits. The method integrates supervised fine-tuning, multi-depth query actions, and cost-aware reinforcement learning, alongside intermediate-state intervention analysis for selective querying. Experimental results demonstrate that FlyBy-4B surpasses Qwen3-14B in performance at lower serving costs, while the 8B variant achieves 51.81% Pass@8. This work establishes a new paradigm for efficient reasoning in small language models.
πŸ“ Abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
Problem

Research questions and friction points this paper is trying to address.

small reasoning models
knowledge bottleneck
test-time computation scaling
parametric knowledge
self-refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small Reasoning Models
Selective Querying
Cost-aware Reinforcement Learning
Knowledge Bottleneck
Test-time Computation
πŸ”Ž Similar Papers
No similar papers found.