🤖 AI Summary
This study addresses the limitation of reasoning models in solving complex problems, particularly their difficulty in effectively extracting reusable solution strategies from known solutions. To overcome this, it proposes a hindsight hierarchical framework coupled with a self-improving loop mechanism that jointly trains three tasks: answer prediction, reverse engineering, and thought-based solving. By reverse-engineering known solutions to generate additional supervisory signals, the approach leverages hindsight to enable closed-loop self-enhancement, which is subsequently applied to the Lean interactive theorem prover. The primary contribution lies in providing formal specifications and concrete instantiations of the core methodology, establishing a novel paradigm for the autonomous iterative optimization of reasoning models, while empirical evaluation remains to be conducted in future work.
📝 Abstract
We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.