Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty large language models face in leveraging intermediate states for iterative refinement during structured reasoning. To this end, we propose FOCUS, a method that introduces a recurrent update mechanism to iteratively optimize solution states. Furthermore, it selects samples situated at the learning frontier from self-generated trajectories for training, thereby maximizing reasoning improvements. Experimental results demonstrate that FOCUS significantly enhances accuracy on Sudoku and maze tasks. Notably, it achieves zero-shot transfer to mathematical reasoning and code execution tasks without fine-tuning, effectively strengthening the model's generalizable reasoning capabilities.
📝 Abstract
A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.
Problem

Research questions and friction points this paper is trying to address.

Structured Reasoning
Intermediate States
Large Language Models
Curriculum Learning
State Revision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Reasoning
Recurrent Updater
Frontier-Oriented Curation
Intermediate States
Zero-shot Transfer