🤖 AI Summary
This study addresses the limited reasoning depth of single-forward-pass decision models by proposing a recurrent decision framework under the System 1.5 thinking paradigm. Built upon pretrained language models and recurrent neural network architectures, the method recursively applies network layers to iteratively refine hidden states before committing to an answer, enabling implicit intermediate reasoning without text generation. Furthermore, a layer-wise scoring mechanism combined with reinforcement learning optimization allows a single model to adaptively accommodate varying computational budgets. Experimental results demonstrate that the proposed model achieves 72% accuracy, surpassing an isomorphic non-recurrent baseline by 13.5 percentage points and effectively overcoming training-depth bottlenecks. Additionally, when employed as a judge model, it improves the generator's F1 score by 7.7 percentage points.
📝 Abstract
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.