🤖 AI Summary
This study addresses the limitation of existing flow-based models in reasoning tasks, where they fail to benefit from additional integration steps, thereby constraining deep reasoning capabilities. To overcome this, we propose ProsQA, a framework built upon flow matching and Transformer architectures. It introduces a stochastic sub-interval unfolding training mechanism to unlock computational potential and incorporates spherical projection techniques to stabilize latent variable dynamics, establishing a theoretical foundation for flow-based reasoning. Furthermore, a parameter-free selection scoring strategy is developed. Experimental results demonstrate that ProsQA achieves 97% accuracy on the ProsQA benchmark and significantly outperforms larger-parameter baseline models on long-horizon reasoning tasks such as Sudoku.
📝 Abstract
Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model's own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.