🤖 AI Summary
This study investigates the unclear reasoning mechanisms and efficiency bottlenecks of continuous-state flow language models (FLMs). Adopting a superposition perspective, it combines theoretical analysis with intermediate-state intervention experiments to compare FLMs against discrete diffusion models on maze planning and Sudoku tasks. The findings reveal that continuous representations intrinsically retain multiple candidate hypotheses, thereby optimizing few-step reasoning. Empirically, continuous states persistently encode information, substantially enhancing few-step inference efficiency. FLMs achieve superior sequence-level accuracy under limited reasoning steps; notably, on the Maze15 task, they attain 95% accuracy while requiring only 63.5% of the original parameter count, underscoring the remarkable parameter efficiency of continuous flow models.
📝 Abstract
Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution accuracy under matched denoising steps. Unlike discrete diffusion, which passes categorical states between denoising steps, FLMs evolve a continuous sequence representation throughout denoising and decodes it into discrete tokens only at the end. Our theoretical analysis shows, from a superposition perspective, how information retained in these continuous states can benefit reasoning. Intermediate-state interventions provide further empirical support for this theoretical account, showing that removing information about alternative candidates reduces subsequent solution recovery. Together, these findings show that FLMs allow evidence for multiple candidates to persist and inform subsequent reasoning before a discrete answer is produced. Furthermore, our experiments on maze planning and Sudoku tasks show that FLMs achieve greater reasoning efficiency in the few-step regime: FLMs achieves higher sequence accuracy than discrete diffusion baselines at matched model sizes and small denoising steps. On maze planning tasks, FLMs can also achieve comparable accuracy with smaller models. For example, on Maze15, FLM reaches the 95\% accuracy target at 64 denoising steps with 36.5\% fewer parameters than MDLM. These findings point to continuous state spaces as a promising foundation for reasoning models that require fewer refinement steps.