🤖 AI Summary
This study addresses the challenge that existing latent reasoning methods struggle to simultaneously achieve usefulness, diversity, interpretability, refinability, and efficiency. To this end, this work proposes FLaRe, a framework that optimizes latent space training strategies based on the flow matching paradigm. Furthermore, it introduces a final-stage training recipe encompassing encoding mechanisms, training positions, and self-verifying chain-of-thought prompting. Experimental results demonstrate that FLaRe comprehensively outperforms existing approaches on arithmetic reasoning benchmarks. Notably, it attains 97% of the accuracy achieved by explicit chain-of-thought methods while requiring only one-quarter of the inference latency. Consequently, FLaRe successfully realizes the synergistic optimization of all five criteria, offering an effective solution for advancing latent reasoning capabilities in large language models.
📝 Abstract
Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.