🤖 AI Summary
This study addresses the limited inference precision and low computational efficiency of continuous diffusion models by proposing the CEDR framework. This method constructs compact latent representations via multi-layer teacher knowledge distillation, enabling asynchronous denoising and efficient inference. It incorporates hierarchical compression, staged curriculum learning, and an adaptive NFT guidance mechanism, while integrating continuous embedding diffusion, the ELF training paradigm, and multi-scale denoising strategies to effectively replace conventional autoregressive Transformer architectures. Experimental results demonstrate that CEDR significantly outperforms baseline models of comparable scale on GSM8K (63.74%), MATH500 (24.6%), and HumanEval (32.85%).
📝 Abstract
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM