🤖 AI Summary
This study addresses the slow generation speed of autoregressive language models by proposing an efficient one-step text generation method. By leveraging noise-data coupling from pretrained models, the approach constructs a continuous flow mapping language model that distills multi-step generation into a single step. Theoretically, it establishes the existence of non-crossing linear paths between Gumbel noise and target sequences, motivating a teacher-guided flow mapping semigroup optimization objective. Integrating a Gumbel-Straight-Flow architecture with knowledge distillation techniques, the proposed method significantly outperforms existing few-step text generation baselines in generation quality across multiple benchmarks.
📝 Abstract
We present Gumbel Straight Flow (GSF), a continuous flow map language model that leverages the noise-data coupling of a pretrained autoregressive language (AR) model. We theoretically demonstrate that the coupling between Gumbel noise and one-hot token sequences induced by an autoregressive model yields non-intersecting linear paths connecting the noise to the sequence representations. To further enhance high-quality few-step path sampling, we use a flow map semigroup objective where the tangent (velocity) condition is guided directly by the AR teacher. Across various benchmarks, including pretraining and downstream tasks, GSF can outperform current few-step language generation baselines.