Gumbel Straight Flow: Distilling Autoregressive Models into One-step Flow Maps

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow generation speed of autoregressive language models by proposing an efficient one-step text generation method. By leveraging noise-data coupling from pretrained models, the approach constructs a continuous flow mapping language model that distills multi-step generation into a single step. Theoretically, it establishes the existence of non-crossing linear paths between Gumbel noise and target sequences, motivating a teacher-guided flow mapping semigroup optimization objective. Integrating a Gumbel-Straight-Flow architecture with knowledge distillation techniques, the proposed method significantly outperforms existing few-step text generation baselines in generation quality across multiple benchmarks.
📝 Abstract
We present Gumbel Straight Flow (GSF), a continuous flow map language model that leverages the noise-data coupling of a pretrained autoregressive language (AR) model. We theoretically demonstrate that the coupling between Gumbel noise and one-hot token sequences induced by an autoregressive model yields non-intersecting linear paths connecting the noise to the sequence representations. To further enhance high-quality few-step path sampling, we use a flow map semigroup objective where the tangent (velocity) condition is guided directly by the AR teacher. Across various benchmarks, including pretraining and downstream tasks, GSF can outperform current few-step language generation baselines.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive Models
Knowledge Distillation
Flow Maps
Few-step Generation
Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gumbel Straight Flow
Flow Matching
Autoregressive Distillation
One-step Generation
Semigroup Objective