No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of logical discontinuities and unnatural state transitions caused by prompt switching in streaming video generation by proposing the SEGUE framework. Specifically, this work introduces the first training-free explicit transition planner, which parses frames and prompts to generate a sequence of transition instructions, ensuring smooth evolution of video states. Additionally, it presents the SPANDMD method, which optimizes training stability for short-period scheduling while preserving full temporal context. By integrating diffusion model distillation with autoregressive generation techniques, the proposed approach supports transferability to frozen models. Experimental results demonstrate that this method ranks first across eight metrics on the OpenTrans-360 benchmark, achieving an overall score of 0.887, and attains leading performance on multiple metrics within StreamAV-Bench.
📝 Abstract
Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goals or taking heuristic shortcuts that bypass necessary intermediate state changes needed for a plausible transition. In this study, we present SEGUE, a novel framework that makes this process explicit and trains the generator to execute these transitions faithfully. At each switch, a training-free planner parses the latest frame and prompts, writes a few segue prompts with roles and durations, and then hands control back to the user's prompt. Furthermore, to address the inherent difficulty of training causal models on short-lived temporal schedules without corrupting preparatory supervision, we introduce SPANDMD, which evaluates each active prompt using the full rollout as temporal context while retaining its DMD residual only within the prompt's assigned span. On OpenTrans-360, a benchmark of 1,800 switches that scores how the old state exits and the new one begins, SEGUE ranks first on all eight transition metrics and raises the overall score over the strongest baseline from 0.866 to 0.887. It also ranks first on four of six instruction-response metrics of StreamAV-Bench, while the planner transfers to frozen autoregressive generators without retraining.
Problem

Research questions and friction points this paper is trying to address.

streaming video generation
mid-stream prompt switching
state-grounded transitions
causal model training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming video generation
Mid-stream prompt switching
Training-free planner
SPANDMD
State-grounded transitions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zejing Rao
University of Chinese Academy of Sciences
K
Ketong Ren
University of Chinese Academy of Sciences
X
Xiaoqiang Liu
Kling AI
Yiping Meng
Yiping Meng
Kuaishou
Computer Vision
Guoxin Zhang
Guoxin Zhang
School of Computer Science, Beijing University of Posts and Telecommunications
Computer VisionPattern Recognition
F
Fan Tang
University of Science and Technology Beijing