Beyond Masks and Trajectories: Flow-Guided Latent Action Injection for Stable Surgical Video Generation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of existing surgical video generation methods on costly auxiliary conditions and their limited scalability by proposing the FLAIR framework. This work pioneers an auxiliary-condition-free architecture that learns motion priors via optical flow and injects them into the latent space of a frozen foundation model, enabling video generation through latent action prediction. Furthermore, the authors construct the SurgActionClip dataset and introduce SurgMetrics, a dedicated evaluation metric for this domain. Experimental results demonstrate that FLAIR generates highly consistent, high-quality surgical videos using text prompts alone. It achieves significantly superior human perceptual alignment compared to conventional approaches, effectively overcoming the conditional dependency bottleneck in surgical video generation.
📝 Abstract
Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with realistic and clinically plausible motion remains challenging. Most existing methods rely on auxiliary conditions, such as masks, trajectories, depth, or reference videos, to achieve visually plausible synthesis. Yet, these auxiliary conditions typically require additional manual annotation or specialized acquisition, making it difficult to scale such methods beyond small, curated datasets. This motivates the need for a reference-free architecture capable of generating high-quality surgical video without requiring auxiliary visual conditions at inference time. We propose FLAIR, a Flow-guided LatentAction Injection framework for Reference-free surgical video generation. FLAIR learns action priors from optical flow of real surgical videos, dynamically predicts corresponding latent action representation from an input prompt, and injects it into a frozen base model to generate surgical videos with improved action consistency. We further construct SurgActionClip-30K, the first large-scale surgical vision dataset comprising action-centric segmented clips and structured caption labels, addressing the persistent lack of fine-grained, action-centric surgical datasets. Lastly, we introduce SurgMetrics, the first surgical domain-specific evaluation metrics for quantifying the quality of generated surgical videos, addressing the persistent absence of clinically grounded evaluation standards in this domain. Extensive experiments demonstrate that FLAIR enables generating high-quality surgical videos using text-only inference without auxiliary conditions, and validation in SurgMetrics demonstrates its strength in alignment with human perception compared to traditional metrics.
Problem

Research questions and friction points this paper is trying to address.

Surgical video generation
Reference-free synthesis
Action consistency
Surgical dataset
Evaluation metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-guided Latent Action Injection
Reference-free Video Generation
Surgical Vision Dataset
Domain-specific Evaluation Metrics
Optical Flow
🔎 Similar Papers
No similar papers found.