VIS-Ground: Video Interactive Storytelling with Contextual Grounding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning implicit dependencies among viewer instructions, source content, and historical states in interactive video generation by introducing the novel concept of context grounding. Methodologically, structured context abstraction transforms multi-source heterogeneous inputs into executable constraint models. Combined with a constraint induction mechanism and a plan-verify-correct-render closed-loop framework, this approach systematically achieves joint dependency modeling to guide constrained video synthesis. Extensive experiments demonstrate that the proposed method attains the highest overall scores across three mainstream video generation backbones, yielding an average absolute improvement of 10.3 points over the strongest baseline.
📝 Abstract
Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer's request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.
Problem

Research questions and friction points this paper is trying to address.

interactive video storytelling
contextual grounding
video generation
cross-context dependencies
heterogeneous inputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contextual Grounding
Interactive Video Generation
Structured Context Abstraction
Constraint Induction
Video Interactive Storytelling
🔎 Similar Papers
No similar papers found.