Score
Designs and implements systems that locate, track, and modify textual content embedded in video frames so the text is replaced or altered across a sequence while matching original visual style (font, color, stroke, shading) and remaining stroke‑level legible. Builds methods to preserve temporal coherence across frames (handling motion, occlusion, and lighting changes) and to render consistent, artifact‑free edits throughout the video.
This study addresses the challenge in video text editing where diffusion models struggle to reproduce precise stroke structures, frequently resulting in garbled characters. To overcome this limitation, we propose a trajectory-aligned glyph rendering approach coupled with deep normalized feature supervision. By introducing frame-wise glyph guidance and multi-depth feature supervision via a frozen recognizer, the generation process is effectively constrained. Furthermore, we construct VTEdit, the first standardized evaluation benchmark incorporating real-world scene trajectories. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in both text accuracy and background consistency, achieving a sentence-level accuracy of 0.9408 while obtaining the highest user preference scores.
Video text editing faces three major challenges: cross-frame accuracy, temporal coherence, and style consistency, with existing methods struggling to achieve stroke-level precision for small-region text modifications. This work proposes the first unified video text editing framework built upon a frozen video diffusion Transformer. It introduces a lightweight text-context adapter, integrates a style encoder with a dual-granularity (line- and character-level) glyph encoder, and designs a glyph-aware spatially focused loss function. Leveraging a three-stage progressive training strategy and a large-scale synthetic dataset, SteerVTE-1M, the method significantly outperforms current baselines in text accuracy, style consistency, and temporal coherence.
This study addresses the persistent challenge in video scene text editing of simultaneously preserving visual fidelity, temporal consistency, and edit locality. To this end, it proposes a systematic solution comprising both evaluation infrastructure and a novel generative model. Methodologically, we construct ViTeX-Bench, a benchmark featuring a real-world paired dataset and an inaugural three-dimensional evaluation protocol. Furthermore, we introduce ViTeX-Edit-14B, an open-source model that integrates OCR calibration, Pareto comparison, and motion-aligned glyph-video conditioned fine-tuning. Experimental results demonstrate that the proposed model achieves a state-of-the-art character accuracy of 0.688 within the field while significantly suppressing text distortion. Ultimately, this work establishes a reproducible research foundation for advancing video text editing.
Existing approaches struggle to jointly handle scene text editing tasks—deletion, generation, and replacement—within a unified framework that simultaneously ensures precise textual appearance control and background integrity. To address this, this work proposes a unified model that decomposes complex text editing into two atomic operations: rendering and erasure. It introduces Overlay-Reference Positional Encoding (ORPE) to achieve pixel-level layout fidelity and exemplar-driven style control, complemented by a Region-Adaptive Suppression (RAS) strategy to ensure clean text removal. The study also establishes TextWand-Bench, the first comprehensive benchmark for general scene text editing. Experimental results demonstrate that the proposed method significantly outperforms both open-source and closed-source models across all three editing tasks in terms of text accuracy, layout-style consistency, and overall image quality.
Existing text-to-video (T2V) models struggle to generate long-duration, motion-rich videos with strong temporal coherence, particularly in modeling implicit temporal logic within prompts and enabling frame-level fine-grained text guidance. To address this, we propose the Cross-Frame Text-Guided Module (CTGM), the first framework integrating a Temporal Information Injector (TII), a Temporal Affinity Refiner (TAR), and a Temporal Feature Booster (TFB) to achieve frame-specific text conditioning and dynamic temporal alignment. Built upon a diffusion-based architecture, our method incorporates latent-space temporal injection, cross-frame text-feature correlation recalibration, and implicit temporal consistency enhancement. Extensive experiments across multiple benchmarks demonstrate significant improvements in motion coherence and semantic fidelity. Both qualitative and quantitative evaluations consistently surpass state-of-the-art methods. The code, pre-trained models, and demonstration videos are publicly released.
Existing text-guided video editing methods often struggle to simultaneously achieve temporal consistency and editability, typically requiring a trade-off between the two. This work proposes EquiEdit, a novel framework that jointly optimizes both aspects to enable high-quality video editing. The approach introduces a temporal Mamba module to enhance inter-frame consistency and incorporates a spectral-transform-based noise injection strategy that preserves the structure of the initial latent-space noise while improving editing flexibility. Built upon a diffusion model architecture, EquiEdit further integrates a temporal-aware scanning mechanism to better capture dynamic content across frames. Experimental results demonstrate that EquiEdit significantly outperforms existing methods in terms of temporal consistency, editability, and fidelity to the input video.
This study addresses the illegibility of visual text in video generation caused by stroke corruption and temporal instability, as well as the neglect of dynamic properties in existing benchmarks. To this end, it proposes VidScribe, a unified diagnostic benchmark spanning four generation paradigms. By constructing a conditionally orthogonal factor space and a trajectory-based gated evaluation suite, this work systematically assesses the text rendering and editing capabilities of eleven systems. The findings reveal that video text capability is not monolithic and identify localized editing as the primary bottleneck. Furthermore, by integrating preference optimization techniques, the study demonstrates that degradation concentrates within structural and temporal factors, and that benchmark-aligned preference optimization significantly enhances visual text generation quality.
This work addresses key challenges in long-form video editing, including maintaining narrative coherence across multiple stages and enabling precise error localization and localized correction. The authors propose an open-source, multimodal multi-agent system that structures the editing process into three phases: asset preparation, editing research, and timeline execution. A novel traceable and replayable editing trajectory mechanism is introduced, allowing for accurate diagnosis of failed segments and selective re-editing without requiring full pipeline re-execution. The system integrates multi-agent collaboration, multimodal analysis, tool invocation logging, intermediate rendering, and a verifiable reward design. Experimental evaluation across 23 themes demonstrates that the approach achieves an average human rating of 3.40 out of 5, significantly outperforming CapCut-Mate and CutClaw in thematic relevance, narrative coherence, and editing fluency.
This study addresses the cumbersome processes of localization, segmentation, and prompt construction in multi-shot generative video editing by proposing an interaction paradigm grounded in multi-level structural parsing. The method transforms videos into malleable hierarchical structures, enabling users to modify elements within a task-centric workspace while AI agents automatically handle intent translation and change propagation. Based on this approach, an interactive system is developed to support both rapid prototyping and end-to-end post-production workflows. User studies and expert evaluations demonstrate that the system significantly enhances efficiency in video comprehension, intent expression, and solution exploration, thereby establishing an effective new paradigm for generative video editing.
This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.