EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of removing video objects entangled with complex real-world effects—such as compositing artifacts, weak object-effect associations, long-tailed physical phenomena, and dynamic interactions—that existing methods struggle to handle. To this end, we propose EffectLearner, a novel framework that introduces semantic reasoning to decouple objects from their associated effects. Our approach leverages a vision-language model for object-effect reasoning, integrates a DiT-based video inpainter with structured prompts to extract effect-aware context, and enhances spatiotemporal coherence through motion-aware mask guidance and motion consistency supervision. We contribute the first paired video dataset focused on complex effects, EffectWorld, and a progressive curriculum training strategy. Experiments demonstrate that our method significantly outperforms state-of-the-art approaches on ROSE-Bench, EffectWorld-Eval, and EffectWorld-Wild, achieving high-quality and temporally consistent removal of both objects and their induced effects.
📝 Abstract
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
Problem

Research questions and friction points this paper is trying to address.

video object removal
object-effect reasoning
real-world scenes
spatiotemporal coherence
compositional effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

object-effect reasoning
video object removal
visual-language model
diffusion transformer
spatiotemporal coherence
🔎 Similar Papers