🤖 AI Summary
This study addresses the limitation of existing unlearning methods for video generation models, which tend to degrade background scenes while neglecting motion concepts. We propose FOMO, the first training-based selective video unlearning framework that prioritizes scene preservation. By leveraging localized representation modification and complementary objective function optimization, FOMO precisely removes target content while employing an auxiliary-data-free mechanism to maintain consistency in non-target information. Notably, this work extends video unlearning to motion behavior elimination for the first time, achieving effective forgetting across harmful content, objects, and motion concepts. Overall, FOMO attains an optimal trade-off between thorough concept removal and robust scene preservation.
📝 Abstract
The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesirable material, making concept unlearning an essential mitigation. Existing methods mainly target static visual concepts, such as objects, identities, or unsafe appearance, largely overlooking motion unlearning. Furthermore, these approaches often pay little attention to preserving the surrounding scene. As a result, successful concept removal may unintentionally alter the background, composition, or overall video dynamics. We argue that effective unlearning should ideally change only what is targeted, while minimizing unnecessary changes to the remaining scene. In this work, we introduce FOMO, to the best of our knowledge the first training-based selective video unlearning method that directly treats preservation of the original scene as a priority. We formulate unlearning around two complementary objectives: what to change and what to preserve. Our method localizes concept-related representations and modifies them, while the preservation mechanism maintains non-target scene information without requiring auxiliary data. Beyond simply erasing unwanted concepts, FOMO explicitly redirects the generation toward a specified safe alternative. We further extend this formulation to motion unlearning, where the concept is defined by temporal behavior rather than a fixed spatial region. Our solution achieves effective unlearning across unsafe content, object, and motion concepts, while achieving the best trade-off between concept removal and scene preservation.
Code: https://github.com/gmum/FOMO
Project Page https://gmum.github.io/FOMO