TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种结合文本和视觉指导的音频移除方法TV-AudioRemover,通过多任务训练和硬混合课程来解决视频中移除对象后其声音残留的问题。
📝 Abstract
Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.
Problem

Research questions and friction points this paper is trying to address.

audio-visual inconsistency
sound removal
multi-modal guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-Visual Guided Sound Removal
Multi-Task Hard-Mixture Curriculum
Natural-Language Instruction
Task Tokens
Modality-Specific Global Guidance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xinyue Guo
MiLM Plus, Xiaomi Inc.
J
Jianxuan Yang
MiLM Plus, Xiaomi Inc.
D
Daiguo Zhou
MiLM Plus, Xiaomi Inc.
J
Jiagao Hu
MiLM Plus, Xiaomi Inc.
Y
Yuxuan Chen
MiLM Plus, Xiaomi Inc.
F
Fei Wang
MiLM Plus, Xiaomi Inc.
Jian Luan
Jian Luan
Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis