Treating Motion as Option with Output Selection for Unsupervised Video Object Segmentation

πŸ“… 2023-09-26
πŸ›οΈ IEEE transactions on circuits and systems for video technology (Print)
πŸ“ˆ Citations: 2
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
In unsupervised video object segmentation, models often suffer from unstable predictions due to over-reliance on motion cues such as optical flow. To address this, we propose the β€œMotion-as-Option” mechanism, which decouples motion information into an optional module: during training, optical flow inputs to the motion encoder are randomly replaced with RGB frames, and an adaptive output selection algorithm dynamically fuses predictions from parallel motion and appearance pathways. This is the first approach to enable non-mandatory modeling of motion cues, integrating motion-appearance collaborative representation learning with stochastic input masking. Our method achieves state-of-the-art performance on DAVIS and FBMS benchmarks, demonstrating significantly improved robustness against anomalous motion disturbances and yielding a 23% gain in prediction stability.
πŸ“ Abstract
Unsupervised video object segmentation aims to detect the most salient object in a video without any external guidance regarding the object. Salient objects often exhibit distinctive movements compared to the background, and recent methods leverage this by combining motion cues from optical flow maps with appearance cues from RGB images. However, because optical flow maps are often closely correlated with segmentation masks, networks can become overly dependent on motion cues during training, leading to vulnerability when faced with confusing motion cues and resulting in unstable predictions. To address this challenge, we propose a novel motion-as-option network that treats motion cues as an optional component rather than a necessity. During training, we randomly input RGB images into the motion encoder instead of optical flow maps, which implicitly reduces the network's reliance on motion cues. This design ensures that the motion encoder is capable of processing both RGB images and optical flow maps, leading to two distinct predictions depending on the type of input provided. To make the most of this flexibility, we introduce an adaptive output selection algorithm that determines the optimal prediction during testing.
Problem

Research questions and friction points this paper is trying to address.

Reducing reliance on motion cues in segmentation
Handling confusing motion cues for stable predictions
Adaptively selecting predictions from multiple input types
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion-as-option network reduces reliance on motion cues
Random RGB input to motion encoder for flexibility
Adaptive output selection algorithm optimizes predictions
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Yonsei University | Kyung Hee University
Suhwan Cho
Suhwan Cho
AI/ML Research Scientist @ GenGenAI
Video Object SegmentationVideo InpaintingVideo GenerationVideo Understanding
Minhyeok Lee
Minhyeok Lee
Yonsei University
Computer Vision and Pattern Recognition
Jungho Lee
Jungho Lee
Yonsei University
Computer vision
M
Myeongah Cho
Department of Software Convergence, Kyung Hee University, Yongin 17104, Korea
S
Sangyoun Lee
School of Electrical and Electronic Engineering, Yonsei University, Seoul 03722, Korea