When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in existing video object segmentation (VOS) benchmarks, which neglect low temporal visibility scenarios and thereby allow models to obtain spurious rewards through empty predictions, rendering evaluation metrics ineffective. To overcome this issue, we introduce the FaVOS benchmark and propose Volumetric J&F, the first spatiotemporal volumetric evaluation metric for VOS. By leveraging spatiotemporal volume modeling and rigorous benchmark design, our approach effectively exposes and mitigates the missing-frame reward trap. Experimental results demonstrate that the proposed metric corrects biases inherent in conventional evaluations and faithfully reflects model segmentation performance under complex temporal conditions. Ultimately, this work establishes a more reliable evaluation paradigm for the VOS community.
📝 Abstract
Video Object Segmentation (VOS) in complex and long videos is increasingly important for real-world applications, where target objects often appear only intermittently within long temporal horizons. However, existing benchmarks largely focus on temporally salient objects that remain visible for most of the video. To address this gap, we introduce FaVOS (A Benchmark for Video Object Segmentation with Fractional Temporal Visibility), a benchmark designed to evaluate VOS methods under low temporal visibility. We show that, in this regime, the standard J&F metric can collapse VOS evaluation into absence classification, because empty predictions receive high rewards on target-absent frames. Consequently, even a trivial empty-mask predictor can outperform strong models such as SAM 3, revealing a fundamental mismatch between current metrics and practical VOS performance. To mitigate this issue, we propose Volumetric J&F, which evaluates mask sequences as spatio-temporal volumes and reduces the dominance of target-absence rewards while preserving sensitivity to segmentation quality and temporal structure. Project page: https://aidaslab.github.io/FaVOS.
Problem

Research questions and friction points this paper is trying to address.

Video Object Segmentation
Evaluation Metric
Temporal Visibility
J&F Metric
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Object Segmentation
Evaluation Metric
Volumetric J&F
Temporal Visibility
Benchmark
🔎 Similar Papers
No similar papers found.
J
Jihwan Hong
AIDAS Laboratory, IPAI & ECE, Seoul National University
W
Woohyeon Park
AIDAS Laboratory, IPAI & ECE, Seoul National University
J
Jaeik Kim
AIDAS Laboratory, IPAI & ECE, Seoul National University
Jaeyoung Do
Jaeyoung Do
Department of Electrical and Computer Engineering, Seoul National University
Generative AI (LLMs)Multi-Modal AI (NLP/Vision)Big Data Systems