Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the localization lag and false positive issues in referring video object segmentation caused by fixed keyframes in long dynamic videos. To this end, we propose an Event-Driven Refresh and Recurrent Memory (EDRRM) mechanism. EDRRM introduces a novel trigger boundary based on EMA-smoothed event scores to precisely capture target transition points and selectively re-invoke the model. Furthermore, it designs an anchor frame retrieval strategy that integrates CLIP similarity with tracking cues to enable efficient recurrent memory. This approach optimizes the Sa2VA inference pipeline, significantly improving the J&F metric while effectively suppressing false positives on benchmarks such as Ref-DAVIS17. By achieving these gains with reduced computational overhead, the proposed method establishes a superior balance between performance and efficiency.
📝 Abstract
Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of initial keyframes and then relies on propagation. In long or dynamic videos, this can cause stale grounding and persistent false positives when the object composition changes (e.g., distractors enter or the target disappears/re-appears). We propose Event-Driven Refresh + Recurrence Memory (EDRRM), an enhancement that selectively re-invokes Sa2VA only at stable change points. EDRRM triggers refresh boundaries using an EMA-smoothed event score computed from tracking-derived cues (births/deaths and coarse composition/layout changes) with temporal constraints. A recurrence memory further retrieves anchor frames via CLIP similarity to re-condition the model on re-appearance events. Experiments on Ref-DAVIS17, MeViS, and ReVOS show that EDRRM achieves a competitive accuracy-efficiency trade-off relative to fixed-window and FrameDiff-SSIM baselines, maintaining comparable or superior J&F scores at substantially lower average refresh-call budgets and reducing false-positive failures. End-to-end runtime analysis further confirms that the overhead introduced by tracking, CLIP-based recurrence matching, and the identifiability gate remains modest relative to the dominant Sa2VA inference cost, thereby validating the efficiency of the proposed pipeline.
Problem

Research questions and friction points this paper is trying to address.

Referring Video Object Segmentation
Stale Grounding
False Positives
Long Videos
Dynamic Videos
Innovation

Methods, ideas, or system contributions that make the work stand out.

Referring Video Object Segmentation
Event-Driven Refresh
Recurrence Memory
Stale Grounding
Tracking-derived Cues
🔎 Similar Papers
No similar papers found.
A
Abu Hanif Muhammad Syarubany
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
J
Jaehyun Jang
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
S
Siwoo Lim
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
S
Seungyeon Ryu
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
Chang D. Yoo
Chang D. Yoo
kaist
machine learningcomputer visionsignal processingspeech enhancementspeech recognition