🤖 AI Summary
This work addresses the high memory overhead of cross-attention and redundant full-frame feature extraction in SAM2 for temporally promptable video segmentation, challenges that cause existing lightweight approaches to suffer significant performance degradation in complex scenes. To tackle these issues, the authors propose Lean-SAM2, a framework integrating three synergistic strategies: Target-Anchored Memory Pruning (TAMP), Temporal Compression with Insurance Mechanism (TCIM), and Temporal Risk-Aware Routing (TARR). These components collectively eliminate computational redundancy while preserving segmentation robustness. Evaluated on LVOSv2, Lean-SAM2 achieves speedups of 1.412× and 1.417× over SAM2.1-Large and Base+, respectively, along with J&F score improvements of 5.0% and 3.6%, substantially outperforming Efficient-SAM2.
📝 Abstract
The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains hindered by heavy memory cross-attention overhead and redundant full-frame visual feature extraction. While recent methods explore efficiency via heuristic memory pruning and window-based sparse routing, they typically suffer from catastrophic performance degradation in complex segmentation scenarios replete with occlusions and distractors. To resolve these limitations, we propose \textbf{Lean-SAM2}, a holistic lightweight framework designed to address the above vulnerabilities while systematically eliminating computational redundancies. Specifically, Lean-SAM2 integrates three collaborative mechanisms: (1) Target-Anchored Memory Pruning (TAMP) safeguards target tokens against deceptive attention by modulating raw attention significance with semantic consistency against prompt-derived foreground anchors; (2) Temporal Condensation with Insurance Memory (TCIM) condenses historical context via a visibility-gated fusion while conditionally archiving high-confidence entries in a parallel insurance bank; and (3) Target-Anchored Risk-Aware Routing (TARR) selectively activates the heavy image encoder for target-related windows based on anchor similarity, utilizing a risk-aware fallback policy to trigger full-frame refreshes during volatile transitions. Extensive evaluations across multiple challenging benchmarks demonstrate that Lean-SAM2 establishes a superior balance between accuracy and efficiency. For example, on the LVOSv2 validation dataset, Lean-SAM2 achieves overall inference speedups of $1.412\times$ and $1.417\times$ on the SAM2.1-Large and SAM2.1-Base+, respectively, significantly outperforming Efficient-SAM2 while boosting the corresponding $\mathcal{J}\&\mathcal{F}$ scores by $5.0\%$ and $3.6\%$. Code is available at https://github.com/DeawhaleQwQ/Lean-SAM2.