Fast SAM2 with Text-Driven Token Pruning

📅 2025-12-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high computational and memory overhead of SAM2 in video object segmentation caused by dense visual tokens, this paper proposes a text-guided post-encoding visual token pruning framework. Without modifying the original model architecture, the method dynamically evaluates and retains critical tokens by jointly leveraging local visual context, text–visual semantic alignment, and uncertainty modeling. It introduces, for the first time, user- or auto-generated textual prompts into the post-encoding importance scoring process, enabling semantic-aware early sparsification prior to temporal propagation. Key design elements include a lightweight routing mechanism, multi-source importance fusion, and decoupling of the image encoder from the memory module. Experiments across multiple benchmarks demonstrate a 42.50% speedup in inference latency and a 37.41% reduction in GPU memory consumption, while maintaining near-equivalent J&F scores—significantly enhancing the deployability of Transformer-based video segmentation models in real-time and edge-computing scenarios.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.Search and Optimization: Sampling/Simulation-based Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSecurity and Privacy: Large-scale security measurements
📝 Abstract
Segment Anything Model 2 (SAM2), a vision foundation model has significantly advanced in prompt-driven video object segmentation, yet their practical deployment remains limited by the high computational and memory cost of processing dense visual tokens across time. The SAM2 pipelines typically propagate all visual tokens produced by the image encoder through downstream temporal reasoning modules, regardless of their relevance to the target object, resulting in reduced scalability due to quadratic memory attention overhead. In this work, we introduce a text-guided token pruning framework that improves inference efficiency by selectively reducing token density prior to temporal propagation, without modifying the underlying segmentation architecture. Operating after visual encoding and before memory based propagation, our method ranks tokens using a lightweight routing mechanism that integrates local visual context, semantic relevance derived from object-centric textual descriptions (either user-provided or automatically generated), and uncertainty cues that help preserve ambiguous or boundary critical regions. By retaining only the most informative tokens for downstream processing, the proposed approach reduces redundant computation while maintaining segmentation fidelity. Extensive experiments across multiple challenging video segmentation benchmarks demonstrate that post-encoder token pruning provides a practical and effective pathway to efficient, prompt-aware video segmentation, achieving up to 42.50 percent faster inference and 37.41 percent lower GPU memory usage compared to the unpruned baseline SAM2, while preserving competitive J and F performance. These results highlight the potential of early token selection to improve the scalability of transformer-based video segmentation systems for real-time and resource-constrained applications.
Problem

Research questions and friction points this paper is trying to address.

Reduces computational and memory costs in SAM2 video segmentation
Prunes irrelevant visual tokens using text guidance before temporal propagation
Maintains segmentation accuracy while improving inference speed and efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-guided token pruning reduces token density
Lightweight routing integrates visual, textual, and uncertainty cues
Early token selection improves inference speed and memory efficiency
🔎 Similar Papers
No similar papers found.
A
Avilasha Mandal
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
A
Avilasha Mandal
Department of Computer Science and Engineering, Indian Institute of Technology, Delhi, New Delhi, 110016, Delhi, India
Chaoning Zhang
Chaoning Zhang
Professor at UESTC (电子科技大学, China)
Computer VisionLLM and VLMGenAI and AIGC Detection
F
Fachrina Dewi Puspitasari
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
X
Xudong Wang
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
J
Jiaquan Zhang
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
C
Caiyan Qin
School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen, 518055, Guangdong, China
G
Guoqing Wang
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
Y
Yang Yang
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China
H
Heng Tao Shen
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 610054, Sichuan, China