Depth-Guided Video Object Counting in Crowded Scenes

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of accurately counting objects in crowded and occluded scenes, where existing RGB-only video object counting methods struggle to distinguish individual instances. To overcome this limitation, we propose Depth-Guided Detector (DG-Det), the first approach to incorporate depth information into this task. Our method leverages a multi-scale RGB-D cross-attention mechanism to effectively fuse depth cues and jointly models explicit occlusion prediction to enhance spatial awareness. Additionally, we introduce a unified cross-frame deduplication framework that substantially reduces redundant counts. Our contributions include the first multi-category RGB-D video counting dataset with depth annotations, publicly released code, and significant performance gains—achieving a 62.01% reduction in mean absolute error (MAE) and a substantial improvement in root mean square error (RMSE).
📝 Abstract
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline. By integrating depth cues with multi-scale RGB-D cross-attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occluded scenes. Furthermore, we introduce a unified de-duplication framework to eliminate cross-frame redundant counting. To facilitate future research, we also release a new RGB-D Video Object Counting dataset featuring depth information and multiple object categories persequence. Extensive experiments demonstrate that our method achieves a 62.01\% reduction in MAE compared to existing baselines, and also produces consistent improvements in RMSE. We provide the source code at https://github.com/streamer-AP/DG-Net and the dataset at https://huggingface.co/datasets/aerospace123/RGBD-VideoCount.
Problem

Research questions and friction points this paper is trying to address.

video object counting
crowded scenes
occlusion
depth information
RGB-D
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-Guided Detection
RGB-D Cross-Attention
Occlusion Prediction
Cross-Frame De-duplication
Video Object Counting
🔎 Similar Papers
No similar papers found.
Y
Yuanjing Xu
Harbin Institute of Technology (Weihai), Weihai, Shandong, China
X
Xinyan Liu
Harbin Institute of Technology (Weihai), Weihai, Shandong, China; City University of Hong Kong, Hong Kong, China
W
Weidong Chen
University of Science and Technology of China, Hefei, Anhui, China
Z
Zixuan Zou
Harbin Institute of Technology (Weihai), Weihai, Shandong, China
L
Linhao Zhang
Harbin Institute of Technology (Weihai), Weihai, Shandong, China
Z
Zhuangzhe Meng
Harbin Institute of Technology (Weihai), Weihai, Shandong, China
Antoni B. Chan
Antoni B. Chan
Professor of Computer Science, City University of Hong Kong
Computer VisionMachine LearningSurveillanceEye Gaze AnalysisComputer Audition
Weigang Zhang
Weigang Zhang
Professor of Computer Science, Harbin Institute of Technology, Weihai
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer Vision