CrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd Counting

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过预训练专家模型的辅助指导提高生成式视觉-语言模型在人群计数任务中的准确性,提出CrowdCue方法,最佳变体达到了62.65的MAE。
📝 Abstract
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained specialist can lift them into useful territory, and through which channel that guidance is best routed. We evaluate Qwen2.5-VL-7B on four widely used crowd counting benchmarks (ShanghaiTech A and B, UCF-QNRF, NWPU-Crowd). Zero-shot prompting rarely produces a parseable count, so LoRA supervised fine-tuning establishes the baseline at overall MAE 81.64. Conditioning on a P2PNet-derived density heatmap as an auxiliary visual signal fails in every encoding we tested, and an adversarial-swap protocol shows the model reads the heatmap but applies it counterproductively. We propose CrowdCue, a family that supplies the same specialist's already-integrated integer count to the VLM as a discrete symbol. The text-channel variant reaches MAE 72.04. The visual-channel variant, which renders the integer as printed digits and supplies it as a second image, reaches MAE 62.65, the strongest result in this paper and well ahead of the cue-supplying specialist alone (84.45 on the same split). In the late-fusion VLM we study, the binding constraint is not the channel but the abstraction level at which the specialist signal is delivered.
Problem

Research questions and friction points this paper is trying to address.

Crowd Counting
Vision-Language Models
Specialist Guidance
Accuracy Improvement
Auxiliary Signal
Innovation

Methods, ideas, or system contributions that make the work stand out.

CrowdCue
vision-language model
specialist-cue conditioning
crowd counting
visual-channel variant
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Moshiur Farazi
Moshiur Farazi
University of Doha for Science and Technology, Australian National University
Computer VisionVision-Language ModelsApplied AI
B
Bekir Ciftler
University of Doha for Science and Technology, Doha, Qatar
A
Abdulhalim Dandoush
University of Doha for Science and Technology, Doha, Qatar
R
Reda Bendraou
University of Doha for Science and Technology, Doha, Qatar