Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

πŸ“… 2026-03-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing zero-shot object counting methods suffer from insufficient fine-grained quantity awareness and spatial sensitivity, and model adaptation often distorts features, compromising generalization. To address these limitations, this work proposes the QICA framework, which enhances performance through synergistic quantity-awareness and spatial cost aggregation. Specifically, it introduces numerical conditional prompts to align semantic recognition with quantitative reasoning and performs spatial aggregation on vision–text similarity graphs to preserve zero-shot transferability. Furthermore, a Synergistic Prompting Strategy (SPS) and a Cost Aggregation Decoder (CAD) are designed, combined with multi-level quantity alignment losses, to significantly strengthen both quantity and spatial perception without requiring visual exemplars. The method achieves state-of-the-art results on FSC-147 and demonstrates exceptional cross-domain generalization in zero-shot evaluations on CARPK and ShanghaiTech-A.

Technology Category

Computer Vision: Visual Reasoning & Symbolic RepresentationsKnowledge Representation and Reasoning: Qualitative ReasoningSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
πŸ“ Abstract
Zero-shot object counting (ZSOC) aims to enumerate objects of arbitrary categories specified by text descriptions without requiring visual exemplars. However, existing methods often treat counting as a coarse retrieval task, suffering from a lack of fine-grained quantity awareness. Furthermore, they frequently exhibit spatial insensitivity and degraded generalization due to feature space distortion during model adaptation.To address these challenges, we present \textbf{QICA}, a novel framework that synergizes \underline{q}uantity percept\underline{i}on with robust spatial \underline{c}ast \underline{a}ggregation. Specifically, we introduce a Synergistic Prompting Strategy (\textbf{SPS}) that adapts vision and language encoders through numerically conditioned prompts, bridging the gap between semantic recognition and quantitative reasoning. To mitigate feature distortion, we propose a Cost Aggregation Decoder (\textbf{CAD}) that operates directly on vision-text similarity maps. By refining these maps through spatial aggregation, CAD prevents overfitting while preserving zero-shot transferability. Additionally, a multi-level quantity alignment loss ($\mathcal{L}_{MQA}$) is employed to enforce numerical consistency across the entire pipeline. Extensive experiments on FSC-147 demonstrate competitive performance, while zero-shot evaluation on CARPK and ShanghaiTech-A validates superior generalization to unseen domains.
Problem

Research questions and friction points this paper is trying to address.

zero-shot object counting
quantity awareness
spatial insensitivity
feature distortion
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-shot object counting
quantity awareness
spatial aggregation
synergistic prompting
feature distortion mitigation
πŸ’Ό Related Jobs
No related jobs found.
D
Da Zhang
Northwestern Polytechnical University; Institute of Artificial Intelligence (TeleAI), China Telecom
Bingyu Li
Bingyu Li
School of Cyber Science and Technology, Beihang University
Internet Infrastructure SecurityCryptography
Feiyu Wang
Feiyu Wang
Fudan University
computer vision
Z
Zhiyuan Zhao
Institute of Artificial Intelligence (TeleAI), China Telecom
J
Junyu Gao
Northwestern Polytechnical University; Institute of Artificial Intelligence (TeleAI), China Telecom