Score
Design and build methods that produce per-pixel spatial localization maps from images by tracing classifier activations or image-level labels to pixels (e.g., class activation mapping, localization extraction, and weakly-supervised localization), including techniques that recover object locations from activations and handle open-set cases. Develop and analyze noise-robust objectives, weak-label-robust training, localization metrics and probing procedures to encourage sharp, spatially grounded predictions and to evaluate localization reliability under label noise, severe class imbalance, and other adversities.
Weakly supervised semantic segmentation (WSSS) suffers from incomplete object localization, blurred boundaries, and bias toward discriminative local regions due to reliance solely on image-level labels. To address these limitations, this paper proposes an instance-guided influence function modeling framework with three key innovations: (1) leveraging instance-level cues to guide Class Activation Map (CAM) generation, thereby improving object completeness; (2) modeling pixel-wise contributions to classification decisions via influence functions to enhance boundary sensitivity; and (3) integrating multi-scale progressive optimization with Conditional Random Field (CRF) post-processing for fine-grained structural recovery. Evaluated on PASCAL VOC 2012, the method achieves 82.3% mean Intersection-over-Union (mIoU), rising to 86.6% after CRF refinement—outperforming state-of-the-art WSSS approaches. Notably, it delivers substantial improvements in object completeness and boundary delineation.
This paper addresses open-world object localization: training models with bounding-box supervision for only a limited set of categories, while requiring them to localize *all* objects—including unseen categories—at inference time. To tackle this challenge, we propose a background-driven object proposal learning paradigm, which—uniquely—treats background discovery as explicit supervisory signal. We formally define background as redundant, low-discriminative image regions and model “objectness” via inverse constraints. Our approach comprises three components: (i) a region-discriminativeness-based background discovery module; (ii) a background suppression loss; and (iii) an end-to-end trainable object proposal network. Evaluated on standard benchmarks, our method significantly outperforms state-of-the-art approaches, achieving substantial gains in both overall and unseen-category recall, as well as cross-category generalization capability.
Traditional Class Activation Mapping (CAM) activates only the most discriminative regions of the target object, resulting in incomplete coverage and poorly defined boundaries—severely limiting performance in pixel-level weakly supervised semantic segmentation (WSSS). To address this, we propose Region-CAM, the first method to introduce Semantic Information Propagation (SIP), which fuses multi-layer features and gradient signals to construct a Semantic Information Map (SIM). Region-CAM further enhances activation map completeness and boundary alignment accuracy via cross-stage semantic propagation and aggregation. The approach requires no additional supervision and is fully compatible with standard classification backbones. On PASCAL VOC 2012, Region-CAM achieves 60.12% mIoU—outperforming baseline CAM by 13.61 percentage points. On MS COCO, it improves mIoU by 16.23% and surpasses LayerCAM by 4.5% in localization accuracy (Loc1).
Small-object detection suffers from significant performance degradation due to inaccurate localization and unstable gradients. This paper identifies that conventional regression-based bounding box localization induces distorted gradients for small objects. To address this, we reformulate bounding box localization as a grid-based classification task—the first such approach—and propose a confidence-driven localization framework. Our method employs two-hot label encoding, confidence distribution prediction, cross-entropy localization loss, and an entropy-based uncertainty loss to jointly correct gradient flow and suppress localization uncertainty. Evaluated on mainstream detectors—including YOLOv8 and RT-DETR—and across three major benchmarks (COCO, VisDrone, and AI-TOD), our approach achieves state-of-the-art performance, notably improving AP for small objects. Moreover, it demonstrates strong generalization across diverse annotation protocols and high-resolution imagery.
To address low localization accuracy in weakly supervised video object localization (WSVOL) for long videos and objects with subtle motion, this paper proposes a collaborative class activation mapping (CAM) method that imposes no inter-frame positional constraints. Our key innovation is the first incorporation of object color consistency as a Conditional Random Field (CRF) loss term into CAM training, enabling direct cross-frame and cross-pixel response constraints and localization refinement—thereby significantly improving robustness to long-range temporal dependencies. The method jointly leverages CAMs, color-space constraints, and a weakly supervised collaborative localization framework, without requiring optical flow or explicit motion modeling. Evaluated on unconstrained video datasets such as YouTube-Objects, our approach achieves new state-of-the-art performance, particularly excelling in scenarios involving large object displacements and extended temporal sequences.
Existing self-attention and related models lack a unified theoretical framework, hindering systematic understanding and extension. This work proposes the “localization method”—a general machine learning framework grounded in localization kernels and local averages—that formalizes local modeling through rigorous theoretical constructs and localization techniques. For the first time, it subsumes state-of-the-art architectures such as Transformers under a unified local modeling paradigm. The framework not only reveals intrinsic connections among diverse models—including kernel methods, MeanShift, Hopfield networks, Locally Linear Embedding (LLE), fuzzy inference, denoising autoencoders, and Transformers—but also introduces scalable hierarchical local and non-local models, thereby establishing a novel paradigm for building data-adaptive learning systems.
Existing attribution methods for image geolocation models struggle to reveal whether predictions rely on human-interpretable, object-level visual cues. This work proposes an object-centric analysis pipeline that first extracts salient regions from attribution maps such as Grad-CAM, then decomposes them into object-like elements using image segmentation. The predictive relevance of these elements is rigorously evaluated through crop-based deletion and insertion tests. This approach enables, for the first time, an object-level interpretation of attribution outcomes. Experiments across three benchmark datasets demonstrate that attribution-guided cropping preserves significantly more predictive information than random cropping, providing strong evidence that geolocation models indeed leverage localized, interpretable object-level cues in their decision-making process.
This study addresses the critical challenge of rapidly locating specific targets in complex visualizations—a fundamental problem in human-computer interaction. Through a large-scale user experiment, it systematically evaluates the efficiency of common visual variables, such as color and size, in single-target localization tasks, offering the first empirically grounded performance comparison within the visualization literature. Introducing the concept of “localization robustness” and integrating eye-tracking and response time data, the research reveals that all visual variables are adversely affected by the number of distractors in displays containing hundreds of objects, thereby challenging the prevailing assumption of preattentive pop-out. The findings demonstrate significant differences in robustness across visual variables under varying target positions and layout configurations, providing empirical foundations and actionable design guidelines for effective visualization.
This work addresses a critical limitation in existing personalized object localization methods, which assume that every query image contains the target object and thus suffer from high false-positive rates when confronted with real-world negative samples lacking the target. To tackle this challenge, the paper introduces a new task—Personalized Object Identification and Localization (POIL)—and presents the first dedicated dataset for it. The authors propose IPLoc-ID, an algorithm that unifies object localization and instance verification within a single autoregressive framework. By integrating contextual reasoning from vision-language models, bounding box prediction, and a self-supervised query mechanism, IPLoc-ID effectively suppresses spurious detections while preserving high localization accuracy, thereby resolving the practical challenges of instance-level recognition and localization in realistic scenarios.
This study addresses the susceptibility of online high-definition map construction to localization errors when using prior maps as supervisory labels, which can lead to label distortion. By introducing three types of synthetic localization noise—Ramp, Gaussian, and Perlin—into the Argoverse 2 dataset, the authors train MapTRv2 variants to systematically quantify, for the first time, the impact of different localization errors on label quality. The findings reveal that heading angle errors are more detrimental than positional errors, with their adverse effects magnified at greater distances. To better align evaluation with real-world driving requirements, a distance-weighted metric is proposed. Experiments demonstrate that model performance degrades superlinearly with increasing noise levels, while access to undistorted ground-truth data significantly enhances performance.