distractor-aware training

Design and implement training procedures, loss terms, and network components that detect and suppress transient or spurious distractors in a model’s attention maps and feature representations. This includes adding auxiliary mask-prediction heads or mask-supervised attention losses, using distractor-mask supervision and cross-view or consistency constraints to separate clean from contaminated features and enforce that attention ignores distractor regions.

distractor-awaretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Inherently Faithful Attention Maps for Vision Transformers

Jun 10, 2025
AA
Ananthu Aniraj
🏛️ Inria | University of Montpellier | Inrae

In vision Transformers, atypical backgrounds induce contextual bias, undermining model robustness and interpretability. To address this, we propose a two-stage differentiable binary attention framework: first, it localizes task-relevant image regions; second, it enforces prediction exclusively from these regions via an end-to-end learned binary mask, effectively suppressing background interference. Our core contributions are (i) intrinsic faithfulness of attention maps—attention weights directly and exclusively govern input visibility—and (ii) a joint optimization mechanism that synergistically enhances region discovery and focused prediction. Evaluated on object-centric tasks—including fine-grained recognition and joint localization-classification—the method significantly improves robustness against spurious correlations and out-of-distribution backgrounds, while boosting localization accuracy and prediction consistency.

Addresses biased representations from out-of-distribution backgroundsEnsures only attended image regions influence predictionsImproves robustness against spurious correlations in object-centric tasks

MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning

Nov 16, 2025
JH
Jingshan Hong
🏛️ Zhejiang University of Technology | Zhejiang Normal University

Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.

Addresses underutilization of discarded pixels in supervised image maskingExploits masked regions as semantic diversity sources rather than ignored dataSolves loss of fine-grained features caused by traditional masking methods

Existing vision models lack subject-awareness, making it difficult to accurately identify and remove distractors in image editing without compromising scene semantic consistency. This work formalizes, for the first time, the task of Subject-Aware Distractor Localization (SADL) and introduces the first real-world benchmark for this task, comprising 1,800 cases with 14,617 annotated candidate objects. The authors propose a two-stage vision-language model (VLM) pipeline grounded in five inclusion factors and three contextual exclusion rules. Evaluation across seven VLMs reveals strong identification capabilities but exposes a systematic over-suppression bias during the exclusion phase. The SADL benchmark serves as a critical diagnostic tool for subject-conditioned reasoning in multimodal systems.

distractor localizationimage compositionsemantic coherence

This work addresses the challenges of multi-concept disentanglement, asynchronous feature fusion and learning, and poor structural alignment in single-image, text-to-image customization—particularly without manual masks. We propose a diffusion-based framework that enables mask-free, concept-aware image generation. Our method introduces: (1) an attention-driven, single-step automatic masking mechanism that leverages self- and cross-attention maps for concept-level segmentation; (2) Uniform and Reweighted sampling strategies to mitigate temporal inconsistency during multi-concept feature extraction; and (3) native support for complex scenarios involving three or more concepts. Extensive experiments demonstrate that our approach preserves strong text-image alignment while significantly improving structural fidelity and concept disentanglement on single images. Ablation studies confirm the effectiveness and robustness of each component, establishing state-of-the-art performance in mask-free, multi-concept customization tasks.

Addresses feature fusion and asynchronous learning issuesDisentangles multiple concepts from single imagesEnables balanced concept acquisition without manual masks

Vision-language models (e.g., CLIP) often rely on spurious correlations—so-called “decision shortcuts”—in fine-grained image classification, leading to poor out-of-distribution generalization. To address this, we propose Test-time Prompt Erasure (TPE), a parameter-free inference-time method that dynamically identifies and suppresses non-causal shortcut features via learnable prompts, thereby strengthening task-invariant causal feature representations. TPE integrates test-time prompt tuning, feature disentanglement analysis, and a gradient-driven masking mechanism for spurious features. Crucially, it requires no model fine-tuning or architectural modification. Evaluated across multiple fine-grained classification benchmarks, TPE reduces average classification error by 12.7% compared to strong baselines, significantly outperforming existing test-time adaptation approaches. It markedly improves model robustness and cross-domain generalization without additional training overhead.

Adaptability LimitationsCLIPVisual Language Models

Latest Papers

What's happening recently
View more

Vision Transformers are prone to background distractions and often rely on spurious correlations for prediction. To address this, this work proposes Inhibitory Self-Attention (ISA), inspired by biological visual suppression mechanisms. ISA is the first self-attention variant to explicitly retain and leverage negative attention scores, circumventing the Softmax-induced constraint of non-negative outputs. By directly suppressing irrelevant features, ISA enhances the model’s selective focus on semantically meaningful regions. Integrated into the Vision Transformer architecture, ISA demonstrates consistent performance gains and improved out-of-distribution generalization across ImageNet-1k, COCO, and multiple robustness benchmarks. Attention visualizations further confirm that ISA produces sharper, more target-concentrated attention maps compared to standard self-attention.

background distractionobject-centric focusself-attention

This work addresses the challenge of removing specified objects in dense scenes, where existing methods often suffer from semantic interference caused by visually similar instances, leading to incomplete removal or duplicated artifacts. To overcome this limitation, the authors propose DORS, a diffusion-based object removal framework that introduces a novel dynamic attention routing mechanism comprising Instance Filtering Attention (IFA) and Context-Guided Routing (CGR). This mechanism enables fine-grained control over the attention space, effectively suppressing interference while preserving scene consistency. Furthermore, the study presents DOR-Bench, the first benchmark specifically designed for evaluating dense object removal. Experimental results demonstrate that DORS significantly outperforms current state-of-the-art methods, achieving superior removal completeness and visual coherence.

attention mechanismdense scenesincomplete removal

This study addresses the oversight of dissociated human attention and recognition mechanisms in existing AI image editing detection. We propose a two-stage cognitive model demonstrating that edited regions drive attentional capture while semantic plausibility determines judgment accuracy. As the first work to introduce pre-attentive and recognition distinctions into this domain, we construct a generative eye-movement prediction framework. Validated through eye-tracking and mixed-effects analyses, the significant dissociation between stages is confirmed. The model achieves attention prediction correlations of 0.77–0.82, and its missed-detection behavioral prediction performance (r=0.52) significantly outperforms linear baselines (r=0.48). These findings establish a novel paradigm for understanding detection blind spots in human-AI interaction, highlighting the critical role of cognitive separation in evaluating synthetic imagery.

AI image editsattention capturedetectability

Hot Scholars

SS

Swakkhar Shatabda

Professor, School of Data and Sciences, BRAC University
optimizationmachine learningcomputational biologybioinformatics
HD

Haodong Duan

Shanghai AI Lab | CUHK | PKU
Computer VisionVideo UnderstandingMultimodal LearningGenerative AI
YN

Yuansheng Ni

University of Waterloo
Artificial IntelligenceNatural Language ProcessingLarge Language Models
WC

Wenhu Chen

Assistant Professor at University of Waterloo
Natural Language ProcessingArtificial IntelligenceDeep Learning