Score
Design and implement training procedures, loss terms, and network components that detect and suppress transient or spurious distractors in a model’s attention maps and feature representations. This includes adding auxiliary mask-prediction heads or mask-supervised attention losses, using distractor-mask supervision and cross-view or consistency constraints to separate clean from contaminated features and enforce that attention ignores distractor regions.
In vision Transformers, atypical backgrounds induce contextual bias, undermining model robustness and interpretability. To address this, we propose a two-stage differentiable binary attention framework: first, it localizes task-relevant image regions; second, it enforces prediction exclusively from these regions via an end-to-end learned binary mask, effectively suppressing background interference. Our core contributions are (i) intrinsic faithfulness of attention maps—attention weights directly and exclusively govern input visibility—and (ii) a joint optimization mechanism that synergistically enhances region discovery and focused prediction. Evaluated on object-centric tasks—including fine-grained recognition and joint localization-classification—the method significantly improves robustness against spurious correlations and out-of-distribution backgrounds, while boosting localization accuracy and prediction consistency.
Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.
Existing vision models lack subject-awareness, making it difficult to accurately identify and remove distractors in image editing without compromising scene semantic consistency. This work formalizes, for the first time, the task of Subject-Aware Distractor Localization (SADL) and introduces the first real-world benchmark for this task, comprising 1,800 cases with 14,617 annotated candidate objects. The authors propose a two-stage vision-language model (VLM) pipeline grounded in five inclusion factors and three contextual exclusion rules. Evaluation across seven VLMs reveals strong identification capabilities but exposes a systematic over-suppression bias during the exclusion phase. The SADL benchmark serves as a critical diagnostic tool for subject-conditioned reasoning in multimodal systems.
This work addresses the challenges of multi-concept disentanglement, asynchronous feature fusion and learning, and poor structural alignment in single-image, text-to-image customization—particularly without manual masks. We propose a diffusion-based framework that enables mask-free, concept-aware image generation. Our method introduces: (1) an attention-driven, single-step automatic masking mechanism that leverages self- and cross-attention maps for concept-level segmentation; (2) Uniform and Reweighted sampling strategies to mitigate temporal inconsistency during multi-concept feature extraction; and (3) native support for complex scenarios involving three or more concepts. Extensive experiments demonstrate that our approach preserves strong text-image alignment while significantly improving structural fidelity and concept disentanglement on single images. Ablation studies confirm the effectiveness and robustness of each component, establishing state-of-the-art performance in mask-free, multi-concept customization tasks.
Vision-language models (e.g., CLIP) often rely on spurious correlations—so-called “decision shortcuts”—in fine-grained image classification, leading to poor out-of-distribution generalization. To address this, we propose Test-time Prompt Erasure (TPE), a parameter-free inference-time method that dynamically identifies and suppresses non-causal shortcut features via learnable prompts, thereby strengthening task-invariant causal feature representations. TPE integrates test-time prompt tuning, feature disentanglement analysis, and a gradient-driven masking mechanism for spurious features. Crucially, it requires no model fine-tuning or architectural modification. Evaluated across multiple fine-grained classification benchmarks, TPE reduces average classification error by 12.7% compared to strong baselines, significantly outperforming existing test-time adaptation approaches. It markedly improves model robustness and cross-domain generalization without additional training overhead.
Vision Transformers are prone to background distractions and often rely on spurious correlations for prediction. To address this, this work proposes Inhibitory Self-Attention (ISA), inspired by biological visual suppression mechanisms. ISA is the first self-attention variant to explicitly retain and leverage negative attention scores, circumventing the Softmax-induced constraint of non-negative outputs. By directly suppressing irrelevant features, ISA enhances the model’s selective focus on semantically meaningful regions. Integrated into the Vision Transformer architecture, ISA demonstrates consistent performance gains and improved out-of-distribution generalization across ImageNet-1k, COCO, and multiple robustness benchmarks. Attention visualizations further confirm that ISA produces sharper, more target-concentrated attention maps compared to standard self-attention.
研究通过在softmax中引入学习的每头sink logit提供弃权机制,以及在每个值上设置门控来过滤噪声,解决了注意力机制中缺乏弃权和噪声过滤的问题。
本文通过将真实缺陷掩码作为训练信号,并结合扩散模型增强,改进了工业检测中缺陷定位的精度。
This work addresses the challenge of removing specified objects in dense scenes, where existing methods often suffer from semantic interference caused by visually similar instances, leading to incomplete removal or duplicated artifacts. To overcome this limitation, the authors propose DORS, a diffusion-based object removal framework that introduces a novel dynamic attention routing mechanism comprising Instance Filtering Attention (IFA) and Context-Guided Routing (CGR). This mechanism enables fine-grained control over the attention space, effectively suppressing interference while preserving scene consistency. Furthermore, the study presents DOR-Bench, the first benchmark specifically designed for evaluating dense object removal. Experimental results demonstrate that DORS significantly outperforms current state-of-the-art methods, achieving superior removal completeness and visual coherence.
This study addresses the oversight of dissociated human attention and recognition mechanisms in existing AI image editing detection. We propose a two-stage cognitive model demonstrating that edited regions drive attentional capture while semantic plausibility determines judgment accuracy. As the first work to introduce pre-attentive and recognition distinctions into this domain, we construct a generative eye-movement prediction framework. Validated through eye-tracking and mixed-effects analyses, the significant dissociation between stages is confirmed. The model achieves attention prediction correlations of 0.77–0.82, and its missed-detection behavioral prediction performance (r=0.52) significantly outperforms linear baselines (r=0.48). These findings establish a novel paradigm for understanding detection blind spots in human-AI interaction, highlighting the critical role of cognitive separation in evaluating synthetic imagery.