Rethinking Feature Reliance Evaluation with Semantically Matched Suppression

📅 2026-07-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing methods for assessing feature dependence, which rely on unnatural conflicting stimuli and thus fail to reflect a model’s genuine reliance on shape versus texture during standard recognition. To overcome this, the authors propose a semantically aligned feature suppression framework that fairly quantifies a model’s dependence on shape and texture while maintaining consistent class-discriminative loss. Applying this framework to compare Vision Transformers (ViTs) and CNNs reveals that ImageNet-pretrained CNNs suffer significant performance degradation under texture suppression, whereas ViTs exhibit consistently higher robustness under both shape and texture suppression. Moreover, ViT internal representations demonstrate greater stability in brain encoding tasks, highlighting their superior feature robustness and stronger alignment with human neural representations.
📝 Abstract
Understanding whether visual recognition models rely on shape, texture, or color is central to interpreting their behavior. Prior cue-conflict studies have strongly influenced the view that CNNs are texture-biased, yet such tests measure cue preference under artificial conflicts rather than feature reliance during natural recognition. We revisit this question through controlled feature suppression and show that performance drops are difficult to interpret unless different suppression operations impose comparable category-level damage. We introduce a semantically matched evaluation framework that compares shape and texture suppression at matched levels of category separability loss. Under this framework, ImageNet-trained CNNs show stronger degradation under texture suppression than under shape suppression, revealing greater texture reliance than suggested by unmatched suppression analyses. Extending the comparison across architectures, we find that Vision Transformers retain higher accuracy than CNNs under both shape and texture suppression. Brain encoding further shows that ViT representations exhibit smaller suppression-induced decreases in neural prediction performance under the tested suppression settings. These findings indicate that semantic comparability is essential for interpreting feature reliance from suppression experiments, and suggest that the robustness advantage of ViTs may be related to representations more compatible with human visual cortex.
Problem

Research questions and friction points this paper is trying to address.

feature reliance
cue-conflict
feature suppression
semantic comparability
visual recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

feature reliance
semantically matched suppression
texture bias
Vision Transformers
brain encoding
N
Ning Jiang
Institute of Medical Technology, Peking University Health Science Center, Beijing 100191, China; National Institute of Health Data Science, Peking University, Beijing 100191, China
T
Tianyi Luo
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Z
Zhengyong Huang
Institute of Medical Technology, Peking University Health Science Center, Beijing 100191, China; National Institute of Health Data Science, Peking University, Beijing 100191, China
Yao Sui
Yao Sui
Harvard Medical School
Computer visionMachine learningNeuroimagingMedical image computing