Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

📅 2026-07-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in existing two-view (CC/MLO) breast cancer classification methods of disentangling view-specific and shared representations, as well as their limited cross-view interaction confined to a single layer. The authors propose a token-centric dual-view learning framework that introduces propagatable fusion tokens into a frozen Vision Transformer backbone, enabling progressive bidirectional interaction through multi-layer cross-attention while avoiding direct feature fusion. This design effectively balances view specificity and complementarity within a unified architecture that jointly handles prompt adaptation and cross-view fusion. Evaluated on the VinDr-Mammo and CMMD datasets, the method significantly outperforms current baselines, achieving a 50.40% F1-score and 0.8090 AUC for BI-RADS classification on VinDr-Mammo, along with a 0.10 absolute AUC improvement in binary classification.
📝 Abstract
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.
Problem

Research questions and friction points this paper is trying to address.

breast cancer classification
multi-view learning
mammography
CC and MLO views
feature fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

token-based fusion
dual-view learning
cross-attention
multi-depth interaction
prompt adaptation
🔎 Similar Papers
2024-08-29Medical Imaging 2025: Digital and Computational PathologyCitations: 1
A
Aysan Ghayouri Pirsoltan
School of Computer Engineering, Iran University of Science and Technology, Tehran, Iran
S
Shima Babakordi
School of Computer Engineering, Iran University of Science and Technology, Tehran, Iran
Mohammad Reza Mohammadi
Mohammad Reza Mohammadi
Assistant Professor of Computer Engineering, Iran University of Science and Technology
Computer VisionMachine LearningDeep Learning