Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing medical report generation methods rely on expert-annotated detection modules, incurring high annotation costs and exhibiting poor cross-dataset generalizability. To address this, we propose a self-supervised anatomical consistency learning framework that requires no expert annotations. Our approach constructs a hierarchical anatomical graph structure and jointly optimizes recursive region reconstruction and region-level contrastive learning. Furthermore, we introduce a prompt-driven attention guidance mechanism to achieve precise multimodal alignment between images and text in both anatomical space and semantic space. The proposed method significantly enhances report interpretability and clinical applicability: it improves vocabulary accuracy by 10%, increases clinical validity by 25%, and outperforms state-of-the-art vision foundation models by 8% on zero-shot visual localization.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologiesGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL) -- a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports -- outperforming state-of-the-art methods by 10% in lexical accuracy and 25% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8% in zero-shot visual grounding.
Problem

Research questions and friction points this paper is trying to address.

Generates medical reports from images without expert annotations
Aligns generated reports with anatomical regions using textual prompts
Improves interpretability through self-supervised anatomical consistency learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses annotation-free anatomical graph for spatial alignment
Employs region-level contrastive learning for semantic alignment
Generates reports using aligned embeddings as visual priors
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Longzhen Yang
Tongji University, Shanghai, China
Z
Zhangkai Ni
Tongji University, Shanghai, China
Ying Wen
Ying Wen
Associate Professor, Shanghai Jiao Tong University
Multi-Agent LearningReinforcement Learning
Y
Yihang Liu
Tongji University, Shanghai, China
L
Lianghua He
Tongji University, Shanghai, China; Shanghai Eye Disease Prevention and Treatment Center, Shanghai, China
H
Heng Tao Shen
Tongji University, Shanghai, China