SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of image-text misalignment, limited zero-shot generalization, and false negatives in contrastive learning arising from the verbosity of radiology reports. To overcome these issues, this work proposes an enhanced sentence-centric vision-language pretraining framework. Methodologically, a large language model is introduced to perform sentence-level structured mapping of reports, thereby expanding positive sample diversity, while a residual modulation mechanism is designed to adaptively optimize semantic feature representations. Extensive evaluations demonstrate that the proposed approach significantly improves generalization capabilities across multiple zero-shot analysis tasks on chest X-rays, outperforming existing mainstream methods.
📝 Abstract
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Pretraining
Zero-Shot Learning
Chest X-Ray Analysis
Contrastive Learning
False Negatives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Pretraining
Zero-Shot Learning
Sentence-Centric Framework
Contrastive Learning
Chest X-Ray Analysis
H
Hangyul Yoon
Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)
H
Hyungyung Lee
Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)
Edward Choi
Edward Choi
KAIST
Machine LearningArtificial IntelligenceHealthcare
Eunho Yang
Eunho Yang
KAIST
Machine LearningStatistics