Clinical Trajectory Alignment for Medical Vision-Language Pre-training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing medical vision-language pretraining in capturing longitudinal clinical changes and the asynchrony of abnormality progression by proposing MedCTA. This framework introduces a novel multi-scale clinical trajectory alignment mechanism that models temporal dynamics from both lesion-level and patient-course perspectives. Furthermore, it leverages an offline large language model to parse structured trends as supervision signals, constructing abnormality-conditioned vision-text trajectories. This approach transcends conventional single-timestamp matching, enabling automated temporal feature learning without manual annotation. Experimental results demonstrate that MedCTA significantly outperforms strong baselines on temporal image classification, cross-modal retrieval, and zero-shot classification tasks, effectively encoding longitudinal evolution semantics.
📝 Abstract
Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.
Problem

Research questions and friction points this paper is trying to address.

Medical Vision-Language Pre-training
Clinical Trajectory Alignment
Longitudinal Clinical Change
Temporal Semantics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Medical Vision-Language Pre-training
Clinical Trajectory Alignment
Longitudinal Clinical Change
Abnormality-level Modeling
LLM-based Supervision