$Δ$Representation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing medical vision-language models in effectively modeling the incremental features of pathological phenotypes relative to normal anatomical structures. To this end, we propose a counterfactual reasoning-based visual phenotype representation learning framework that disentangles lesion and normal anatomical representations through a counterfactual mechanism. The method integrates fine-grained geometric supervision with incremental computation to facilitate phenotype clustering, while incorporating spatial relationship modeling to optimize lesion localization. Experimental evaluations on the ReXGroundingCT and LIDC-IDRI datasets demonstrate that the proposed approach significantly improves both lesion localization accuracy and phenotype characterization performance.
📝 Abstract
Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf{$Δ$Representation}, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbf{BaseAnatomy}, a geometry-supervised representation learning module, and \textbf{$Δ$Phenotype}, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. $Δ$Phenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textit{ReXGroundingCT} and \textit{LIDC-IDRI} demonstrate that $Δ$Representation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at https://anonymous.4open.science/r/deltarep-CF6D.
Problem

Research questions and friction points this paper is trying to address.

Medical Vision-Language Models
Pathological Phenotype Representation
Counterfactual Reasoning
Lesion Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Reasoning
Geometry Supervision
Representation Learning
Medical VLMs
Phenotype Characterization
🔎 Similar Papers
No similar papers found.