š¤ AI Summary
This study addresses the challenge of distinguishing whether predictions in medical world models originate from patient-specific information or population-level average patterns. To this end, it proposes P³, a novel auditing framework that evaluates personalized evidence across three dimensions: patient history, spatial support, and gain beyond prior. Furthermore, the work develops Cancer JEPA, a model built upon the Joint Embedding Predictive Architecture that integrates low-rank regression with lesion-constrained neural correction to achieve precise prediction of dynamic contrast-enhanced breast MRI. Empirical evaluations reveal that while the model leverages patient-specific information to reduce prediction errors, certain performance gains do not significantly surpass population priors. These findings establish rigorous definitional criteria for personalization in medical artificial intelligence, clarifying the boundary between genuine individualized modeling and reliance on aggregate statistical patterns.
š Abstract
Longitudinal models forecast how a patient's imaging state evolves, but accuracy does not show whether the patient's observed trajectory drives the prediction. A population-average forecast may be useful but cannot establish a patient-specific world-model claim. We introduce Patient, Place, Prior (P$^3$), an audit asking whether a forecast benefits from the patient's longitudinal imaging history (Patient), benefits from patient-matched externally supplied spatial support (Place), and gains predictive value beyond a population-average prediction under matched support and context (Prior). We also propose Cancer JEPA, a one-step model that forecasts frozen representations of future breast dynamic contrast-enhanced MRI examinations during neoadjuvant therapy. It adds a lesion-constrained neural correction, trained with an occlusion-based latent objective, to a patient-conditioned low-complexity reduced-rank regression baseline. This factorization permits a post-hoc P$^3$ audit of the frozen model. In a validation cohort previously used in development, forecast error is lower when the neural correction receives the patient's history rather than another patient's and patient-matched lesion occupancy maps rather than substituted maps. However, the descriptive 95% interval comparing the correction computed from patient history with the population-average neural correction includes zero. P$^3$ thus separates input use from evidence of patient-specific predictive value beyond a population-level pattern.