🤖 AI Summary
This study addresses critical limitations in existing medical world models—particularly their deficiencies in causal reasoning, uncertainty quantification, and prospective validation—and proposes a systematic framework for building trustworthy models in clinical settings. Drawing on a comprehensive review of 1,455 publications and a focused analysis of 98 core studies, the work establishes the first clear theoretical boundaries and empirical criteria for such models, centering on four key capabilities: patient state representation, temporal dynamics modeling, intervention simulation, and physician-in-the-loop planning. Introducing clinical digital twins as an integrative paradigm, the framework incorporates structured narrative synthesis, multimodal physiological modeling, uncertainty calibration, and safety-constrained planning. The authors identify 14 rigorously defined empirical studies demonstrating preliminary feasibility in trajectory prediction and intervention comparison, while also highlighting key bottlenecks and pathways toward credible clinical deployment.
📝 Abstract
Medical world models offer a framework for extending medical artificial intelligence beyond static prediction by representing evolving patient states and modelling how they change over time and in response to clinical interventions. This Review defines the conceptual boundaries, technical foundations, application domains, and evidence requirements of the field through a structured narrative synthesis with reproducible evidence mapping.We screened 1,455 unique records and assembled a corpus of 98 sources, including 14 studies that met a strict empirical definition of a medical world model. The field is organised around four capabilities: patient state representation, temporal dynamics modelling, intervention-conditioned simulation, and clinician-supervised planning. Evidence spans medical imaging, longitudinal electronic health records, treatment response modelling, physiological and multimodal state modelling, ultrasound and surgical interaction, and population and health-system simulation; clinical digital twins are treated as a cross-cutting integration framework.Current studies provide early evidence of technical feasibility for trajectory forecasting and comparison of candidate interventions, but most remain retrospective, task-specific, or preclinical. The evidence base is further limited by incomplete longitudinal intervention data, inconsistent action semantics, limited causal identifiability, long-horizon error accumulation, inadequate uncertainty estimation, and limited external validation. Clinical translation will therefore depend on precise intervention representations, robust causal and mechanistic grounding, calibrated trajectory-level uncertainty, safety-constrained planning, and prospective multicentre validation against clinically meaningful endpoints.