Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing controllable surgical video generation methods often suffer from anatomical distortions, instrument appearance drift, and temporal inconsistencies due to the direct fusion of heterogeneous visual conditions. To address these issues, this work proposes a hierarchical surgical anchoring mechanism that integrates anchor-relative multimodal experts—each dedicated to processing edges, depth, and optical flow—and employs a staged fusion strategy to deliver precise control signals to a video diffusion model. Built upon the Wan2.2 backbone and evaluated on the Cholec80-SurgWAM benchmark, the proposed approach significantly outperforms current baselines in generation quality, temporal coherence, and multimodal controllability. It represents the first unified surgical world modeling framework that simultaneously preserves anatomical structure, ensures complementary information integration, and maintains consistent cross-modal interactions.
📝 Abstract
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.
Problem

Research questions and friction points this paper is trying to address.

surgical world model
multimodal control
anatomical distortion
temporal consistency
controllable video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Surgical World Model
Multimodal Control
Hierarchical Surgical Anchor
Anchor-Relative Modality Experts
Controllable Video Generation
🔎 Similar Papers
No similar papers found.
Rulin Zhou
Rulin Zhou
The Chinese University of Hong Kong Shenzhen Research Institute
Deep LearningMedical Image Processing
Wanhao Liu
Wanhao Liu
University of Science and Technology of China
AI for Science
G
Guoheng Ma
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR 999077, China and Shenzhen Loop Area Institute, Shen Zhen 518055, China
L
Liangjin Shao
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR 999077, China and Shenzhen Loop Area Institute, Shen Zhen 518055, China
Q
Qiujie Song
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR 999077, China and Shenzhen Loop Area Institute, Shen Zhen 518055, China
Y
Yidu Wang
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR 999077, China and Shenzhen Loop Area Institute, Shen Zhen 518055, China
Guankun Wang
Guankun Wang
The Chinese University of Hong Kong
Computer visionImage analysis
T
Tong Chen
School of Electrical and Computer Engineering, Faculty of Engineering, the University of Sydney, Sydney, NSW 2006, Australia
Long Bai
Long Bai
Research Assistant, Institute of Computing Technology, Chinese Academy of Sciences
Event-Centric AnalysisKnowledge GraphNatural Language Processing
Luping Zhou
Luping Zhou
School of Electrical and Computer Engineering, University of Sydney
Medical ImagingComputer VisionMachine Learning
Hongliang Ren
Hongliang Ren
Chinese University of Hong Kong | National University of Singapore | JHU/Harvard(RF) | CUHK(PhD)
Biorobotics & intelligent systemsmedical mechatronicscontinuumsoft flexible robots/sensorsmultisensory perception