🤖 AI Summary
Existing controllable surgical video generation methods often suffer from anatomical distortions, instrument appearance drift, and temporal inconsistencies due to the direct fusion of heterogeneous visual conditions. To address these issues, this work proposes a hierarchical surgical anchoring mechanism that integrates anchor-relative multimodal experts—each dedicated to processing edges, depth, and optical flow—and employs a staged fusion strategy to deliver precise control signals to a video diffusion model. Built upon the Wan2.2 backbone and evaluated on the Cholec80-SurgWAM benchmark, the proposed approach significantly outperforms current baselines in generation quality, temporal coherence, and multimodal controllability. It represents the first unified surgical world modeling framework that simultaneously preserves anatomical structure, ensures complementary information integration, and maintains consistent cross-modal interactions.
📝 Abstract
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.