CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

๐Ÿ“… 2026-08-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of modeling complex interactions between asymmetric surgical roles and uncalibrated viewpoints in dual-endoscope proceduresโ€”a setting where existing visual world models fall short. To this end, we propose CrossScope, a dual-stream surgical world model that introduces, for the first time, a role-asymmetric future prediction paradigm for dual-endoscope settings. CrossScope employs a geometry-guided residual interaction mechanism to enable goal-directed cross-view evidence routing: motion cues from the primary endoscope guide prediction in the secondary view, while pose-aligned appearance features from the secondary endoscope reciprocally aid the primary viewโ€™s prediction. This design preserves view-specific representations while fostering effective cross-view collaboration. Experiments on our newly collected synchronized dual-endoscope ERCP dataset demonstrate that CrossScope significantly outperforms current baselines in visual fidelity, structural consistency, target localization accuracy, and motion coherence.
๐Ÿ“ Abstract
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.
Problem

Research questions and friction points this paper is trying to address.

dual-scope
role-asymmetric
surgical video prediction
multi-observer
future dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

role-asymmetric modeling
dual-scope prediction
geometry-guided interaction
surgical world model
cross-view evidence routing
๐Ÿ”Ž Similar Papers
No similar papers found.