Score
Designs and implements algorithms and representations that use explicit spatial geometry (e.g., 3D pose, scene layout, reconstructed surfaces) to track and localize object instances over time and across views. Builds methods that enforce geometric consistency across frames—predicting instance-consistent masks and poses, maintaining identity under occlusion, and using geometry as an implicit memory or constraint for localization in reconstructed space.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
This work addresses the challenge that existing video segmentation methods, which rely on explicit appearance-based memory, struggle to maintain spatial consistency across viewpoints and over time under large viewpoint changes and prolonged occlusions. To overcome this limitation, the authors propose a unified framework that, for the first time, leverages spatially aligned geometric representations from a feed-forward 3D reconstruction model as implicit memory for promptable 3D instance tracking. By introducing a cross-modal spatial encoder that fuses visual and textual prompts into a shared geometric space, the framework enables end-to-end spatial reconstruction and consistent mask prediction. Experiments on the newly introduced large-scale dataset InsTrack demonstrate state-of-the-art performance in cross-view consistency, promptable tracking, video object segmentation, and 3D reconstruction.
This paper addresses text-driven scene-consistent image generation: synthesizing images that simultaneously preserve geometric/appearance fidelity to a reference scene graph and accurately realize textual descriptions of target entities and their spatial relationships. To overcome the trade-off between these competing objectives in existing methods, we propose a geometry-guided diffusion framework comprising: (1) a multi-view geometric modeling pipeline for constructing scene-consistent training data; (2) self-supervised spatial regularization incorporating cross-view geometric constraints; and (3) a scene-text joint attention mechanism. Notably, this work is the first to explicitly integrate geometric priors into attention optimization for text-to-image generation. On our newly established benchmark, our method achieves +12.6% CLIP-Scene score, +9.3% TIFA score, and 78.4% human preference rate, demonstrating strong capability in generating complex geometric compositions.
Existing methods for scene-consistent video generation often suffer from error accumulation due to reliance on external memory, non-differentiable operations, or decoupled multi-model architectures, leading to degraded consistency. This work proposes a “Geometry-as-Context” framework that integrates geometric information as dynamic context within an autoregressive video generation process, enabling end-to-end training through alternating estimation of current-view geometry and rendering of novel views. Key innovations include a camera-gated attention mechanism to enhance pose awareness, interleaved training of geometry and RGB sequences, and stochastic dropping of geometric context during training to support pure RGB inference. Experiments demonstrate that the proposed method significantly outperforms existing approaches under both unidirectional and round-trip camera trajectories, achieving substantial improvements in scene consistency and camera control accuracy.
Existing video world models struggle to effectively retain information about scenes that have moved out of view during long-horizon rollouts, as both explicit and implicit memory approaches are often hindered by retrieval errors, redundant storage, or insufficient geometric modeling. This work proposes GIM-World, a novel framework that uniquely integrates explicit cross-view geometric constraints into an implicit memory architecture. It employs a lightweight Transformer encoder to compress historical observations into a fixed-size set of memory tokens and introduces a geometry head that queries camera poses to distill 3D scene structure from a frozen foundation model. An information-guided pruning strategy ensures memory compactness while preserving long-term geometric and visual consistency. Evaluated on the MIND dataset, GIM-World significantly outperforms current state-of-the-art explicit and implicit memory baselines.
This study addresses the challenges of inefficient historical visual memory retrieval and accumulated 3D reconstruction errors in long-horizon camera-controlled video generation by proposing the GEAR framework. This method leverages geometric information as explicit addresses to route attention toward frame-level visual memories, thereby preventing error accumulation caused by global fusion. Furthermore, it introduces a novel geometry-correspondence attention mechanism with an implicit octree structure, combined with token-level matching and visibility evidence accumulation, to achieve precise historical feature injection while effectively eliminating occlusion interference. Built upon diffusion models, GEAR enables the generation of videos up to one minute in length, achieving state-of-the-art performance in visual quality, camera control accuracy, and revisit consistency.
Existing image manipulation localization methods rely primarily on 2D cues and suffer significant performance degradation when tampered regions are seamlessly blended with the background. This work proposes a geometry-aware localization framework that, for the first time, incorporates 3D geometric cues—such as depth and surface normals derived from monocular reconstruction—into the task. By assessing the reliability of these 3D cues, the method employs a multi-scale fusion mechanism to selectively integrate them with RGB features. The proposed approach achieves notably improved localization accuracy with minimal additional computational overhead, demonstrating that trustworthy 3D geometric information effectively complements conventional 2D forensic cues.
This work addresses the challenge of maintaining semantic, structural, and geometric consistency under extreme viewpoint variations in cross-view localization. The authors propose CROSS, a novel framework that reframes cross-view localization as a joint learning task beyond pose estimation, integrating 3D grounding alignment, structure-aware matching, and relative hypothesis ranking to cohesively model semantic, structural, and geometric consistency. Notably, CROSS introduces structure learning as an intrinsic constraint to preserve semantic integrity—avoiding the pitfalls of point-wise matching—and leverages the powerful 2D representation capabilities of vision foundation models to enhance geometric reasoning. Evaluated on KITTI and VIGOR benchmarks, the method achieves state-of-the-art performance, significantly improving semantic stability, structural reliability, and geometric transferability across drastically different viewpoints.
This study addresses the challenges of localizing geometric inconsistencies in multi-view images and the limited generalizability of existing methods. To this end, we introduce DeformView, the first wide-baseline multi-view inconsistency benchmark dataset, and propose DEFECt3R, a lightweight classifier designed for pixel-level inconsistency localization. By leveraging cross-view feature matching, DEFECt3R effectively captures inter-view geometric relationships. Furthermore, a hard negative supervision strategy is incorporated to explicitly train the model, thereby suppressing false alarms. Experimental results demonstrate that DEFECt3R significantly improves localization accuracy for geometric inconsistencies while substantially reducing the false positive rate, establishing a new benchmark for multimedia forensics.
This work addresses the challenge of simultaneously preserving geometric consistency and motion fidelity in novel view synthesis from monocular video. To this end, the authors propose a motion-aware training framework that leverages multi-view point tracking to provide joint geometric and motion supervision, explicitly reinforcing correspondences across views and over time. A key insight is that certain attention layers in diffusion models inherently encode strong correspondence cues; building on this observation, the method introduces an auxiliary multi-view tracking head trained jointly with the main generation pipeline to align spatio-temporal features across views under camera conditioning. Experiments demonstrate that the proposed approach significantly improves geometric consistency on multiple benchmarks while maintaining state-of-the-art accuracy in camera trajectory estimation.