Score
Designs and implements algorithms and tools that perform image- or video-level edits conditioned on explicit geometric specifications—such as depth boxes, orientation boxes, object-level constraints, or object-state instructions—so edits place, orient, or modify objects according to those geometry cues. Builds pipelines to specify, propagate, and enforce geometric constraints across viewpoints and time, aligning geometry streams for consistent, appearance-agnostic object-level editing and state changes.
This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.
Existing image editing methods struggle to precisely control large-scale object motion and viewpoint changes. This work proposes a structured editing framework based on 3D bounding boxes, where users define target transformations by specifying input and output 3D boxes, which the system then formulates as a geometric constraint problem. The approach introduces an innovative “box-based thinking” interface—using color-coded faces to explicitly represent orientation—and leverages a depth-aligned planar floor as a global reference frame. Combined with depth-aware rendering, a two-stage training strategy (utilizing synthetic multi-object scenes and real-world Objectron videos), and a structure-conditioned generative model, the method directly processes real images and significantly outperforms existing techniques in large-scale 3D editing tasks while preserving object identity and plausibly reconstructing occluded regions.
Existing methods for image and video object manipulation struggle to simultaneously preserve background content, maintain geometric consistency across viewpoints, and offer fine-grained user control. This work proposes Ctrl&Shift, an end-to-end diffusion framework that decomposes manipulation into two stages—object removal followed by camera-pose-guided reference-based inpainting—enabling geometrically consistent editing within a unified diffusion process without explicit 3D modeling. By integrating explicit camera pose control, reference-guided inpainting, and a multi-task, multi-stage training strategy, the method effectively disentangles background, identity, and pose signals. This design preserves generalization to real-world scenes while supporting precise geometric manipulation. Experiments demonstrate that Ctrl&Shift significantly outperforms existing geometry-based and diffusion-based approaches in terms of generation fidelity, viewpoint consistency, and user controllability.
This work addresses the challenge of sketch-based structural editing of 3D scenes in video, particularly under large camera motions (e.g., significant rotation or scaling), where maintaining cross-view consistency, preserving unedited regions, and ensuring geometric fidelity from 2D sketches to 3D outputs remain difficult. We propose a 3D-aware video editing framework comprising three key components: (1) dense stereo matching to jointly estimate scene point clouds and camera parameters; (2) a point-cloud-guided geometric editing module coupled with a 3D-aware mask propagation strategy, enabling sparse sketch/mask inputs to be mapped onto depth-consistent 3D components; and (3) integration of first-frame image editing with video diffusion models to synthesize temporally coherent 3D videos. Extensive experiments demonstrate that our method significantly outperforms existing approaches in view consistency, geometric accuracy, and detail realism.
3D editing faces core challenges including cross-view inconsistency, geometric distortion, and reliance on manually annotated 3D masks. To address these, we introduce 3DEditVerse—the first large-scale paired benchmark for 3D editing—and propose 3DEditFormer, a conditional Transformer model. Our method employs a dual-guided attention mechanism to decouple edit regions from structural priors and incorporates time-varying adaptive gating to jointly enforce locality and multi-view consistency. Built upon an image-to-3D generation framework, it unifies pose-driven geometric editing with foundation-model-guided appearance editing—eliminating the need for precise 3D masks. Extensive evaluation across multiple benchmarks demonstrates state-of-the-art performance in editing fidelity, structural integrity, inference speed, and cross-view consistency. This work significantly advances scalable and practical 3D editing.
Existing text-driven video editing methods struggle to accurately model spatial relationships and physical constraints in complex scenes, often resulting in ambiguous editing targets and distorted outputs. This work proposes a novel "Plan–Guide–Edit" framework that introduces chain-of-thought reasoning into video editing for the first time. Leveraging a multimodal large language model, the approach performs structured semantic reasoning to generate precise editing instructions annotated with bounding boxes and object attributes, which are then executed by a diffusion model to achieve high-fidelity, spatiotemporally consistent edits. By explicitly bridging semantic intent with spatial execution, the method significantly improves localization accuracy and physical plausibility in multi-object, complex scenarios, surpassing strong baselines with substantially less training data and achieving state-of-the-art performance.
为解决图像编辑中相机参数控制问题,提出CameraEditor框架,将空间问题转化为时序预测任务,并通过动态全景裁剪和插入过渡帧等方法提高几何精度和内容一致性。
This work addresses the challenge of achieving pixel-level precision in sketch-based image editing, which is hindered by the absence of high-quality datasets that jointly encode geometric constraints and semantic instructions. To overcome this limitation, we propose SI-Edit, a novel framework accompanied by SI-Data—the first high-quality dataset specifically designed for instruction-guided sketch editing. SI-Data comprises quadruplets of images, sketches, natural language instructions, and corresponding edited results, automatically generated using multimodal large language models to enable spatial-semantic co-learning. By jointly modeling geometric sketches and semantic instructions, our method achieves high-fidelity local deformations and significantly outperforms existing approaches in both structural preservation and alignment with user intent.
This study addresses the limitations of existing video generation methods, which lack precise control over camera and object motion while exhibiting insufficient geometric consistency. To this end, this work proposes a video world model grounded in an explicit 4D scene representation. The model introduces geometric constraints through rigid 3D mesh reconstruction to suppress artifacts in unobserved regions, and integrates depth-aware video rendering with a motion adapter to achieve precise motion control. Furthermore, an automatic annotation pipeline, the RealCOD-Rigid dataset, and the IG-IoU evaluation metric are constructed to support this research. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in both visual quality and the precision of camera and object motion control.
为解决3D几何编辑的劳动密集问题,提出了一种基于视觉、无需训练的多代理系统ViSculpt,通过模仿人类艺术家的工作流程直接在Blender中编辑现有3D网格。