Score
Designs and implements methods that separate scene content into foreground objects and background layers, producing per-point, per-patch, or per-pixel foreground/background labels, masks, or decoupled geometry for independent editing or recomposition. This competence covers conditional 3D point or patch classification, foreground-background segmentation and scene factorization to remove duplicate object geometry from backgrounds and enable stable object-background separation.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
To address the inefficiency and inaccuracy of Gaussian splatting in reconstructing individual objects, this paper proposes an object-centric 2D Gaussian splatting paradigm. Methodologically: (i) object masks guide targeted reconstruction and enable automatic background removal; (ii) an occlusion-aware Gaussian pruning strategy dynamically eliminates occluded and redundant Gaussians; (iii) the 2D Gaussian rasterization pipeline is optimized, and a lightweight mesh generation mechanism is introduced. The key contribution is the first shift of Gaussian representation from scene-level to object-level modeling—achieving comparable rendering quality while reducing model size to 4% of the baseline and accelerating training by 71%. Moreover, the framework supports plug-and-play appearance editing and physics-based simulation, significantly enhancing efficiency, compactness, and controllability.
Existing 3D inpainting and object removal methods suffer from geometric distortions and texture inconsistencies under unconstrained viewpoints (i.e., arbitrary camera poses and trajectories). This paper introduces the first geometry-guided, multi-view test-time adaptive optimization framework. Our key contributions are: (1) object-mask-based fine-grained inpainting region detection to enhance robustness under complex viewpoints; (2) a geometry-prior-integrated 3D reconstruction pipeline coupled with a multi-view consistency refinement network; and (3) transfer adaptation of pre-trained image inpainting models, followed by test-time self-supervised fine-tuning. Evaluated on a newly constructed diverse unconstrained benchmark, our method significantly outperforms state-of-the-art approaches, achieving geometrically accurate, photorealistic, and cross-view consistent 3D inpainting—both in forward-facing and arbitrary-view settings.
本文针对任意语义概念的精确分割问题,提出了一种无需训练的上下文分割框架FoRIS,通过逐步细化前景响应达到精准分割。
To address image ambiguity in few-shot image classification caused by multi-object coexistence or complex backgrounds, this paper proposes a localization-aware preprocessing method that integrates the Segment Anything Model (SAM) with object-centric cropping—requiring no fine-grained pixel-level annotations—to automatically extract foreground regions and emphasize discriminative local features. The method significantly reduces reliance on manual annotation while enhancing feature discriminability. Evaluated on standard few-shot benchmarks—including Mini-ImageNet, Tiered-ImageNet, and CUB—it achieves state-of-the-art performance, yielding average accuracy improvements of 2.3–4.7 percentage points. Comprehensive experiments demonstrate that the proposed approach exhibits strong generalization capability and computational efficiency, offering a novel and robust paradigm for representation learning in few-shot settings.
本文提出PredErase,一种无需训练的方法,通过预测性潜在引导解决物体及其影响移除问题,改进了现有冻结模型在物体移除时的不足。
Existing approaches struggle to jointly handle scene text editing tasks—deletion, generation, and replacement—within a unified framework that simultaneously ensures precise textual appearance control and background integrity. To address this, this work proposes a unified model that decomposes complex text editing into two atomic operations: rendering and erasure. It introduces Overlay-Reference Positional Encoding (ORPE) to achieve pixel-level layout fidelity and exemplar-driven style control, complemented by a Region-Adaptive Suppression (RAS) strategy to ensure clean text removal. The study also establishes TextWand-Bench, the first comprehensive benchmark for general scene text editing. Experimental results demonstrate that the proposed method significantly outperforms both open-source and closed-source models across all three editing tasks in terms of text accuracy, layout-style consistency, and overall image quality.
Existing 2D instance segmentation models often produce fragmented and inconsistent masks across multiple views, hindering reliable 3D scene understanding. This work proposes a multi-cue guided approach for generating cross-view consistent 2D instance masks by integrating semantic, geometric, and structural information, along with an identity matching mechanism to align instances across viewpoints. These consistent masks are then leveraged to guide the optimization of a 3D Gaussian Splatting feature field, marking the first effective incorporation of instance-level consistency into the Gaussian Splatting framework. Experiments demonstrate that the proposed method significantly improves cross-view mask consistency and 3D instance segmentation stability while preserving high-quality photometric reconstruction, thereby enabling robust downstream editing tasks.
本文提出了一种通过对象级网格喷溅的方法,解决3D场景重建中缺乏对象级结构及纹理污染问题,提高了网格保真度和新颖视角合成质量。
This work addresses the challenge of achieving high-fidelity 3D geometric reconstruction and robust semantic understanding without requiring camera parameters. To this end, it proposes the first unified framework that jointly trains 3D feedforward reconstruction and 3D panoptic segmentation. The approach incorporates geometric priors during feature initialization and enables end-to-end co-optimization through combined geometric and semantic losses. A set-based mask decoder is employed, compatible with both online and full pairwise attention architectures. The method achieves state-of-the-art performance in 3D panoptic segmentation on ScanNet, ScanNet200, and ScanNet++. Ablation studies confirm that joint training mutually enhances both reconstruction accuracy and semantic understanding.