🤖 AI Summary
This work addresses the challenge of selecting target objects in 3D Gaussian splatting scenes under sparse-view settings by proposing a training-free, interactive 3D object selection method. Operating directly on raw Gaussian primitives, the approach generates superpoints via geometric clustering and constructs a graph with continuity-aware edge weights. Integrating sparse user-provided scribbles with visibility-aware 3D lifting, it achieves globally consistent selection through graph-cut energy minimization and supports multi-round, cross-view human-in-the-loop refinement. To our knowledge, this is the first method to enable lightweight 3D selection using only sparse views and minimal user annotations, significantly reducing both interaction and computational overhead while attaining accuracy comparable to state-of-the-art multi-view SAM-based approaches—making it well-suited for real-world 3D editing and asset extraction tasks.
📝 Abstract
Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gaussian representation to embed per-object labels, or build dense multi-view SAM observations, both requiring heavy computation and dense viewpoint coverage that is rarely available in practice. We present GaussianSelector, a training-free framework for interactive 3D object selection from sparse views and sparse scribble guidance. Operating directly on native Gaussian primitives, we coarsen dense Gaussians into geometrically coherent superpoints and construct a continuity-weighted graph using appearance and spatial cues. Sparse user scribbles are lifted into 3D via visibility-aware transmittance coverage, and selection is solved as a global graph-cut energy minimization that propagates sparse evidence to a complete 3D object. This design naturally supports multi-round refinement, where users iteratively correct the selection from additional viewpoints to progressively improve the result. Experiments demonstrate that GaussianSelector achieves competitive selection quality against state-of-the-art multi-view SAM-based methods, while requiring significantly fewer interaction views and substantially lower computational overhead. These properties make it well suited for human-in-the-loop 3D scene editing and 3D asset extraction in real-world deployment scenarios.