🤖 AI Summary
This study addresses the challenges of geometric fitting and conflicts between semantic and physical constraints in 3D scene object insertion. We propose a VLM-driven, scene-aware object insertion framework that introduces a scene-grounding representation to translate high-level semantics into explicit 3D constraints. The method leverages VLM reasoning to generate structured fitting cues that guide conditional object generation, and integrates generative AI, 3D reconstruction, and collision detection for mesh optimization, supporting rigid, scaled, and elastic geometric fitting modes. Experimental results demonstrate that the proposed approach achieves a spatial relationship success rate of 69.7% and a support success rate of 91.7%, significantly outperforming baseline methods. These findings confirm its effectiveness in addressing the "make-it-fit" task within complex 3D scenes.
📝 Abstract
Inserting objects into existing 3D scenes requires more than selecting a plausible location:
the inserted object must also fit local geometry while preserving semantic intent and physical plausibility.
Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces.
We introduce \textbf{ElasticFit}, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation.
Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting).
These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting.
ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding.
In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8\% to 69.7\% and support success from 48.3\% to 91.7\% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.