π€ AI Summary
This work addresses the joint furniture removal task for single-view panoramic images and textured meshes. We propose a cross-modal guidance method leveraging a Simplified Furniture-Free Mesh (SDM) as a geometric prior. Specifically, depth and normal maps are rendered from the SDM and processed with Canny edge detection; these geometric cues are then integrated into multi-view panoramic image inpainting via ControlNet. Concurrently, mesh topology is optimized and planar structures are extended to ensure geometric consistency. Our key contribution is the first explicit integration of the SDM as a βgeometric X-rayβ into the image generation pipeline, effectively mitigating structural incompleteness caused by single-view occlusions. Experiments demonstrate that our method produces high-resolution, geometrically consistent furniture-free panoramas and textured meshes, significantly outperforming NeRF-based approaches (which suffer from blurriness and low resolution) and RGB-D inpainting methods (prone to hallucination) in both detail fidelity and structural plausibility.
π Abstract
We present a pipeline for generating defurnished replicas of indoor spaces represented as textured meshes and corresponding multi-view panoramic images. To achieve this, we first segment and remove furniture from the mesh representation, extend planes, and fill holes, obtaining a simplified defurnished mesh (SDM). This SDM acts as an ``X-ray'' of the scene's underlying structure, guiding the defurnishing process. We extract Canny edges from depth and normal images rendered from the SDM. We then use these as a guide to remove the furniture from panorama images via ControlNet inpainting. This control signal ensures the availability of global geometric information that may be hidden from a particular panoramic view by the furniture being removed. The inpainted panoramas are used to texture the mesh. We show that our approach produces higher quality assets than methods that rely on neural radiance fields, which tend to produce blurry low-resolution images, or RGB-D inpainting, which is highly susceptible to hallucinations.