Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D scene inpainting methods rely on precise camera poses, limiting their applicability to in-the-wild pose-free multi-view images. To address this limitation, this work proposes FreeInpaint, a framework that generates 3D-consistent inpainted scenes without requiring precomputed camera poses. The proposed method extends a feed-forward 3D foundation model by introducing a learnable masked attention mechanism to preserve spatial anchors. Furthermore, it incorporates a support token refinement strategy to inject diffusion priors, effectively tackling the challenge of high-fidelity completion under occlusion. Experimental results demonstrate that FreeInpaint achieves superior inpainting quality and efficient inference across multiple datasets.
📝 Abstract
3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed. The project page is https://rorisis.github.io/FreeInpaint/.
Problem

Research questions and friction points this paper is trying to address.

3D scene inpainting
pose-free
multi-view images
cross-view correspondence
severe occlusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pose-Free 3D Inpainting
Feed-Forward Framework
Learnable Mask Attention
Support Token Refinement
3D Foundation Model
🔎 Similar Papers
💼 Related Jobs
No related jobs found.