UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D generation models often suffer from geometric distortions, overly smoothed textures, and color inconsistencies when processing unconstrained multi-view images. This work proposes a training-free, plug-and-play framework that introduces, for the first time, a voxel-level image reliability routing mechanism. By leveraging Simultaneous Focus Cross-Attention (SFC-Attn) and a Voxel Reference Score (VRS), the method dynamically guides each voxel during denoising to attend to the most trustworthy input view. Without relying on external matching or segmentation models, it effectively enforces multi-view consistency, significantly improving 3D generation quality from multiple images. The approach substantially mitigates geometric artifacts and color confusion, thereby unlocking the full potential of single-image 3D foundation models in multi-view scenarios.
📝 Abstract
Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step. Revisiting recent single-image 3D foundation models, we show that explicitly routing each voxel to its most informative image is sufficient to unlock strong performance on inconsistent multi-image inputs. Based on this observation, we propose UMI3D, a training-free and plug-and-play framework that restructures cross-attention for unconstrained multi-image 3D generation. Its core, Simultaneous Focus Cross-Attention (SFC-Attn), activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. To enable this routing, we derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel--image affinity that requires no external matching, segmentation, or correspondence models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Project Page: UMI3D-Project.github.io.
Problem

Research questions and friction points this paper is trying to address.

3D generation
multi-image inputs
cross-attention
unconstrained images
voxel-image affinity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simultaneous Focus Cross-Attention
Voxel Reference Score
multi-image 3D generation
training-free framework
cross-attention routing
🔎 Similar Papers
No similar papers found.