π€ AI Summary
This study addresses the challenges of local-view-based active 3D reconstruction and long-horizon keyframe selection under limited computational budgets. We propose a training-free framework that unifies both tasks by leveraging generative model evidence. The core innovation lies in a novel mechanism for uncertainty estimation and information gain derivation based on cross-attention evidence, which integrates evidence-space reasoning, latent token uncertainty propagation, and occlusion-aware aggregation. Experimental results demonstrate that our method reduces the Chamfer distance by up to 12.7% and accelerates inference by 1.5Γ, while achieving comparable reconstruction accuracy using only 14% of the input views.
π Abstract
How can a 3D reconstruction system acquire and retain useful information to understand the geometry of a scene from partial views under a limited computation budget? Existing active view acquisition methods typically estimate uncertainty over observed or instantiated geometry, limiting their ability to reason about unseen structure, while long-horizon reconstruction methods often retain redundant observations. We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection by leveraging evidence provided by a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence associated with 3D latent tokens and derives an Evidential Information Gain to guide both view acquisition and keyframe selection based on the expected reduction in posterior entropy. Matisse supports multi-object scenes through occlusion-aware, object-balanced aggregation and propagates uncertainty through intermediate latents to avoid full reconstruction during planning. Matisse reduces Chamfer distance by 12.7%, 3.8%, and 9.2% on GSO30, YCB-V, and Replica, respectively, relative to the best baseline on each dataset, and achieves a $1.50\times$ end-to-end speedup over the best active reconstruction baseline on GSO30 with the same reconstruction backend. In the GSO30 keyframe selection experiment for long-horizon reconstruction, Matisse achieves comparable Chamfer distance using 14% of the input views compared with Stream3D.