🤖 AI Summary
This study addresses the challenges of component-to-surface assignment ambiguity and over-prediction in single-image multi-surface 3D structure recovery. To this end, it proposes a joint estimation framework based on image-conditioned multi-Bernoulli deep sets. The core innovation lies in the design of an auxiliary-free ExactMB objective function, which enables the joint inference of surface existence and metric depth by marginalizing over one-to-one assignment relationships. Experimental results demonstrate that the proposed method significantly reduces over-prediction by 88.2% and 80.5% on the LD-Real and MD-3K benchmarks, respectively, while maintaining excellent ordinal accuracy. The source code has been made publicly available.
📝 Abstract
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-free Exact Multi-Bernoulli objective (ExactMB) learns depth and presence by marginalizing one-to-one assignments to complete, distinct targets. Our analysis shows that matching expected count can leave component-surface assignment unresolved. We extend real and synthetic layered-depth benchmarks to evaluate depth accuracy, recovered support, and overprediction. Compared to depth stacking, ExactMB reduces overprediction by a relative 88.2% on LD-Real and 80.5% on MD-3K while retaining most ordinal accuracy, with comparable conditional metric-depth error on LD-Syn. Further ablation studies show that ordered assignment improves depth-accurate recall and precision over marginalization, whereas the count-regularized configuration achieves higher deeper-rank precision than ordered assignment at lower recall. Our code will be publicly released.