Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

📅 2024-12-05
🏛️ arXiv.org
📈 Citations: 4
Influential: 0
📄 PDF

career value

225K/year
🤖 AI Summary
This work addresses the failure of conventional stereo matching and monocular depth estimation under challenging scenarios—including textureless regions, occlusions, specular, and transparent objects—where ground-truth depth supervision is unavailable. We propose the first zero-shot cross-domain stereo matching framework. Methodologically, we design a dual-branch collaborative architecture: a geometric branch enforcing disparity consistency and a vision foundation model (VFM) branch incorporating monocular priors; introduce a geometry-guided cost volume fusion mechanism for cross-modal feature alignment; and construct MonoTrap, an optically plausible synthetic dataset enabling purely synthetic, zero-shot training. Without access to real-world annotations, our method achieves state-of-the-art zero-shot performance across multiple benchmarks, significantly outperforming existing stereo and monocular approaches on challenging cases such as specular and transparent surfaces.

Technology Category

Application Category

📝 Abstract
We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching with learned contextual cues. Following this design, our framework introduces novel cost volume fusion mechanisms that effectively handle critical challenges such as textureless regions, occlusions, and non-Lambertian surfaces. Through our novel optical illusion dataset, MonoTrap, and extensive evaluation across multiple benchmarks, we demonstrate that our synthetic-only trained model achieves state-of-the-art results in zero-shot generalization, significantly outperforming existing solutions while showing remarkable robustness to challenging cases such as mirrors and transparencies.
Problem

Research questions and friction points this paper is trying to address.

Combines geometric constraints with monocular depth priors
Handles textureless regions, occlusions, and non-Lambertian surfaces
Achieves zero-shot generalization with synthetic-only training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines geometric constraints with monocular depth VFMs
Uses dual-branch architecture for stereo and monocular fusion
Introduces novel cost volume fusion for challenging surfaces
🔎 Similar Papers
No similar papers found.