Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the object structural distortion and background shift caused by implicit geometry modeling in reference-based image synthesis by proposing the D2R framework. This method achieves geometry-guided RGB generation by predicting synthetic depth for unobserved scenes, with core innovations including supervising a frozen depth estimator via encoder feature correction, designing an independent renderer, and introducing a reference-conditioned rectification mechanism. Furthermore, the AnyInsertion++ benchmark is constructed to evaluate cross-category generalization. Experiments demonstrate that D2R significantly outperforms existing baselines under both paired and cross-category settings, reducing AbsRel by 43.7%, improving PSNR by 2.4 dB, and decreasing CLIP distance by 55%.
๐Ÿ“ Abstract
Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene's geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: https://shjo-april.github.io/Depth2RGB/
Problem

Research questions and friction points this paper is trying to address.

object compositing
depth estimation
scene geometry
reference-based synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-to-RGB
Frozen Depth Estimator
Geometry-Guided Compositing
Encoder-Feature Supervision
AnyInsertion++
๐Ÿ”Ž Similar Papers
No similar papers found.