R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of insufficient viewpoint coverage in sparse real-world image collections, which degrades the quality of real-to-simulation (R2S) scene reconstruction and hinders policy transfer. To overcome this limitation, the authors propose a dual-agent collaborative framework that couples a behavior-aware robotic agent with a geometry-aware agent to generate behaviorally feasible missing views as pseudo-observations under a fixed data acquisition budget, integrating them into visual assets while iteratively refining the scene’s collision surfaces for improved simulation fidelity. By jointly leveraging behaviorally plausible queries and geometry constraints anchored to actual captures, the method achieves accurate and efficient reconstruction while preserving robotic dynamics consistency. Experiments on three Replica scenes demonstrate a PSNR of 19.062 dB using only six input views—significantly outperforming baselines—and boost the success rate of a real Unitree G1 robot in a sitting task to 82.5% ± 6.8%, vastly exceeding GaussGym’s 10.0% ± 10.5%.
📝 Abstract
Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-environment real-image capture-count efficiency, and sparse human capture can leave behavior-scoped robot views under-supported. Camera-controlled synthesis can fill missing views, but its use in R2S requires behavior-admissible queries and capture-anchored structural conditioning. We present R2S-EGO, which couples a simulator-derived robot proxy that represents the behavior-scoped executable query domain with a capture-anchored geometry proxy that supplies scene-specific structural conditions. Within this domain, fixed- budget selection targets current support deficits for which geometry support is available. The generated observations are assimilated as pseudo-observations to refine the visual asset, while real captures remain anchors. The fused geometry proxy also supplies the scene collision surface, which is refreshed between rounds. Together, these updates refine the existing simulation scene while its robot dynamics and control stack stay fixed. Across 48 frozen Unitree G1 ego views in three Replica scenes, six-view R2S-EGO reaches 19.062 dB PSNR, compared with 14.226 dB for the strongest reported R2S baseline. Across five paired policy-training seeds, R2S-EGO achieves 82.5% +/- 6.8% real-G1 sitting success, compared with 10.0% +/- 10.5% for GaussGym.
Problem

Research questions and friction points this paper is trying to address.

real-to-sim
sparse capture
ego-centric views
scene representation
behavior-scoped views
Innovation

Methods, ideas, or system contributions that make the work stand out.

real-to-sim
dual-proxy refinement
sparse capture
ego-centric synthesis
geometry-conditioned rendering