RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of paired data and geometric evaluation benchmarks for video rendering-to-real translation by proposing the first end-to-end evaluation framework based on reconstructed digital twins. The framework establishes a benchmark comprising twelve scenes, providing editable 3D replicas, proxy renderings, and real target videos. By integrating neural reconstruction with visual geometry techniques, it automates data processing through static-dynamic decoupling and camera registration. Furthermore, combining multi-view verification with scene decomposition enables automated assessment of appearance fidelity and geometric-dynamic consistency. The project releases 1,496 frames of paired data with an 85.7% retention rate, achieving a depth si-RMSE of 0.2170 and an instance mIoU of 0.3673. These contributions significantly reduce manual modeling costs while effectively identifying causes of generation failures.
📝 Abstract
Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video depicting the same scene evolution, paired with an editable, geometrically registered 3D replica. Such data has traditionally required substantial manual modeling, calibration, and animation effort. We introduce RenderBench, a benchmark of 12 reconstructed real-world scenes spanning large-scale indoor environments and egocentric viewpoints, with both static and dynamic settings. Our construction pipeline combines visual geometry, neural reconstruction, and assisted 3D authoring. Each scene is decomposed into static objects and dynamic actors, registered to the capture cameras, and accepted only after multi-view geometric and temporal validation. Each evaluation unit contains appearance reference images, a held-out real target video, an editable digital twin, a matched proxy render, and renderer-native scene annotations. We evaluate transfer models against paired real target videos, retain PAI-Bench-C-compatible structural projections, and use scene annotations to localize failures by object, visibility, articulation, and motion. The first release retains 12 of 14 registered samples (85.7%), comprising 1,496 paired real-proxy frames. All released scenes pass file-integrity and environment-edit audits, while proxy diagnostics yield a depth si-RMSE of 0.2170 and instance mIoU of 0.3673. RenderBench provides paired real observations and editable scene state for assessing both appearance fidelity and preservation of geometry and dynamics.
Problem

Research questions and friction points this paper is trying to address.

render-to-real video transfer
benchmark
digital twins
video generation evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Render-to-Real Video Transfer
Digital Twins
Neural Reconstruction
Benchmark
Scene Decomposition
🔎 Similar Papers
No similar papers found.