Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of deep foundation models in reconstructing dynamic 4D scenes by proposing a training-free dynamic 4D reconstruction framework. The method leverages vision-language models to generate semantic prompts, which are combined with SAM 3 to precisely segment dynamic objects, thereby decoupling the scene into a static background and dynamic point clouds. Furthermore, it extends Depth Anything 3 to dynamic scenarios without fine-tuning, mining implicit motion cues to achieve instance-level separation of static and dynamic components. Experimental results demonstrate that the proposed framework improves dynamic segmentation accuracy by 5.5%, accelerates pose estimation by 13×, reduces memory consumption by 4–8×, and enables denser temporal sampling.
📝 Abstract
Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encodes cross-frame matching, a property absent in depth-only models like DA3. We present Dyna3, a training-free framework that extends DA3 for 4D dynamic scene reconstruction without any fine-tuning. Our key insight is that DA3's cross-view features, though trained only for depth consistency, implicitly encode motion-discriminative signals when combined with best-match feature search across frames. Its static surfaces find consistent matches globally, while dynamic objects cannot. We further adopt vision-language models (VLM) to automatically generate scene-specific semantic prompts for SAM 3, enabling precise instance-level segmentation that distinguishes which objects move from what objects exist. For reconstruction, we decouple the scene into a cross-frame aligned static background and per-frame dynamic point clouds. Experiments on four datasets demonstrate that Dyna3 surpasses correspondence-trained methods with +5.5pp J-Mean over state-of-the-art VGGT4D on dynamic object segmentation, while achieving up to 13x faster pose estimation and 3x faster 4D reconstruction with 4 to 8x lower memory. Dyna3 could therefore enable much denser temporal sampling that prior methods cannot support.
Problem

Research questions and friction points this paper is trying to address.

4D reconstruction
depth foundation models
dynamic scenes
training-free
motion segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free 4D Reconstruction
Depth Foundation Models
Vision-Language Models
Dynamic-Static Decoupling
Instance Segmentation
💼 Related Jobs
No related jobs found.
X
Xinhao Xiang
IFM Lab, University of California, Davis, CA, USA
Weiyang Li
Weiyang Li
Chongqing University
Physical layer securityUAVUnmanned System
Z
Zhijie Zheng
IFM Lab, University of California, Davis, CA, USA
A
Abhijeet Rastogi
IFM Lab, University of California, Davis, CA, USA
Jiawei Zhang
Jiawei Zhang
University of California, Davis (UC Davis)
Machine LearningFoundation Model