A Compact Explicit 4D Representation for Dynamic Scenes

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of compactly representing dynamic scenes while jointly modeling temporal geometry and appearance rendering. We propose Sparc4D, a feedforward autoencoder that encodes monocular videos into sparse 4D states. By sharing static features and compressing time-varying ones, it decodes 2D Gaussian surfels for scene representation. Leveraging spatially anchored temporal slots, Sparc4D achieves a fourfold feature compression ratio while preserving fine texture details through source pixel reprojection, thereby enabling zero-shot cross-dataset transfer without fine-tuning. Experimental results demonstrate that the model requires only 0.95 MB of storage on average for 32-frame sequences while attaining a PSNR of 21.70 dB. These findings indicate that Sparc4D outperforms MoVieS and maintains lossless reconstruction quality even after aggressive compression.
📝 Abstract
A compact dynamic-scene representation must retain both the surfaces seen over time and the appearance needed to render them from new viewpoints. We present Sparc4D, a feed-forward autoencoder that encodes a monocular video with known cameras into a sparse 4D scene state. Static features are shared across the clip, while spatially anchored temporal slots compress time-varying features. A sparse decoder produces 2D Gaussian surfels, while stored source pixels preserve fine texture through geometric re-projection. The state includes one full source frame and dynamic-region pixels sampled every fourth frame, alongside learned features and sparse occupancy. For a 32-frame MultiCamVideo clip, it averages 0.95M 32-bit-equivalent values on random windows and 0.92M on the first-32 protocol. On first-32, Sparc4D reaches 21.70\,dB, compared with 20.40\,dB for MoVieS. On randomly placed windows, their PSNR scores are comparable. With stored texture disabled, temporal slots compress the time-varying feature state by a median $4.0\times$ and reduce the mean state from 1.04M to 0.42M values, with essentially unchanged target-view reconstruction quality. Without fine-tuning on real data, Sparc4D transfers to DyCheck and Neu3D, where stored texture improves LPIPS while slightly reducing PSNR.
Problem

Research questions and friction points this paper is trying to address.

dynamic scenes
4D representation
compact representation
novel view synthesis
monocular video
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Representation
Feed-forward Autoencoder
2D Gaussian Surfels
Temporal Compression
Dynamic Scene Reconstruction