🤖 AI Summary
This study addresses the challenge in dynamic 3D Gaussian Splatting where global video motion causes misalignment between the canonical space and individual frames, thereby imposing excessive burden on deformation models. To mitigate this, we propose constructing an affine-aligned atlas space that explicitly absorbs global motion via per-frame affine transformations prior to canonical Gaussian construction, effectively decoupling global motion from local deformations. Notably, this strategy modifies only the canonical construction phase, enabling seamless integration into existing frameworks with negligible parameter overhead. Our approach substantially improves reconstruction quality for sequences involving significant camera motion, offering an efficient and generalizable optimization paradigm for dynamic scene modeling.
📝 Abstract
Gaussian splatting has recently emerged as an efficient representation for images and videos due to its explicit structure and fast rendering capability. Existing Gaussian-based video representations often decompose a video into canonical Gaussians and temporal deformation. However, when a video contains large global motion such as camera movement, the canonical representation may become misaligned with individual frames, increasing the burden on the temporal deformation model. In this paper, we propose an affine-atlas canonical Gaussian representation, which constructs canonical Gaussians in a larger affine-aligned atlas space. Frame-wise affine transforms absorb global motion before canonical Gaussian construction, reducing the gap between the canonical representation and target frames. Since the proposed method only modifies the canonical construction stage, it can be integrated into existing canonical-Gaussian-based methods with negligible additional parameter cost. Experiments show that our method improves reconstruction quality especially for sequences with large camera motion.