Octree-based Video Representation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy in existing video models that rely on uniform grids, which struggle to accommodate the non-uniform distribution of spatiotemporal visual complexity. To this end, we propose OctVideo, a pioneering octree-based adaptive sparse video representation. By recursively partitioning spatiotemporal volumes, it enables hierarchical modeling with coarse granularity for smooth regions and fine granularity for detailed areas, combined with a Conv1D VAE and lightweight residual modules for efficient reconstruction. This approach significantly reduces spatiotemporal redundancy while supporting zero-shot generalization to high-resolution scenarios. On Kinetics-400, OctVideo achieves a PSNR of 36.12 dB with only 38.2M parameters, attaining state-of-the-art encoding-decoding speed and computational efficiency alongside competitive video recognition performance.
📝 Abstract
Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.
Problem

Research questions and friction points this paper is trying to address.

Video Representation
Computational Efficiency
Spatio-temporal Redundancy
Octree
Video Reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Octree-based Video Representation
Spatio-temporal Partitioning
Conv1D VAE
Sparse Structure Processing
Efficient Video Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rungui Zhou
Peking University
C
Chuanzhi Zhou
Peking University
Y
Yuk-Kit Hou
Peking University
Peng-Shuai Wang
Peng-Shuai Wang
Assistant Professor, Peking University
Geometry processing3D Deep LearningComputer graphics