🤖 AI Summary
This study addresses the computational redundancy in existing video models that rely on uniform grids, which struggle to accommodate the non-uniform distribution of spatiotemporal visual complexity. To this end, we propose OctVideo, a pioneering octree-based adaptive sparse video representation. By recursively partitioning spatiotemporal volumes, it enables hierarchical modeling with coarse granularity for smooth regions and fine granularity for detailed areas, combined with a Conv1D VAE and lightweight residual modules for efficient reconstruction. This approach significantly reduces spatiotemporal redundancy while supporting zero-shot generalization to high-resolution scenarios. On Kinetics-400, OctVideo achieves a PSNR of 36.12 dB with only 38.2M parameters, attaining state-of-the-art encoding-decoding speed and computational efficiency alongside competitive video recognition performance.
📝 Abstract
Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.