Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

📅 2026-05-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the degradation in tracking accuracy caused by conventional frame-level sampling in static videos, which often discards fine-grained motion information. To overcome this limitation, the authors propose a tile-based polyomino modeling approach that enables efficient trajectory extraction through spatiotemporal joint pruning. The method employs a three-stage pipeline—tile classification, integer linear programming (ILP)-based pruning, and canvas packing—and supports user-specified detectors while adaptively adjusting the sampling strategy under given accuracy constraints. Experiments across seven static video datasets demonstrate that, with trajectory accuracy loss capped at 5%, the system achieves up to 17.4× higher throughput than state-of-the-art methods and 68.8× improvement over the full-frame, frame-by-frame baseline.
📝 Abstract
Track materialization converts raw video into reusable object tracks that downstream queries can run against without rerunning tracking, but extracting those tracks efficiently and with high fidelity remains expensive. Prior systems reduce cost through temporal frame sampling, erasing the inter-frame motion that fine-grained tracking requires. In stationary video, however, large portions of each frame contain no objects of interest, and the remaining regions tolerate different sampling rates. We present Tetris, a track-extraction system that decomposes videos into a tile-based polyomino data model, enabling fine-grained spatiotemporal pruning that reduces detector calls with minimal fidelity loss. Tetris runs three operators upstream of the user-provided detector: a classifier identifies relevant tiles and groups them into polyominoes, an integer linear program (ILP) prunes redundant polyominoes under a user-specified accuracy constraint, and a packer assembles the survivors into canvases that minimize detector calls. Across 7 stationary-video datasets, Tetris stays within a 5% tracking accuracy loss of a full-frame, every-frame reference pipeline, whereas prior systems exceed this bound on 3 of the 7 datasets. At this 5% bound, Tetris achieves up to 17.4x higher throughput than prior systems and up to 68.8x higher than the reference pipeline. The project page is at https://tetris-db.github.io .
Problem

Research questions and friction points this paper is trying to address.

video object tracking
track materialization
spatiotemporal sampling
high-fidelity tracking
efficient video processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

tile-level sampling
polyomino data model
spatiotemporal pruning
integer linear programming
video object tracking
🔎 Similar Papers
No similar papers found.
C
Chanwut Kittivorawong
U. of California, Berkeley
A
Alena Chao
U. of California, Berkeley
C
Charlie Si
U. of California, Berkeley
A
Alvin Cheung
U. of California, Berkeley