🤖 AI Summary
This study addresses the challenge of distinguishing newly observed objects, novel entities, and completed events in continuous video counting by proposing StaMina, a framework for causal video counting. The method introduces a state-conditioned update mechanism that supports object recognition through recurrent visual context modeling. It further employs differentiable recursion to learn event transitions along constrained paths, thereby maintaining visibility states, persistent identities, and records of completed events. Additionally, a multi-source data pipeline is constructed to generate counting trajectories. Evaluated on the SVCBench benchmark, an 8B-parameter model achieves a Gaussian accuracy of 38.2 in streaming scenarios, outperforming Counting-SFT by 10.9 points and significantly enhancing online video understanding capabilities.
📝 Abstract
Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/