When Should the Count Change? Learning State Maintenance for Causal Video Counting

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of distinguishing newly observed objects, novel entities, and completed events in continuous video counting by proposing StaMina, a framework for causal video counting. The method introduces a state-conditioned update mechanism that supports object recognition through recurrent visual context modeling. It further employs differentiable recursion to learn event transitions along constrained paths, thereby maintaining visibility states, persistent identities, and records of completed events. Additionally, a multi-source data pipeline is constructed to generate counting trajectories. Evaluated on the SVCBench benchmark, an 8B-parameter model achieves a Gaussian accuracy of 38.2 in streaming scenarios, outperforming Counting-SFT by 10.9 points and significantly enhancing online video understanding capabilities.
📝 Abstract
Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/
Problem

Research questions and friction points this paper is trying to address.

Video Counting
State Maintenance
Causal Video
Event Transition
Continuous Counting
Innovation

Methods, ideas, or system contributions that make the work stand out.

State Maintenance
Causal Video Counting
Differentiable Recurrence
Trajectory Supervision
Phase Conditioning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pengyiang Liu
Beihang University
D
Dongyue Lyu
Beihang University
Junbo Niu
Junbo Niu
Peking University
Foundation Model
Z
Zhongyue Shi
Beihang University
J
Jiahao Xie
Beihang University
Si Liu
Si Liu
Fred Hutchinson Cancer Center
GenomicsBiostatisticsAnomaly DetectionOpen Category Detection