arXiv · 2609.35416
When Should the Count Change? Learning State Maintenance for Causal Video Counting
Abstract
Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi, Jiahao Xie, Si Liu. 2026-09-28. When Should the Count Change? Learning State Maintenance for Causal Video Counting. https://arxiv.org/abs/2609.35416
Cite the original work for its findings. Save a collection to share your selection of sources.