🤖 AI Summary
This work addresses the limitations of existing deep detector–based traffic perception methods, which rely on pretrained models and struggle in scenarios lacking annotations or involving open-set vehicle categories. The authors propose a lightweight, purely geometric, and category-agnostic traffic perception framework that extracts moving objects via background subtraction and thresholding, then counts vehicles using geometric rules applied along virtual induction lines—eliminating the need for object detection models, training data, or trajectory tracking. A self-calibrating pre-calibration mechanism is further introduced to recover lane geometry and count lane occupancy boundaries through patch statistics, while also estimating average per-vehicle speed at no additional computational cost. Evaluated on four video sequences, the method achieves counting accuracies of 83.3%–100%, with 91% accuracy in real-world deployment, substantially outperforming a patch-tracking baseline (37.5%) under comparable computational constraints.
📝 Abstract
Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing in practice: a detector already trained on the class one wants to count. We revisit a purely geometric traffic-sensing pipeline for Single Board Computers in which detection is class-agnostic: moving objects come from background subtraction and thresholding, and counting is decided by a geometric rule on an imaginary line across the road, a software inductive loop detector. With no object model, training set or per-object trajectory, it runs faster than real time on Raspberry Pi class hardware. Two counting rules are described: a constant average speed rule, whose expected accuracy is derived analytically as about 86% under a Gaussian speed distribution, and a self-calibrating pre-calibration rule that recovers the lane geometry from blob statistics and counts edges of lane occupancy, additionally yielding per-vehicle average speed at no extra cost. Over four videos the latter counts with 83.3%-100% accuracy; in a field deployment it reaches 91% against 37.5% for a blob-tracking baseline under the same compute budget. We report the observations of that period in detail: the resolution floor below which accuracy collapses, the frame rate floor at which vehicles alias past the counting line, the gap between short curated clips and long uncontrolled footage, and the trade-off between Python (easier to tune, 100% CPU) and C++ (40% CPU, thermally viable). These are properties of the sampling geometry, not of the hardware of the time, and still constrain edge deployments. We close by arguing where motion-based, class-agnostic detection remains the right tool: open-set classes with no annotated data, tight power budgets, privacy-constrained installations, and the cold start of mining training crops to bootstrap a learned detector.