Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of aggregate metrics to reveal failure modes and deployment suitability in pedestrian tracking by proposing a tracking-only evaluation protocol tailored for edge deployment. By sharing detections to isolate variables, we construct a reproducible, fine-grained evaluation framework that directly quantifies critical capabilities such as initialization, missing-observation continuation, and identity recovery, thereby overcoming the limitations of HOTA. Benchmarking on the JackRabbot dataset using the PedRefTrack tracker and a Jetson Orin platform reveals that continuing through missing observations is a core weakness, and multiple trackers fail to maintain 10Hz real-time operation in crowded scenes. This work precisely identifies robustness and real-time performance bottlenecks, delineating the practical deployment boundaries for pedestrian tracking systems.
📝 Abstract
Mobile robots operating among pedestrians need trajectories that become available quickly, remain spatially credible through missed observations, preserve identity, and fit within an embedded computing budget. Aggregate tracking scores provide limited insight into when and how trajectories fail, while varying detector inputs can confound tracker and detector quality. We present a deployment-focused, tracker-only evaluation protocol that uses shared detections to isolate tracker behavior and directly evaluates initialization, detector-gap continuation, identity recovery, close-neighbor association, and load-dependent tracker-step runtime, while Higher Order Tracking Accuracy (HOTA) is retained as a complementary aggregate measure. We apply the protocol to the JackRabbot Dataset and Benchmark (JRDB) using six open-source trackers and our lightweight Pedestrian Reference Tracker (PedRefTrack), together with a GT-assisted variant that estimates the remaining tracker-side gap under idealized association and motion. Under fixed detections, the non-GT trackers span only 24.26%-29.67% HOTA yet exhibit markedly different capability profiles. After 1.0 s without detector support, no tracker without GT assistance maintains spatially correct, same-identity output in more than half of eligible cases, making missing-observation continuation the dominant limitation among the tested properties. Close-neighbor failures are smaller and increase mainly at the shortest separations. Tracker-step runtime on an NVIDIA Jetson Orin is heavy-tailed and load-sensitive, causing several trackers to fall below the 10 Hz real-time target in crowded frames. The protocol provides a reproducible way to characterize tracker behavior and deployment suitability in pedestrian-centric environments. Code and evaluation scripts are released at https://github.com/SCAI-Lab/tracker_eval.
Problem

Research questions and friction points this paper is trying to address.

pedestrian tracking
tracking evaluation
deployment suitability
embedded computing
identity preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deployment-focused evaluation
Tracker-only protocol
Pedestrian tracking
Embedded computing
HOTA
🔎 Similar Papers
No similar papers found.
D
Dominik Wojcikiewicz
Spinal Cord Injury Artificial Intelligence Lab (SCAI), D-HEST, ETH Zurich, Switzerland; and Digital Health Care and Rehabilitation, Swiss Paraplegic Research (SPF), Switzerland
Diego Paez-Granados
Diego Paez-Granados
Lecturer and Head of SCAI Lab, ETH Zurich
Machine LearningAssistive RoboticsRehabilitationHuman-Robot InteractionShared Control