Score
Designs and executes systematic comparative evaluations and benchmarks of tracking algorithms, including building test datasets and simulations, defining performance metrics and noise models, and running controlled experiments to measure accuracy, robustness, and failure modes across methods. Analyzes results to diagnose algorithm weaknesses, produce reproducible benchmarking protocols, and recommend the most appropriate tracking approach for given scenarios.
This paper addresses the challenges of evaluating robustness and ensuring reproducibility in Real-Time Localization and Tracking Systems (RTLS) under radar spoofing attacks. To this end, we propose the first modular and reproducible benchmarking framework. Methodologically, we design a decoupled dual-stream architecture—comprising a clean stream and a spoofing detection stream—to model three canonical radar spoofing types: drift, ghost, and mirror attacks. The framework integrates JPDA and GNN trackers, introduces a realistic offset metric to quantify assignment errors, and enables attack interpretability via trajectory offset visualization, clustering overlay, and spoofing injection timelines. Our contributions include an open-source benchmark framework, standardized evaluation protocols, and automated analytical tools—collectively enhancing transparency, comparability, and community verifiability in anti-spoofing tracking research.
Existing visual SLAM methods exhibit severe generalization deficits across diverse applications (e.g., XR, IoT, autonomous driving, UAVs, human pose tracking) and heterogeneous environments (indoor/outdoor, static/dynamic scenes, varying motion patterns), stemming from deep coupling among algorithm design, environmental characteristics, and platform motion dynamics. Method: We propose the first three-dimensional challenge taxonomy—“algorithm–environment–motion”—and systematically evaluate state-of-the-art methods (ORB-SLAM2/3, VINS-Fusion) on multi-source benchmarks (TUM, EuRoC, ARKitScenes, UAV-Human), quantifying performance via absolute trajectory error (ATE), relative pose error (RPE), and tracking loss. Contribution/Results: No method achieves robust cross-domain or intra-domain heterogeneous generalization. To address this, we introduce a principled co-optimization pathway comprising input representation disentanglement, intermediate information reuse, and output dynamic validation—establishing a reproducible benchmark and foundational design principles for universal visual localization.
Real-time online object tracking in videos constitutes a core task in computer vision, with wide-ranging applications including video surveillance, motion capture, and robotics. Deployed tracking systems usually lack formal safety assurances to convey when tracking is reliable and when it may fail, at best relying on heuristic measures of model confidence to raise alerts. To obtain such assurances we propose interpreting object tracking as a sequential hypothesis test, wherein evidence for or against tracking failures is gradually accumulated over time. Leveraging recent advancements in the field, our sequential test (formalized as an e-process) quickly identifies when tracking failures set in whilst provably containing false alerts at a desired rate, and thus limiting potentially costly re-calibration or intervention steps. The approach is computationally light-weight, requires no extra training or fine-tuning, and is in principle model-agnostic. We propose both supervised and unsupervised variants by leveraging either ground-truth or solely internal tracking information, and demonstrate its effectiveness for two established tracking models across four video benchmarks. As such, sequential testing can offer a statistically grounded and efficient mechanism to incorporate safety assurances into real-time tracking systems.
Existing point tracking methods face significant challenges in real-world scenarios—including high motion complexity, frequent occlusions, and large object diversity—yet lack a systematic benchmark for evaluating robustness and failure modes. To address this, we introduce ITTO, the first high-challenge dynamic point tracking benchmark, comprising first-person real-world videos and multi-source data. We propose a multi-stage manual annotation protocol to precisely characterize motion patterns, occlusion events, and appearance variations. ITTO introduces a novel performance analysis protocol stratified by motion complexity and, for the first time, systematically exposes critical failure points of mainstream trackers—particularly in post-occlusion re-identification. Experiments reveal substantial performance degradation of state-of-the-art methods on ITTO, highlighting deficiencies in long-term occlusion handling and complex dynamic modeling. These findings provide concrete diagnostic insights and quantitative evaluation standards to guide algorithmic improvement.
Detecting and tracking camouflaged objects—such as military camouflage or biological mimics—remains challenging due to low contrast, ambiguous boundaries, and complex backgrounds. To address this, we introduce COTD, the first dedicated benchmark for camouflaged object tracking, comprising 200 sequences and 80,000 frames, and systematically evaluate 20 state-of-the-art trackers, revealing substantial performance degradation under low-contrast and edge-ambiguous conditions. To overcome these limitations, we propose HiPTrack-MLS: a Transformer-CNN hybrid architecture featuring hierarchical feature enhancement, high-fidelity pixel-level attention, and adaptive mask-guided multi-scale similarity learning. On COTD, HiPTrack-MLS achieves 70.2% AUC—surpassing 70% for the first time—and outperforms prior art by 12.6%. Both the dataset and source code are fully open-sourced to foster standardized, reproducible research in camouflaged object tracking.
Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.
Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.
This work addresses the limitations of existing evaluation methods that focus solely on final outcomes, which fail to distinguish reliable reasoning from accidental success or diagnose process-level flaws in long-horizon tasks. To this end, we propose ClawTrack, a dual-dimensional evaluation framework that jointly assesses task completion (Task Score) and reasoning process quality (Process Score). Spanning 320 tasks across eight domains, ClawTrack introduces fine-grained, stepwise scoring along four dimensions, enabling the first interpretable, process-level evaluation of autonomous agent reasoning trajectories. Our Process Grader combines rule-based logic with large language models, incorporating 12,541 task-specific scoring criteria and integrating over 25 deterministic simulation environments. Validation across 21 models and more than 16,000 trials demonstrates that process scores effectively attribute success or failure, filter out spurious successes, and—when used to select high-quality reasoning trajectories—significantly boost performance across model scales, with consistent results across different evaluator models.