Score
Design loss functions and supervision schemes that enforce spatiotemporal consistency of predicted trajectories by aligning predicted and reference trajectories at pointwise, windowed, and whole-trajectory scales. This work includes finite-difference penalties on velocity and acceleration and heading, pixel- or feature-level temporal alignment, trajectory registration to local targets, and stepwise constraints to prevent discontinuities and preserve smooth object motion.
Existing evaluation metrics for visual object tracking lack a comparable, continuous-time measure for the trajectory function of time (FoT), relying instead on discrete-frame assessments that fail to characterize arbitrary-time states or disentangle distinct error types (e.g., localization, false positives, missed detections). Method: We propose Star-ID—the first spatiotemporally aligned trajectory integral distance—defining a rigorous, comparable FoT metric over continuous spacetime. Star-ID strictly distinguishes temporally aligned versus misaligned trajectory segments and analytically decouples detection and localization errors. It introduces time-averaged metrics and a theoretical error decomposition model, supported by a multi-object numerical validation framework. Contribution/Results: We provide formal theoretical analysis and demonstrate—via both single- and multi-object simulations—that Star-ID significantly enhances physical interpretability and fine-grained discriminative power in tracking evaluation, enabling precise, continuous-time performance assessment.
To address poor generalization and robustness of trajectory prediction models caused by the absence of ground-truth trajectory annotations, this paper proposes the first fully annotation-free, end-to-end training framework for trajectory prediction. Methodologically: (1) a high-accuracy self-supervised pose and velocity estimation module replaces ground-truth inputs; (2) an input-quality-aware noise-resilient training paradigm explicitly models how input noise degrades prediction reliability; (3) the Trajectron++ architecture is enhanced with few-shot generalization strategies. Contributions include: the first ground-truth-agnostic trajectory prediction training framework; uncovering the intrinsic relationship between input quality and prediction stability; and achieving, across diverse environments, a 42% improvement in prediction stability under noise, few-shot performance approaching that of fully supervised baselines, and significantly enhanced trajectory smoothness and cross-scenario generalization.
In pedestrian trajectory prediction, supervised learning suffers from long-tailed data distributions and struggles to model anomalous behaviors such as abrupt stops or sharp turns. To address this, we propose the first self-supervised framework that explicitly and jointly models position, velocity, and acceleration. Our method introduces a hierarchical velocity/acceleration feature injection architecture, enforces physical kinematic consistency via a novel self-supervised mechanism, and integrates a pseudo-label generation strategy to enable cooperative prediction and dynamic coupling constraints among the three motion variables. Crucially, the framework requires no ground-truth velocity or acceleration annotations—only raw trajectory coordinates are needed for training. Evaluated on ETH-UCY and Stanford Drone datasets, our approach achieves state-of-the-art performance, demonstrating significant improvements in both prediction accuracy and robustness for anomalous motions.
This work addresses the representational limitations and modeling bias of linear trajectory models in autonomous driving motion prediction. We systematically evaluate their fitting performance across vehicle, cyclist, and pedestrian trajectories. To mitigate overfitting and improve generalization, we propose the first empirical Bayes-based joint estimation framework that simultaneously infers the observation noise distribution and model parameter priors from heterogeneous real-world trajectory data—enabling prior-driven, regularized modeling. Empirical results demonstrate that linear models achieve high fidelity (mean fitting error < 0.2 m over 2 s) despite minimal complexity, outperforming sophisticated nonlinear baselines. Cross-modal statistical analysis further reveals strong generalization robustness across traffic participants. Our findings provide theoretical foundations and methodological support for lightweight, interpretable, and production-deployable motion prediction systems.
This work addresses the challenge of inferring smooth trajectories from discrete, unpaired observational snapshots by proposing an efficient method that circumvents the need for costly preprocessing or trajectory simulation inherent in existing approaches. The method lifts the interpolation problem into phase space, where it learns an explicit conditional acceleration field via regression and leverages stochastic differential equations to generate smooth trajectories consistent with the given marginal distributions. Requiring only positional data for training—without trajectory alignment or simulation—it achieves substantial gains in computational efficiency and scalability. Experimental results demonstrate that the proposed approach matches or outperforms current state-of-the-art methods across multiple benchmark tasks, confirming its effectiveness and competitiveness.
This work addresses the challenges of small object detection with event cameras, where sparse, asynchronous signals and weak responses are highly susceptible to noise, leading to temporal discontinuity and unstable predictions. To overcome these limitations, the authors propose PACT, a physics-guided convection consistency modeling framework that, for the first time, introduces a convection conservation mechanism to model event evolution as a motion-driven feature transport process. Specifically, features are propagated along an estimated velocity field using a differentiable convection operator, enabling motion-aware feature extraction and noise suppression, while trajectory-level consistency constraints preserve temporal continuity of weak responses. Experiments demonstrate that PACT achieves a 20.72% improvement in IoU and a 15.03% gain in accuracy on standard event-based datasets, with computational efficiency comparable to existing methods.