ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of identity error propagation and uncertainty loss caused by early erroneous associations in cross-camera tracking. To this end, it proposes a Recurrent Event World Model that leverages recurrent neural networks to construct persistent belief states, preserving matching candidate uncertainties while dynamically updating beliefs via Bayesian posterior estimation. Furthermore, by integrating language-guided feature fusion with multimodal evidence association, the method explicitly incorporates association uncertainty into future predictions, effectively maintaining trajectory diversity and guiding identity decisions. Experimental results demonstrate that the proposed approach achieves HOTA scores of 65.19 and 45.36 on the CityFlowV2 and MTMMC datasets, respectively, significantly improving both identity continuity and prediction accuracy.
📝 Abstract
Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.
Problem

Research questions and friction points this paper is trying to address.

Multi-Camera Tracking
Language-Guided Tracking
Identity Association
Uncertainty Propagation
Camera Handoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Event World Model
Language-Guided Multi-Camera Tracking
Association Uncertainty
Structured Posterior Update
Recurrent Belief
💼 Related Jobs
No related jobs found.
Haoyang Wu
Haoyang Wu
Intel Labs China
wireless sensingwireless network architecturesoftware-defined radio
S
Shoudong Han
Huazhong University of Science and Technology
C
Chaoyue Li
Huazhong University of Science and Technology
S
Sijia Chen
Huazhong University of Science and Technology
Z
Zhenyang Xie
Jiangxi University of Water Resources and Electric Power
W
Wang sihan
Zhongnan University of Economics and Law