🤖 AI Summary
This study addresses the reliance on expensive annotations and the sensitivity to noisy pseudo-labels in RGB-T tracking by proposing ESMTrack, an end-to-end self-supervised framework. Requiring only first-frame annotation, the method employs a tri-branch architecture that integrates visual grounding with a temporal triplet loss. Furthermore, it introduces a novel APCE-based adaptive weighting mechanism for modality reliability, which, combined with forward-backward consistency filtering and contrastive learning, achieves effective modality decoupling and eliminates offline pseudo-label bias. Extensive experiments demonstrate that ESMTrack attains state-of-the-art performance across five major benchmarks while maintaining strong generalization capability and real-time inference speed.
📝 Abstract
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.