End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on expensive annotations and the sensitivity to noisy pseudo-labels in RGB-T tracking by proposing ESMTrack, an end-to-end self-supervised framework. Requiring only first-frame annotation, the method employs a tri-branch architecture that integrates visual grounding with a temporal triplet loss. Furthermore, it introduces a novel APCE-based adaptive weighting mechanism for modality reliability, which, combined with forward-backward consistency filtering and contrastive learning, achieves effective modality decoupling and eliminates offline pseudo-label bias. Extensive experiments demonstrate that ESMTrack attains state-of-the-art performance across five major benchmarks while maintaining strong generalization capability and real-time inference speed.
📝 Abstract
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.
Problem

Research questions and friction points this paper is trying to address.

RGB-T tracking
self-supervised learning
end-to-end optimization
modality dominance bias
pseudo-labeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised RGB-T Tracking
End-to-End Optimization
Triplet Loss
Modality Decoupling
Cross-Modal Contrastive Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shenglan Li
School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology, China; Mine Digitization Engineering Research Center of the Ministry of Education, China
Rui Yao
Rui Yao
China University of Mining and Technology
Computer VisionMachine Learning
K
Kunyang Sun
School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology, China; Mine Digitization Engineering Research Center of the Ministry of Education, China
Hong Jia
Hong Jia
Lecturer (Assistant Professor), University of Auckland; University of Melbourne
On-Device MLHuman-Centred AIMobile ComputingMobile Health
Y
Yong Zhou
School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology, China; Mine Digitization Engineering Research Center of the Ministry of Education, China
J
Javen Qinfeng Shi
Australian Institute for Machine Learning, Adelaide University, Australia
Xinyu Zhang
Xinyu Zhang
The University of Auckland; The University of Adelaide
Computer VisionMachine LearningGenerative AI