Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the representation gap between template and search frames caused by sensor discrepancies in cross-modal object tracking. To this end, we propose TSDA-Track, a novel framework that introduces the first template-search domain adaptation mechanism based on a Transformer architecture. By employing a multi-stage strategy comprising adversarial pre-alignment and contrastive post-encoding alignment, the method effectively mitigates modality bias during training and enhances cross-modal feature correspondence, enabling unified inference without modality-specific branches. Zero-shot evaluations on the LasHeR and RGBT234 datasets demonstrate that TSDA-Track significantly outperforms existing state-of-the-art methods. Furthermore, experiments validate its applicability in aerial tracking scenarios, highlighting its robustness across diverse sensing conditions.
📝 Abstract
Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking where paired modalities are simultaneously available, cross-modal tracking requires localization when template and search frames originate from different active modalities. Accordingly, we introduce TSDA-Track, a Template-Search Domain Adaptation framework to reduce modality discrepancy during training. We investigate two feature alignment strategies. Pre-AFA TSDA-Track applies adversarial alignment before transformer's template-search interaction to suppress modality-specific bias. Enc-CFA TSDA-Track applies contrastive alignment to encoder representations after interaction to strengthen target-level cross-modal correspondence. Both variants retain a shared inference pipeline without modality-specific branches. Experiments on LasHeR, and zero-shot evaluations on RGBT234 and GTOT under multiple cross-modal protocols demonstrate improvements over representative state-of-the-art trackers. For instance, under the modality-switch protocol on RGBT234, Pre-AFA TSDA-Track achieves an SR/PR of 43.2/56.0, compared with 36.8/50.0 for ToMP-101 baseline. In addition, a study on Anti-UAV-024 further verifies the applicability of TSDA-Track to aerial tracking. Our study highlights the effectiveness of feature alignment domain adaptation for cross-modal tracking.
Problem

Research questions and friction points this paper is trying to address.

Cross-Modal Object Tracking
Domain Adaptation
Modality Discrepancy
Template-Search Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Object Tracking
Domain Adaptation
Feature Alignment
Adversarial Learning
Contrastive Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fereshteh Aghaee Meibodi
Department of Electrical and Computer Engineering, University of Victoria
A
Amir Mehdi Soufi Enayati
Department of Mechanical Engineering, University of Victoria
S
Shadi Alijani
Department of Mechanical Engineering, University of Victoria
Homayoun Najjaran
Homayoun Najjaran
University of Victoria
ControlRoboticsAutomation