Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance degradation in RGBT tracking caused by incomplete and unstable feature representations under modality dropout. To tackle this challenge, we propose the Spatio-Temporal Conditional Denoising Transformer (SCDT), a unified framework that integrates spatial cues, fine-grained short-term temporal correlations, and global long-term modality evolution context. SCDT employs a conditional denoising strategy to adaptively reconstruct missing modalities and enhance weak modality features. A key innovation is the dynamic noise modulation mechanism, which enables a single model to flexibly handle both modality-complete and modality-missing scenarios without architectural modifications. Extensive experiments on three mainstream RGBT benchmarks demonstrate that SCDT significantly outperforms state-of-the-art methods, confirming its robustness and effectiveness.
📝 Abstract
Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.
Problem

Research questions and friction points this paper is trying to address.

modality-missing
RGBT tracking
multimodal feature representation
temporal context
feature reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatio-Temporal Conditional Denoising
Modality-Missing RGBT Tracking
Temporal Context Integration
Noise-Modulated Adaptation
Unified Multimodal Representation
🔎 Similar Papers
2023-12-25International Journal of Computer VisionCitations: 0