🤖 AI Summary
This study addresses the challenges of oriented vehicle detection in low-light UAV scenarios, where spatially heterogeneous modality reliability and weak small-object features degrade performance. To this end, we propose a Reliability-Conditioned Differential Representation Network that introduces a novel joint mechanism for evidence selection and complementary recovery based on local modality reliability. Specifically, degradation-aware learning estimates modality credibility to guide adaptive evidence selection, while uncertainty-guided differential recovery extracts complementary discriminative information from ambiguous regions, yielding a unified multimodal representation. Experiments demonstrate that the proposed method achieves mAP50 accuracies of 85.3% and 73.9% on the DroneVehicle and VEDAI datasets, respectively. These results confirm its effectiveness in resolving modality fusion difficulties and boundary ambiguity for small objects under low-light conditions.
📝 Abstract
Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially varying modality reliability, while the small visual extent of vehicles further weakens boundaries, orientation cues, and thermal responses. Accordingly, selecting trustworthy observations based on local modality reliability while further exploiting complementary discriminative information in regions with ambiguous modality preference is key to constructing effective multimodal representations. Based on this insight, we propose ReDiffNet, a reliability-conditioned differential representation network in which modality reliability guides both evidence selection and complementary recovery. Specifically, degradation-aware reliability learning estimates relative spatial reliability, uncertainty-guided differential recovery exploits cross-modal differences to recover complementary cues in ambiguous regions, and reliability-conditioned reconstruction integrates retained and recovered evidence into a unified representation. ReDiffNet achieves 85.3% and 73.9% mAP50 on DroneVehicle and VEDAI, respectively, supporting its effectiveness.