🤖 AI Summary
This work addresses the challenge of applying vision foundation models—typically trained on RGB imagery—to object detection in non-visible spectral domains such as infrared (IR) and synthetic aperture radar (SAR). To bridge this gap, the authors propose the first integration of the flow-matching foundation model FLUX.1 Kontext with low-rank adaptation (LoRA), enabling high-quality cross-spectral image translation using only 100 paired images per domain. The translated images are leveraged to generate labeled synthetic data for training lightweight detectors, YOLOv11n and DETR. Experiments on the KAIST and M4-SAR datasets demonstrate that the synthetic data substantially improves mean average precision (mAP) for downstream detection tasks. Furthermore, the study validates that the Learned Perceptual Image Patch Similarity (LPIPS) metric serves as an effective proxy for predicting detection performance, offering a practical tool for evaluating translation quality without requiring extensive retraining.
📝 Abstract
Foundation models for vision are predominantly trained on RGB data, while many safety-critical applications rely on non-visible modalities such as infrared (IR) and synthetic aperture radar (SAR). We study whether a single flow-matching foundation model pre-trained primarily on RGB images can be repurposed as a cross-spectral translator using only a few co-measured examples, and whether the resulting synthetic data can enhance downstream detection. Starting from FLUX.1 Kontext, we insert low-rank adaptation (LoRA) modules and fine-tune them on just 100 paired images per domain for two settings: RGB to IR on the KAIST dataset and RGB to SAR on the M4-SAR dataset. The adapted model translates RGB images into pixel-aligned IR/SAR, enabling us to reuse existing bounding boxes and train object detection models purely in the target modality. Across a grid of LoRA hyperparameters, we find that LPIPS computed on only 50 held-out pairs is a strong proxy for downstream performance: lower LPIPS consistently predicts higher mAP for YOLOv11n on both IR and SAR, and for DETR on KAIST IR test data. Using the best LPIPS-selected LoRA adapter, synthetic IR from external RGB datasets (LLVIP, FLIR ADAS) improves KAIST IR pedestrian detection, and synthetic SAR significantly boosts infrastructure detection on M4-SAR when combined with limited real SAR. Our results suggest that few-shot LoRA adaptation of flow-matching foundation models is a promising path toward foundation-style support for non-visible modalities.