Few-Shot LoRA Adaptation of a Flow-Matching Foundation Model for Cross-Spectral Object Detection

📅 2026-01-07
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of applying vision foundation models—typically trained on RGB imagery—to object detection in non-visible spectral domains such as infrared (IR) and synthetic aperture radar (SAR). To bridge this gap, the authors propose the first integration of the flow-matching foundation model FLUX.1 Kontext with low-rank adaptation (LoRA), enabling high-quality cross-spectral image translation using only 100 paired images per domain. The translated images are leveraged to generate labeled synthetic data for training lightweight detectors, YOLOv11n and DETR. Experiments on the KAIST and M4-SAR datasets demonstrate that the synthetic data substantially improves mean average precision (mAP) for downstream detection tasks. Furthermore, the study validates that the Learned Perceptual Image Patch Similarity (LPIPS) metric serves as an effective proxy for predicting detection performance, offering a practical tool for evaluating translation quality without requiring extensive retraining.

Technology Category

Computer Vision: Diffusion Models for VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Foundation models for vision are predominantly trained on RGB data, while many safety-critical applications rely on non-visible modalities such as infrared (IR) and synthetic aperture radar (SAR). We study whether a single flow-matching foundation model pre-trained primarily on RGB images can be repurposed as a cross-spectral translator using only a few co-measured examples, and whether the resulting synthetic data can enhance downstream detection. Starting from FLUX.1 Kontext, we insert low-rank adaptation (LoRA) modules and fine-tune them on just 100 paired images per domain for two settings: RGB to IR on the KAIST dataset and RGB to SAR on the M4-SAR dataset. The adapted model translates RGB images into pixel-aligned IR/SAR, enabling us to reuse existing bounding boxes and train object detection models purely in the target modality. Across a grid of LoRA hyperparameters, we find that LPIPS computed on only 50 held-out pairs is a strong proxy for downstream performance: lower LPIPS consistently predicts higher mAP for YOLOv11n on both IR and SAR, and for DETR on KAIST IR test data. Using the best LPIPS-selected LoRA adapter, synthetic IR from external RGB datasets (LLVIP, FLIR ADAS) improves KAIST IR pedestrian detection, and synthetic SAR significantly boosts infrastructure detection on M4-SAR when combined with limited real SAR. Our results suggest that few-shot LoRA adaptation of flow-matching foundation models is a promising path toward foundation-style support for non-visible modalities.
Problem

Research questions and friction points this paper is trying to address.

cross-spectral object detection
foundation model adaptation
few-shot learning
non-visible modalities
flow-matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Shot LoRA
Flow-Matching Foundation Model
Cross-Spectral Translation
Synthetic Data for Detection
Non-Visible Modality Adaptation
💼 Related Jobs
No related jobs found.
M
Maxim Clouser
Yrikka Inc.
Kia Khezeli
Kia Khezeli
Cornell University
J
John Kalantari
Yrikka Inc.