Inference-Time Scaling of Diffusion Models for Infrared Data Generation

📅 2025-11-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Infrared image annotation data is scarce, hindering the development of downstream vision models. To address this, we propose an inference-time scaling framework that fine-tunes the FLUX.1-dev diffusion model on few-shot infrared data and integrates a domain-adapted CLIP-based verifier to dynamically enforce contrastive scoring and text–image alignment guidance during sampling. Our method requires no additional generator training; instead, it leverages a lightweight verifier to enhance generation quality. On the KAIST dataset, it achieves a 10% reduction in FID over the unguided baseline, significantly improving both image fidelity and text–image consistency. To our knowledge, this is the first work to combine inference-time CLIP guidance with parameter-efficient fine-tuning for low-data infrared image generation. It effectively bridges the domain gap between visible-light and infrared modalities and establishes a novel paradigm for generative modeling in resource-constrained imaging domains.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Vertical and domain-specific searchWeb Mining and Content Analysis: Large pretrained models with web dataEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
Infrared imagery enables temperature-based scene understanding using passive sensors, particularly under conditions of low visibility where traditional RGB imaging fails. Yet, developing downstream vision models for infrared applications is hindered by the scarcity of high-quality annotated data, due to the specialized expertise required for infrared annotation. While synthetic infrared image generation has the potential to accelerate model development by providing large-scale, diverse training data, training foundation-level generative diffusion models in the infrared domain has remained elusive due to limited datasets. In light of such data constraints, we explore an inference-time scaling approach using a domain-adapted CLIP-based verifier for enhanced infrared image generation quality. We adapt FLUX.1-dev, a state-of-the-art text-to-image diffusion model, to the infrared domain by finetuning it on a small sample of infrared images using parameter-efficient techniques. The trained verifier is then employed during inference to guide the diffusion sampling process toward higher quality infrared generations that better align with input text prompts. Empirically, we find that our approach leads to consistent improvements in generation quality, reducing FID scores on the KAIST Multispectral Pedestrian Detection Benchmark dataset by 10% compared to unguided baseline samples. Our results suggest that inference-time guidance offers a promising direction for bridging the domain gap in low-data infrared settings.
Problem

Research questions and friction points this paper is trying to address.

Addresses scarcity of annotated infrared data for vision models
Enhances synthetic infrared image generation quality using diffusion models
Bridges domain gap in low-data infrared settings through inference-time guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inference-time scaling with domain-adapted CLIP verifier
Fine-tuning FLUX.1-dev model using parameter-efficient techniques
Guiding diffusion sampling for improved infrared generation quality
🔎 Similar Papers