RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark

📅 2025-03-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
RAW image visual understanding is constrained by the sRGB pretraining paradigm, limiting performance on native RAW data. Method: This paper proposes an end-to-end RAW adaptation framework featuring a learnable ISP input adapter and model-level adapters operating in synergy to decouple physics-based ISP modeling from semantic understanding. It introduces RAW-native data augmentation and cross-domain robust training, built upon a lightweight adapter architecture, differentiable ISP modeling, and multi-granularity feature alignment. Contribution/Results: We establish RAW-Bench—the first benchmark covering 17 realistic RAW degradations—and demonstrate state-of-the-art performance with 62% fewer parameters and 2.3× faster inference. The method exhibits strong generalization and robustness under challenging conditions including low-light, rain, fog, and motion blur.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
In the computer vision community, the preference for pre-training visual models has largely shifted toward sRGB images due to their ease of acquisition and compact storage. However, camera RAW images preserve abundant physical details across diverse real-world scenarios. Despite this, most existing visual perception methods that utilize RAW data directly integrate image signal processing (ISP) stages with subsequent network modules, often overlooking potential synergies at the model level. Building on recent advances in adapter-based methodologies in both NLP and computer vision, we propose RAW-Adapter, a novel framework that incorporates learnable ISP modules as input-level adapters to adjust RAW inputs. At the same time, it employs model-level adapters to seamlessly bridge ISP processing with high-level downstream architectures. Moreover, RAW-Adapter serves as a general framework applicable to various computer vision frameworks. Furthermore, we introduce RAW-Bench, which incorporates 17 types of RAW-based common corruptions, including lightness degradations, weather effects, blurriness, camera imaging degradations, and variations in camera color response. Using this benchmark, we systematically compare the performance of RAW-Adapter with state-of-the-art (SOTA) ISP methods and other RAW-based high-level vision algorithms. Additionally, we propose a RAW-based data augmentation strategy to further enhance RAW-Adapter's performance and improve its out-of-domain (OOD) generalization ability. Extensive experiments substantiate the effectiveness and efficiency of RAW-Adapter, highlighting its robust performance across diverse scenarios.
Problem

Research questions and friction points this paper is trying to address.

Adapting pre-trained visual models to RAW images.
Integrating ISP modules with downstream vision tasks.
Benchmarking RAW-based corruptions for model evaluation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses learnable ISP modules as input-level adapters
Employs model-level adapters for ISP-task synergy
Introduces RAW-Bench with 17 corruption types
🔎 Similar Papers
No similar papers found.