FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trust and security risks arising when deployed AI image generators are silently replaced with lower-quality models. We propose FARE, a black-box integrity auditing framework that detects such substitutions using only output images, without requiring access to model weights. Specifically, FARE constructs an acceptance region by extracting forensic artifact features from an authenticated generator, and incorporates hard example mining to amplify distinctive artifacts and tighten decision boundaries, thereby enhancing sensitivity to subtle model alterations. Experimental results demonstrate that FARE effectively identifies substitution attacks across diverse scenarios, significantly outperforming existing baselines at strict operating points. Furthermore, the framework exhibits robustness against precise-model and decision-only adversarial attacks.
📝 Abstract
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator---using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under the exact-model and decision-only attacks evaluated in this work.
Problem

Research questions and friction points this paper is trying to address.

bait-and-switch
image generator auditing
integrity verification
black-box API
forensic detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Forensic Acceptance Region Estimation
Bait-and-Switch Detection
Image Generator Artifacts
Hard Sample Mining
Integrity Auditing