π€ AI Summary
This study addresses the insufficient robustness of ensemble models against component failures and the evaluation bias arising from single-corpus assessments in AI-generated image detection. Through rigorous controlled experiments and preprocessing path analysis, it reveals the limitations of stacking ensembles that rely on in-domain retraining. Revising earlier conclusions, this work proposes an architecture-diversity-based majority voting rule, validated via a gradient-boosted meta-learner and McNemarβs test. It demonstrates that uncalibrated ensembles underperform optimal individual models, thereby establishing the necessity of domain adaptation. Furthermore, the proposed voting rule effectively prevents cascading failures without retraining while exposing deficiencies in abstention handling and potential JPEG compression biases. All code and corrected results have been made publicly available.
π Abstract
We built an ordinary stacking ensemble for AI-generated image detection -- three open detectors producing five scores, fused by a gradient-boosted meta-learner that treats a detector's failure as missing data -- deployed it, and then evaluated it against three controls it should have faced first. This paper reports what the controls found, including where they overturned our own earlier conclusions. Fusion is worth its cost only when refitted on the target domain. The shipped meta-learner, fitted on a separate corpus, does not beat its best single member on 2000 StyleGAN faces (AUC 0.9896 vs 0.9961; McNemar p = 1.000). But a stacker refitted in-domain beats that member plus a post-hoc calibrator (Delta-AUC = +0.0025, [+0.0014, +0.0038]; p = 3.4e-4). An earlier draft claimed the calibrated single detector won outright; that comparison mixed regimes and we correct it here. One corpus is not an evaluation. On 80 screenshots every model's AUC interval contains 0.5. We can say nothing stronger: the difference between the ensemble's drop and its best member's is [-0.185, +0.115]. Abstention is real, correlated, and mishandled. With four of five detectors silent and the survivor reporting"real", the system returns P(AI) = 0.9985, because an all-NaN input scores 0.9995 in a learner never fitted with missingness. The three AIDE checkpoints fail together, sharing one preprocessing path. A quorum rule requiring two distinct architectures prevents both failures with no retraining. We also find Corpus A carries a class-conditional JPEG bias severe enough to separate the classes from the header alone, which limits every in-domain number we report. Code, harness, hash-identified artifacts and all corrections are released.