🤖 AI Summary
This study addresses the robustness challenge of pixel-level image provenance tracing under adversarial editing by investigating the theoretical limits and security risks of passive verification mechanisms. Methodologically, we derive minimax theoretical upper bounds and construct an interface leakage model based on total variation distance and adversarial distribution shift analysis. Our core contribution reveals that the success of black-box attacks fundamentally stems from verifier information leakage rather than distributional overlap. Experiments confirm that public verifiers such as CLIP are vulnerable to targeted attacks. Accordingly, we advocate decoupling the evaluation of statistical performance ceilings from information leakage risks, thereby establishing a new paradigm for designing secure provenance tracing systems.
📝 Abstract
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.