π€ AI Summary
This study addresses the evaluation challenge in image generation where outputs are visually realistic yet factually inconsistent. To this end, it constructs a benchmark comprising 1,274 sample pairs and proposes the PERSIST-Agent framework. This framework introduces a novel paradigm for evaluating world consistency in single images without predefined claims, integrating an agent-based architecture with iterative retrieval augmentation. Through a persistent verification state mechanism, it enables iterative validation of semantic consistency while performing self-optimizing prompt engineering under fixed model weights. Experimental results demonstrate that PERSIST-Agent significantly improves pairwise accuracy on 8B-parameter models, effectively reveals detector biases, and narrows the gap between visual realism and factual accuracy.
π Abstract
Advances in image generation have made visual authenticity increasingly difficult to assess. Although image forensics now examines both generation artifacts and higher-level visual inconsistencies, a plausible image can still contradict real-world facts or rules. We introduce WorldFact-Bench to evaluate image-world consistency from a single image, without a predefined claim or verification target. The benchmark contains 1,274 source-aligned real-fake pairs across four verification regimes and ten semantic domains. Each pair introduces a specific, evidence-supported factual conflict while seeking to preserve non-target content and visual plausibility. Images are evaluated independently, and pair accuracy requires both members of a pair to be classified correctly. We further propose PERSIST-Agent, which organizes iterative verification around a persistent state linking candidate facts, visual observations, evidence, and verification statuses. This state guides subsequent inspection and retrieval while retaining unresolved candidates. With backbone weights fixed, harness self-optimization refines the agent's prompts and execution rules through validation feedback. Experiments reveal strong label biases in several detectors and uneven gains from retrieval. On the evaluated 8B backbones, PERSIST-Agent improves pair accuracy over both direct judgment and retrieval-augmented baselines, while ablations support the role of persistent verification state. These findings highlight the value of state-guided verification and the remaining gap between visual plausibility and factual correctness.