TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
This study addresses the limitation of existing evaluation metrics in capturing visual inconsistencies that violate physical laws, anatomical structures, or spatial relationships in text-to-image models. To this end, we propose TerraVis, a framework that establishes world consistency as an independent evaluation dimension. TerraVis constructs a hierarchical taxonomy of violations spanning object, interaction, and scene levels, and designs a multi-stage automated detection and scoring pipeline leveraging multimodal large language models. Experimental results demonstrate that TerraVis achieves the highest correlation with human judgments on mainstream benchmarks, effectively revealing significant real-world logical deficiencies that persist even in traditionally high-scoring models.