TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing evaluation metrics in capturing visual inconsistencies that violate physical laws, anatomical structures, or spatial relationships in text-to-image models. To this end, we propose TerraVis, a framework that establishes world consistency as an independent evaluation dimension. TerraVis constructs a hierarchical taxonomy of violations spanning object, interaction, and scene levels, and designs a multi-stage automated detection and scoring pipeline leveraging multimodal large language models. Experimental results demonstrate that TerraVis achieves the highest correlation with human judgments on mainstream benchmarks, effectively revealing significant real-world logical deficiencies that persist even in traditionally high-scoring models.
📝 Abstract
Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationships. Such failures are not well captured by existing fidelity, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations spanning object-, interaction-, and scene-level failures, and employs a multi-stage evaluation framework to identify and quantify them. Given an image, TerraVis first uses an MLLM to assess its eligibility for evaluation, then detects violations across 18 taxonomy-defined types and classifies them as minor or major to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis achieves the strongest correlation with human judgments of world consistency among existing metrics. Our benchmark results further show that models that achieve strong performance on conventional metrics can still exhibit substantial world-consistency failures. These findings highlight world consistency as a complementary evaluation dimension and demonstrate that TerraVis enables systematic quantification, diagnosis, and comparison of such failures. Our code is publicly available at https://github.com/ShyFoo/TerraVis.
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Grounded Visual Consistency
Text-to-Image Generation
MLLM Workflows
Evaluation Framework
Taxonomy of Violations