OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current text-to-image generation models frequently violate fundamental physical commonsense, yet prevailing evaluation benchmarks lack the granularity needed to accurately diagnose their physical reasoning capabilities. To address this, this work introduces OmniPhys—the first fine-grained evaluation benchmark grounded in a physics knowledge graph—and proposes OmniPrompt, an iterative optimization framework that formulates physical alignment as a discrete optimization problem. By leveraging a multi-image feedback buffer and batched meta-policy updates, OmniPrompt enables noise-robust collective optimization. Coupled with a dual-path validation protocol and alignment to PhET simulations, the approach effectively mitigates gradient hallucination. Extensive experiments across twelve mainstream text-to-image models uncover pervasive physical reasoning bottlenecks and demonstrate substantial improvements in physical consistency across diverse backbone architectures, validating the transferability and efficacy of the proposed meta-policy.
📝 Abstract
While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptions, failing to diagnose the mastery of specific physical principles. Moreover, the high stochasticity of generative processes causes current prompt optimization methods to suffer from gradient hallucinations, where optimizers are misled by transient visual artifacts rather than systemic flaws. To address these challenges, we introduce OmniPhys, a rigorous benchmark of 1,551 samples grounded in a Physical Knowledge Graph. By aligning PhET simulations with standard curricula, OmniPhys operationalizes a knowledge-to-scenario pipeline that performs diagnostic stress tests via a dual-path verification protocol. We further propose OmniPrompt, an iterative framework that treats physical alignment as a discrete optimization problem. For each query, OmniPrompt aggregates K stochastic images into a per-query feedback buffer. Across training, it further merges feedback from batches of B queries before each meta-policy update, filtering seed and query-local noise. Evaluations across 12 representative text-to-image models reveal universal physical bottlenecks. Results demonstrate that OmniPrompt significantly enhances physical consistency across diverse backbones, proving the transferability and efficacy of our evolved meta-policies. The code and data are available at https://github.com/zjukg/OmniPhys
Problem

Research questions and friction points this paper is trying to address.

physical commonsense
text-to-image generation
benchmarking
prompt optimization
gradient hallucination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical Commonsense
Knowledge Graph
Text-to-Image Generation
Prompt Optimization
Discrete Optimization
🔎 Similar Papers
2024-06-09arXiv.orgCitations: 3