🤖 AI Summary
This study addresses the low reliability, fragmented evidence, and computational bottlenecks encountered when large language models validate hypotheses concerning metal-organic frameworks. To overcome these challenges, this work proposes a fault-aware agent architecture. A multi-task diagnostic benchmark is first constructed to precisely identify failure modes. Subsequently, retrieval-augmented generation, machine learning potentials, and workflow orchestration techniques are integrated to systematically optimize structural parsing, literature retrieval, evidence synthesis, and computational verification. Experimental results demonstrate that the proposed approach significantly outperforms direct reasoning and retrieval-based baselines, substantially enhancing hypothesis validation performance. Furthermore, the associated benchmark dataset has been made publicly available to facilitate future research.
📝 Abstract
Large language models are increasingly used as reasoning components in AI-driven materials Co-Scientists, yet the reliability of the resulting verification pipeline remains unclear. Metal-organic frameworks (MOFs) provide a particularly challenging setting because structures may appear under different identifiers, synthesis outcomes depend strongly on experimental conditions, evidence is distributed across heterogeneous sources, and some hypotheses require computation rather than literature alone. We introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification. T-MOF-1-3 are evaluated under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, while T-MOF-4 separately evaluates computational verification. Guided by these diagnosed failure modes, we develop MOF-Verify, a failure-aware agentic harness that targets structural, literature, evidence-sufficiency, and computational bottlenecks before producing a final verdict. Across multiple backbone LLMs, MOF-Verify substantially improves hypothesis-verification performance over direct inference and retrieval-based baselines. Benchmark datasets are released at https://github.com/IMMS-Ewha/MOF-Verify-Benchmark.