MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low reliability, fragmented evidence, and computational bottlenecks encountered when large language models validate hypotheses concerning metal-organic frameworks. To overcome these challenges, this work proposes a fault-aware agent architecture. A multi-task diagnostic benchmark is first constructed to precisely identify failure modes. Subsequently, retrieval-augmented generation, machine learning potentials, and workflow orchestration techniques are integrated to systematically optimize structural parsing, literature retrieval, evidence synthesis, and computational verification. Experimental results demonstrate that the proposed approach significantly outperforms direct reasoning and retrieval-based baselines, substantially enhancing hypothesis validation performance. Furthermore, the associated benchmark dataset has been made publicly available to facilitate future research.
📝 Abstract
Large language models are increasingly used as reasoning components in AI-driven materials Co-Scientists, yet the reliability of the resulting verification pipeline remains unclear. Metal-organic frameworks (MOFs) provide a particularly challenging setting because structures may appear under different identifiers, synthesis outcomes depend strongly on experimental conditions, evidence is distributed across heterogeneous sources, and some hypotheses require computation rather than literature alone. We introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification. T-MOF-1-3 are evaluated under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, while T-MOF-4 separately evaluates computational verification. Guided by these diagnosed failure modes, we develop MOF-Verify, a failure-aware agentic harness that targets structural, literature, evidence-sufficiency, and computational bottlenecks before producing a final verdict. Across multiple backbone LLMs, MOF-Verify substantially improves hypothesis-verification performance over direct inference and retrieval-based baselines. Benchmark datasets are released at https://github.com/IMMS-Ewha/MOF-Verify-Benchmark.
Problem

Research questions and friction points this paper is trying to address.

Metal-organic frameworks
Hypothesis verification
Large language models
Failure diagnosis
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Failure-Aware Agentic Harness
Hypothesis Verification
Metal-Organic Frameworks
Diagnostic Benchmark
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Donghyun Lee
Donghyun Lee
Research Institute, National Cancer Center Korea
Medical Data ProcessingMedical Artificial Intelligence
Taehoon Lee
Taehoon Lee
Project Leader @ AI R&D Center, SK Telecom
AI RoboticsMachine LearningComputer Vision
G
Geonhee Ahn
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea
Jieun Kim
Jieun Kim
Associate professor, Hanyang University
UI/UX designInclusive DesignHuman computer interaction
Jihyun Park
Jihyun Park
Assistant Professor, Ewha Womans University, Department of Architecture
IEQIndoor AirPOEBuilding PerformanceHuman Factors
S
Suyeon Cho
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea
Y
Yoona Kim
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea
C
Chaerim Shin
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea
H
Hoi Ri Moon
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea
Jonggeol Na
Jonggeol Na
Ewha Womans University
process systems engineeringmachine learningmultiscale modeling
S
Sukho Hong
Lymeric, Seoul, Republic of Korea
Jihwan Oh
Jihwan Oh
KAIST
Human-AI AlignmentData-Centric AILarge Language ModelMachine Learning
S
Soo Kyung Kim
Institute for Multiscale Matters and System (IMMS), Ewha Womans University, Seoul, Republic of Korea