Automated Extraction of Material Properties using LLM-based AI Agents

📅 2025-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Material discovery is hindered by the scarcity of large-scale, machine-readable experimental datasets linking crystal structures to thermoelectric properties. To address limitations of existing databases—including small size, heavy reliance on manual curation, and theoretical bias—this work introduces an LLM-based agent workflow that automatically extracts thermoelectric performance and crystal structure data from nearly 10,000 scientific publications. Our approach innovatively integrates dynamic token allocation, zero-shot multi-agent coordination, and conditional table parsing, achieving high extraction accuracy (F1 > 0.92) while reducing inference cost by 40%. We construct the largest LLM-curated thermoelectric dataset to date—comprising 27,822 temperature-resolved, unit-standardized records—and release an open-source platform supporting semantic search and interactive data export. This significantly enhances scalability and practicality for data-driven materials discovery.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalNatural Language Processing: Information ExtractionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
The rapid discovery of materials is constrained by the lack of large, machine-readable datasets that couple performance metrics with structural context. Existing databases are either small, manually curated, or biased toward first principles results, leaving experimental literature underexploited. We present an agentic, large language model (LLM)-driven workflow that autonomously extracts thermoelectric and structural-properties from about 10,000 full-text scientific articles. The pipeline integrates dynamic token allocation, zeroshot multi-agent extraction, and conditional table parsing to balance accuracy against computational cost. Benchmarking on 50 curated papers shows that GPT-4.1 achieves the highest accuracy (F1 = 0.91 for thermoelectric properties and 0.82 for structural fields), while GPT-4.1 Mini delivers nearly comparable performance (F1 = 0.89 and 0.81) at a fraction of the cost, enabling practical large scale deployment. Applying this workflow, we curated 27,822 temperature resolved property records with normalized units, spanning figure of merit (ZT), Seebeck coefficient, conductivity, resistivity, power factor, and thermal conductivity, together with structural attributes such as crystal class, space group, and doping strategy. Dataset analysis reproduces known thermoelectric trends, such as the superior performance of alloys over oxides and the advantage of p-type doping, while also surfacing broader structure-property correlations. To facilitate community access, we release an interactive web explorer with semantic filters, numeric queries, and CSV export. This study delivers the largest LLM-curated thermoelectric dataset to date, provides a reproducible and cost-profiled extraction pipeline, and establishes a foundation for scalable, data-driven materials discovery beyond thermoelectrics.
Problem

Research questions and friction points this paper is trying to address.

Automated extraction of material properties from scientific literature
Addressing limited machine-readable experimental materials datasets
Developing scalable AI workflow for materials discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-driven workflow autonomously extracts material properties
Integrates dynamic token allocation and multi-agent extraction
Uses GPT models for accurate large-scale data curation
🔎 Similar Papers
No similar papers found.
S
Subham Ghosh
Mehta Family School of Data Science and Artificial Intelligence, Indian Institute of Technology Roorkee
A
Abhishek Tewari
Mehta Family School of Data Science and Artificial Intelligence, Department of Metallurgical and Materials Engineering, Indian Institute of Technology Roorkee