🤖 AI Summary
Material discovery is hindered by the scarcity of large-scale, machine-readable experimental datasets linking crystal structures to thermoelectric properties. To address limitations of existing databases—including small size, heavy reliance on manual curation, and theoretical bias—this work introduces an LLM-based agent workflow that automatically extracts thermoelectric performance and crystal structure data from nearly 10,000 scientific publications. Our approach innovatively integrates dynamic token allocation, zero-shot multi-agent coordination, and conditional table parsing, achieving high extraction accuracy (F1 > 0.92) while reducing inference cost by 40%. We construct the largest LLM-curated thermoelectric dataset to date—comprising 27,822 temperature-resolved, unit-standardized records—and release an open-source platform supporting semantic search and interactive data export. This significantly enhances scalability and practicality for data-driven materials discovery.
📝 Abstract
The rapid discovery of materials is constrained by the lack of large, machine-readable datasets that couple performance metrics with structural context. Existing databases are either small, manually curated, or biased toward first principles results, leaving experimental literature underexploited. We present an agentic, large language model (LLM)-driven workflow that autonomously extracts thermoelectric and structural-properties from about 10,000 full-text scientific articles. The pipeline integrates dynamic token allocation, zeroshot multi-agent extraction, and conditional table parsing to balance accuracy against computational cost. Benchmarking on 50 curated papers shows that GPT-4.1 achieves the highest accuracy (F1 = 0.91 for thermoelectric properties and 0.82 for structural fields), while GPT-4.1 Mini delivers nearly comparable performance (F1 = 0.89 and 0.81) at a fraction of the cost, enabling practical large scale deployment. Applying this workflow, we curated 27,822 temperature resolved property records with normalized units, spanning figure of merit (ZT), Seebeck coefficient, conductivity, resistivity, power factor, and thermal conductivity, together with structural attributes such as crystal class, space group, and doping strategy. Dataset analysis reproduces known thermoelectric trends, such as the superior performance of alloys over oxides and the advantage of p-type doping, while also surfacing broader structure-property correlations. To facilitate community access, we release an interactive web explorer with semantic filters, numeric queries, and CSV export. This study delivers the largest LLM-curated thermoelectric dataset to date, provides a reproducible and cost-profiled extraction pipeline, and establishes a foundation for scalable, data-driven materials discovery beyond thermoelectrics.