🤖 AI Summary
The high carbon footprint of density functional theory (DFT) calculations conflicts with the limited reliability of purely data-driven machine learning (ML) models in photovoltaic (PV) material discovery. Method: We propose and systematically evaluate multiple ML/DFT hybrid screening strategies, establishing the first evaluation framework that jointly quantifies prediction accuracy and CO₂ emissions. Leveraging PV-specific descriptors and customized model architectures, we quantitatively assess performance–environment trade-offs across strategies. Contribution/Results: We find that direct bandgap and power-conversion-efficiency prediction outperforms indirect approaches based on absorption spectra; moreover, data-driven models can surpass the accuracy of commonly used DFT functionals. The optimal hybrid strategy retains >95% of DFT-level accuracy while reducing computational carbon emissions by up to 70%. Our study identifies several low-emission, high-throughput screening pathways, providing a reusable methodology and practical guidelines for green, high-throughput materials discovery.
📝 Abstract
Computational screening has become a powerful complement to experimental efforts in the discovery of high-performance photovoltaic (PV) materials. Most workflows rely on density functional theory (DFT) to estimate electronic and optical properties relevant to solar energy conversion. Although more efficient than laboratory-based methods, DFT calculations still entail substantial computational and environmental costs. Machine learning (ML) models have recently gained attention as surrogates for DFT, offering drastic reductions in resource use with competitive predictive performance. In this study, we reproduce a canonical DFT-based workflow to estimate the maximum efficiency limit and progressively replace its components with ML surrogates. By quantifying the CO$_2$ emissions associated with each computational strategy, we evaluate the trade-offs between predictive efficacy and environmental cost. Our results reveal multiple hybrid ML/DFT strategies that optimize different points along the accuracy--emissions front. We find that direct prediction of scalar quantities, such as maximum efficiency, is significantly more tractable than using predicted absorption spectra as an intermediate step. Interestingly, ML models trained on DFT data can outperform DFT workflows using alternative exchange--correlation functionals in screening applications, highlighting the consistency and utility of data-driven approaches. We also assess strategies to improve ML-driven screening through expanded datasets and improved model architectures tailored to PV-relevant features. This work provides a quantitative framework for building low-emission, high-throughput discovery pipelines.