Supervised Pretraining for Material Property Prediction

📅 2025-04-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Addressing the scarcity of large-scale labeled data and the limited generalization capability of self-supervised learning (SSL) in materials property prediction, this paper proposes a proxy-label-based supervised graph neural network pretraining framework. The method follows a two-stage paradigm: proxy-label supervised pretraining followed by task-adaptive fine-tuning. Key contributions include: (1) the first introduction of class-level categorical information as proxy labels for supervised pretraining in materials science, significantly enhancing downstream multi-task generalization; and (2) a graph-structure-preserving noise augmentation strategy that injects perturbations while maintaining physical plausibility and topological consistency. Evaluated on six diverse materials property prediction tasks, the proposed approach reduces mean absolute error (MAE) by 2.0%–6.67% over state-of-the-art SSL baselines, establishing new performance benchmarks.

Technology Category

Machine Learning: Graph-based Machine LearningData Mining & Knowledge Management: Graph Mining, Social Network Analysis & CommunityNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Accurate prediction of material properties facilitates the discovery of novel materials with tailored functionalities. Deep learning models have recently shown superior accuracy and flexibility in capturing structure-property relationships. However, these models often rely on supervised learning, which requires large, well-annotated datasets an expensive and time-consuming process. Self-supervised learning (SSL) offers a promising alternative by pretraining on large, unlabeled datasets to develop foundation models that can be fine-tuned for material property prediction. In this work, we propose supervised pretraining, where available class information serves as surrogate labels to guide learning, even when downstream tasks involve unrelated material properties. We evaluate this strategy on two state-of-the-art SSL models and introduce a novel framework for supervised pretraining. To further enhance representation learning, we propose a graph-based augmentation technique that injects noise to improve robustness without structurally deforming material graphs. The resulting foundation models are fine-tuned for six challenging material property predictions, achieving significant performance gains over baselines, ranging from 2% to 6.67% improvement in mean absolute error (MAE) and establishing a new benchmark in material property prediction. This study represents the first exploration of supervised pertaining with surrogate labels in material property prediction, advancing methodology and application in the field.
Problem

Research questions and friction points this paper is trying to address.

Developing self-supervised learning for material property prediction
Improving model robustness with graph-based augmentation techniques
Enhancing accuracy in predicting novel material functionalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised pretraining with surrogate labels
Graph-based augmentation for robust learning
Fine-tuning foundation models for property prediction
💼 Related Jobs
No related jobs found.
C
Chowdhury Mohammad Abid Rahman
Lane Department of Computer Science and Electrical Engineering, West Virginia University
A
Aldo H. Romero
Department of Physics and Astronomy, West Virginia University
P
P. Gyawali
Lane Department of Computer Science and Electrical Engineering, West Virginia University