🤖 AI Summary
This study addresses the issue that existing calibration methods for node classification in graph neural networks overlook local homophily, resulting in miscalibrated confidence estimates. To this end, we propose HoTS, a homophily-aware temperature scaling post-hoc calibrator. This work provides the first theoretical derivation of the inverse relationship between temperature and local homophily, constructing the framework based on the contextual stochastic block model, entropy-based logit concentration, and local homophily estimation techniques. Extensive experiments across 18 benchmark datasets demonstrate that HoTS achieves the lowest average expected calibration error (4.79%) and yields the most reliable confidence rankings, thereby validating its effectiveness and generalization capability.
📝 Abstract
For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-aware modules without a principled structural form. We study how local graph structure should enter node-level calibration. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse-temperature is governed by a homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately inversely with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three temperature parameters, preserves the predicted class, and learns the strength of the structural correction from calibration data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of 4.79%, the best average rank, and the most reliable confidence ranking in selective classification. Code is available at https://github.com/inu0104/HoTS.