🤖 AI Summary
This study addresses the limitation that graph structures cannot determine the sign and magnitude of residual dependencies, resulting in uncontrolled propagation of prediction errors. To mitigate this issue, we propose a Bayesian residual covariance model that derives an exact identity for the expected squared error change under linear updates, along with a quadratic excess error formulation, enabling optimal correction via the posterior mean. Furthermore, four positive semi-definite covariance families are introduced to perform model averaging, thereby accommodating structural uncertainty. Empirical evaluations demonstrate that the proposed approach significantly reduces squared errors across most graph neural network benchmarks while minimizing worst-case recall gaps. The results further confirm that model averaging consistently outperforms the selection of any single covariance family, establishing its efficacy in enhancing predictive reliability within graph-based learning frameworks.
📝 Abstract
Observing a prediction error at one node can help correct predictions elsewhere, but the benefit depends on the residual dependence between nodes. A graph suggests where that dependence might occur, yet does not establish its sign or strength. We develop a Bayesian model of residual covariance to determine how revealed errors should update a fixed feature-based predictor. For a fixed set of revealed nodes, we derive an exact identity for the change in expected squared error under linear residual propagation. The identity characterizes the optimal linear update and expresses the excess error of any other update as an exact quadratic. Under covariance uncertainty, the Bayes-optimal fixed linear update uses the posterior mean covariance. We average over four positive-semidefinite covariance families representing no dependence, positive dependence, negative dependence, and dependence whose sign can alternate with graph distance. Across eight graph benchmarks, four show posterior support within the distance-profile family for negative covariance between neighbors and alternating signs with distance. Among the evaluated methods, ours achieves the smallest worst-case recall gap to the best method on each graph. The method improves squared error over residual propagation on five of eight graphs, but performs substantially worse on one dense heterophilous graph. In these experiments, selecting the highest-posterior family performs similarly to model averaging.