🤖 AI Summary
This work addresses regression prediction from heterogeneous, noisy sensor data in the absence of ground-truth labels by proposing the Neural Conjugate Aggregation Model (NCAM). NCAM integrates neural networks with conjugate Gaussian inference within a hierarchical Bayesian framework to unsupervisedly learn each sensor’s bias and reliability, enabling uncertainty decomposition and posterior aggregation for the target variable. To mitigate structural non-identifiability, the model incorporates sensor anchoring and variance regularization. Coupled with locally adaptive Monte Carlo conformal prediction, NCAM yields heteroscedastic prediction intervals that simultaneously offer Bayesian interpretability and finite-sample coverage guarantees. Experiments demonstrate that NCAM significantly outperforms baseline methods—including mean aggregation, probabilistic PCA, and Kalman filtering—on both synthetic and real-world air quality datasets, while providing well-calibrated uncertainty estimates.
📝 Abstract
We study regression-based data fusion under uncertainty, where multiple noisy and biased measurement sources are available but ground-truth labels are absent during training. This setting arises in sensor networks, simulation ensembles, and scientific monitoring systems where supervision is costly or infeasible. We propose the Neural Conjugate Aggregation Model (NCAM), a hierarchical Bayesian framework that combines neural networks with conjugate Gaussian inference for unsupervised multi-source fusion. NCAM learns source-specific bias and reliability conditioned on contextual covariates, yielding an analytically tractable posterior over a latent target variable with decomposed epistemic and aleatoric uncertainty. Structural non-identifiability is resolved through sensor anchoring and variance regularization, enabling stable and interpretable posterior aggregation. To complement Bayesian uncertainty with finite-sample guarantees, we integrate locally adaptive Monte Carlo conformal prediction, producing heteroscedastic prediction intervals with coverage guarantees under exchangeability assumptions. Experiments on synthetic and real-world air-quality datasets demonstrate improved predictive accuracy and well-calibrated uncertainty compared to unsupervised baselines, including mean aggregation, probabilistic PCA, and Kalman filtering.