CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of cross-domain heterogeneous calibration biases in probabilistic event prediction using large language models. To this end, it proposes an end-to-end calibration framework based on causal temporal hypergraphs. The method decomposes prediction into three stages—evidence weighting, aggregation, and fusion—introducing a time-decay modulation mechanism governed by causal topological distance, direction-aware Noisy-OR logical aggregation with deduplication, and adaptive source fusion. These innovations overcome the limitations of conventional post-hoc calibration. Experimental results demonstrate that the proposed approach significantly reduces expected calibration error and Brier scores while improving predictive accuracy across cross-domain benchmarks, outperforming existing state-of-the-art methods.
📝 Abstract
Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at https://github.com/QwenQKing/Chain.
Problem

Research questions and friction points this paper is trying to address.

LLM forecasting
probability calibration
calibration bias
causal-temporal hypergraph
trustworthiness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal-Temporal Hypergraph
Probability Calibration
Large Language Models
Event Forecasting
Noisy-OR Aggregation