build time-series forecasts

Designs, implements, and evaluates models and end-to-end pipelines that produce point and probabilistic forecasts for temporally ordered numeric metrics, including multi-horizon and hierarchical/aggregated series. This work covers model selection and training (statistical, machine‑learning or hybrid), uncertainty quantification and calibration, generation of forecasts for metrics like demand, revenue, cost, consumption and resource/budgeted usage, and the automation of reporting and deployment.

buildtime-seriesforecasts

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitation of existing forecasting systems that rely predominantly on point predictions and thus fail to adequately characterize uncertainty for informed decision-making. To overcome this, the authors propose a hybrid framework that extends point forecasts from classical models—such as Theta, exponential smoothing, and ARIMA—into probabilistic forecasts by integrating error post-processing with model-specific, horizon-dependent uncertainty scaling. The approach calibrates forecast errors using historical simulation, conformal prediction, quantile regression, and GARCH-based methods, and systematically evaluates in-sample versus out-of-sample calibration performance. Empirical results on the M4 dataset demonstrate an average 4.6% reduction in Continuous Ranked Probability Score (CRPS). In-sample calibration consistently outperforms out-of-sample calibration, particularly over longer forecast horizons, thereby validating the effectiveness and practical utility of the proposed framework.

forecast errorsin-sample calibrationpost-processing

This work addresses the challenges of probabilistic forecasting in large-scale hierarchical time series, which often suffer from high computational costs, poor scalability, or reliance on strong modeling assumptions. The authors propose e2eTD, a novel end-to-end scalable probabilistic top-down approach that directly forecasts only the top-level smoothed series—approximately 0.3% of the total—and propagates predictions downward by modeling the joint distribution of historical disaggregation proportions. Bottom-level coherent probabilistic forecasts are then obtained by summing samples from the disaggregated paths, ensuring hierarchical consistency without requiring specialized hardware. Evaluated on the M5 and Favorita benchmarks, e2eTD achieves the lowest weighted quantile loss (ranking 11th out of 892 teams in the M5 competition) and demonstrates remarkable efficiency, processing hierarchies with 40,000 and 300,000 series in just 5 and 20 minutes, respectively, on a standard laptop.

forecast coherencehierarchical forecastinglarge-scale time series

This study addresses the challenge of uncertainty quantification in aggregated time series forecasting, particularly for annual totals and year-over-year growth rates. It proposes a simulation-augmented multi-step split conformal prediction method (SA-MSCP), which generates future trajectories via block bootstrap resampling from cross-validated residuals and constructs calibrated prediction intervals using empirical quantiles. By innovatively integrating a simulation-augmentation mechanism into the multi-step split conformal prediction framework, the method significantly improves empirical coverage for both aggregate totals and their growth rates, yielding more reliable uncertainty estimates without compromising predictive accuracy.

aggregated forecastingconformal predictionprediction intervals

HoTPP Benchmark: Are We Good at the Long Horizon Events Forecasting?

Jun 20, 2024
IK
Ivan Karpukhin
🏛️ Sber AI Lab | Skoltech

Existing event prediction research primarily focuses on single-step next-event forecasting; long-horizon, multi-step joint prediction (of both time and type) remains unexplored. Method: We introduce HoTPP, the first long-horizon temporal point process (TPP) benchmark, covering critical domains such as finance and healthcare, and propose T-mAP—a theoretically grounded metric for systematically evaluating models’ long-term predictive capability. Contribution/Results: Empirical analysis reveals that mainstream marked temporal point process (MTPP) models consistently underperform simple baselines (e.g., Poisson or historical frequency) in long-horizon forecasting and suffer from mode collapse. We quantitatively identify autoregressive sampling and intensity-based loss functions as key bottlenecks limiting long-range performance. To foster reproducibility and advancement, we open-source a unified evaluation framework, implementations of state-of-the-art models, and comprehensive experimental results—paving the way from short-horizon “myopic” modeling toward genuine long-horizon event prediction.

Addressing shortcomings in existing evaluation metrics for predictionsAnalyzing mode collapse and prediction diversity in current methodsEvaluating long-horizon event forecasting in MTPP models

Latest Papers

What's happening recently
View more

This work addresses the inconsistency among multi-granularity forecasts in online hierarchical time series prediction by proposing a reconciliation method that explicitly models hierarchical relationships through a graph structure. The approach characterizes forecast residuals using a matrix normal distribution and formulates a multivariate linear regression framework, integrating ridge regression, Bayesian estimation, and shrinkage principles. An efficient online recursive inference mechanism is developed to enable adaptive forecast reconciliation and uncertainty quantification. The method is validated on a district heating load forecasting task, demonstrating its effectiveness. To support practical deployment, the authors release PyOnlineForecast, an open-source toolkit for online hierarchical forecasting.

forecast coordinationhierarchical reconciliationlinear models

Existing approaches to modeling the joint probability distribution of irregular multivariate time series often struggle to simultaneously maintain model expressiveness and marginal consistency, leading to unreliable or contradictory predictions. This work proposes CircuITS, the first method to introduce probabilistic circuits into this task, constructing a structured, interpretable joint probabilistic model that supports efficient inference. By leveraging the compositional structure of probabilistic circuits, CircuITS flexibly captures complex dependencies among variables while rigorously preserving consistency between the joint distribution and its marginals. Empirical evaluation on four real-world datasets demonstrates that CircuITS significantly outperforms state-of-the-art methods in both joint and marginal density estimation.

consistent marginalizationirregular multivariate time seriesjoint probabilistic modeling

This work addresses the instability in multi-step probabilistic forecasting, where erratic fluctuations in predictions can undermine downstream decision-making and system reliability. The authors propose a distribution-free probabilistic forecasting method that models the conditional quantile function using neural-network-parameterized regression splines. Their approach jointly optimizes predictive accuracy and temporal stability by incorporating an explicit stability regularizer into the training objective, which penalizes discrepancies between successive forecast updates. Crucially, the framework allows for region-specific weighting—enabling enhanced robustness in critical areas such as distribution tails or the central region. This is the first distribution-agnostic method to co-optimize stability and accuracy, and it demonstrates consistent efficacy across two datasets with markedly different statistical properties, significantly reducing prediction instability while preserving high forecast quality.

distribution-free forecastingforecast stabilityforecast updates

This work addresses the lack of established scaling laws and extensible foundation models in time series forecasting by proposing a unified training framework that yields a family of foundation models ranging from 4M to 2.5B parameters. Through a consistent architecture, large-scale data, standardized training protocols, and an innovative u-muP hyperparameter transfer method, the study provides the first systematic empirical validation that model performance in time series forecasting consistently improves with scale. The resulting models achieve state-of-the-art results across three major benchmarks—BOOM, GIFT-Eval, and TIME—and five checkpoints are released under the Apache 2.0 license to support further research and reproducibility.

forecasting benchmarksfoundation modelsmodel generalization

Hot Scholars

CG

Chenjuan Guo

Professor, East China Normal University
Data AnalyticsMachine Learning
JS

Jimeng Shi

Postdoctoral researcher, University of Illinois Urbana-Champaign
AI for Science & EngineeringLLMsDiffusion ModelsText Mining
HK

Harrison Katz

Forecasting Data Science, Airbnb
StatisticsForecastingBayesian Data AnalysisCompositional Time Series
FS

Flora Salim

Professor, CSE, UNSW
Machine LearningTime SeriesSpatiotemporalUbiComp
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery