The Nixtlaverse: An Open-Source Ecosystem for Forecasting

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the incompatibility of data interfaces among statistical, machine learning, and neural network model families in forecasting software, which necessitates redundant development efforts. To overcome this, we propose a standardized integration design paradigm and construct an open-source Python ecosystem built upon shared data contracts. By unifying long-format panel data and keyed forecast outputs while preserving model-specific implementations, the framework facilitates cross-engine hybrid modeling and rolling-origin evaluation. Furthermore, by integrating hierarchical weighting metrics with sparse reconciliation algorithms, we validate the multi-engine forecast reconciliation performance and memory efficiency on the M5 competition dataset. This approach effectively eliminates framework barriers and has achieved broad academic reuse and community adoption.
πŸ“ Abstract
Large forecasting applications often combine statistical, machine-learning, and neural models. These families solve the same problem but differ in fitted state, training procedures, and how they parallelize work. Forecasting software must therefore either hide these differences behind a single estimator interface, or keep the families in separate packages, forcing users to rewrite data preparation and evaluation for every package. We present the Nixtlaverse, an ecosystem of open-source Python libraries for time series forecasting, as a case study of a third design: all libraries share the same long-format panel data and keyed forecast outputs, while every model family keeps its own specialized implementation. We demonstrate this design through three use cases on the public M5 competition data. First, we evaluate statistical, machine-learning, and neural models, and an external engine from a separate ecosystem, in a single rolling-origin evaluation with per-series and hierarchy-weighted metrics. Second, we profile runtime and peak memory from 100 to 30,490 series and locate each family's bottleneck: statistical fitting scales approximately linearly in the number of series, feature construction dominates machine-learning memory, and neural training time is nearly independent of panel size under a fixed training budget. Third, we reconcile the forecasts of multiple engines, including the external one, over all 42,840 series of the M5 hierarchy, with sparse reconciliation where dense implementations exhausted memory. These use cases establish the costs, boundaries, and utility of shared data and output contracts. The Nixtlaverse has seen substantial public distribution, scholarly reuse, and adoption through other forecasting frameworks, and is released under permissive open-source licenses with public datasets, reproducible examples, and verifiable benchmark artifacts.
Problem

Research questions and friction points this paper is trying to address.

time series forecasting
ecosystem design
model interoperability
data standardization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time Series Forecasting
Open-Source Ecosystem
Sparse Reconciliation
Long-format Panel Data
Scalability Profiling
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Olivier Sprangers
Olivier Sprangers
Nixtla, San Francisco, CA, United States
M
Max Mergenthaler Canseco
Nixtla, San Francisco, CA, United States
M
Marco Peixeiro
Nixtla, San Francisco, CA, United States
S
Saul Caballero Ramirez
Nixtla, San Francisco, CA, United States
M
Mariana Menchero GarcΓ­a
Nixtla, San Francisco, CA, United States
J
Jing-Qiang Goh
Nixtla, San Francisco, CA, United States
H
Han Wang
Nixtla, San Francisco, CA, United States
N
Nikhil Gupta
Nixtla, San Francisco, CA, United States
R
Rogelio Melo
Nixtla, San Francisco, CA, United States
S
Senbong Gee
Nixtla, San Francisco, CA, United States
C
Cristian Challu
Nixtla, San Francisco, CA, United States