Score
Designs and implements evaluation metrics, diagnostic procedures, and scoring pipelines that quantify forecasting performance across horizons for time-series and sequence forecasts, covering both point and probabilistic outputs. Builds cumulative and goal-dependent metrics and visualizations to compute horizon-wise and aggregated errors, detect masking by pointwise scores, diagnose recursive feedback–driven error growth, and assess calibration, sharpness, and other verification properties.
This work addresses the common oversight in traditional time series forecasting, where evaluation focuses predominantly on point accuracy while neglecting temporal consistency—i.e., the stability of predictions for the same future timestamp when issued from different origins. To remedy this, the authors propose the Accuracy and Consistency Score (AC Score), a differentiable, user-weighted evaluation metric that explicitly incorporates stability into the assessment of multi-step probabilistic forecasts and serves as an end-to-end training objective. By optimizing the AC Score within a seasonal ARIMA framework, experiments on the M4 Hourly dataset demonstrate that the approach achieves comparable or superior point forecast accuracy while reducing prediction volatility for identical target timestamps by 75%.
Current time-series forecasting evaluation suffers from a fundamental flaw: conventional metrics conflate model performance with intrinsic data predictability, yielding biased assessments. To address this, we propose a spectral-coherence-based predictability-aligned evaluation framework. It introduces the Spectral Coherence Predictability (SCP) score and the Linear Utilization Ratio (LUR) diagnostic tool—revealing, for the first time, the phenomenon of “predictability drift.” Our method integrates fast Fourier transform (O(N log N)), frequency-resolved analysis, and linear system modeling to quantify task-inherent difficulty and assess how efficiently models exploit available information. Experiments demonstrate complementary strengths: complex models excel on low-SCP subtasks, while linear models dominate high-SCP regimes. This framework shifts evaluation from static ranking to predictability-aware dynamic diagnosis, enabling principled model selection and targeted improvement. (149 words)
This study addresses the limitation of existing forecasting systems that rely predominantly on point predictions and thus fail to adequately characterize uncertainty for informed decision-making. To overcome this, the authors propose a hybrid framework that extends point forecasts from classical models—such as Theta, exponential smoothing, and ARIMA—into probabilistic forecasts by integrating error post-processing with model-specific, horizon-dependent uncertainty scaling. The approach calibrates forecast errors using historical simulation, conformal prediction, quantile regression, and GARCH-based methods, and systematically evaluates in-sample versus out-of-sample calibration performance. Empirical results on the M4 dataset demonstrate an average 4.6% reduction in Continuous Ranked Probability Score (CRPS). In-sample calibration consistently outperforms out-of-sample calibration, particularly over longer forecast horizons, thereby validating the effectiveness and practical utility of the proposed framework.
This paper addresses the lack of a quantitative predictability metric prior to time-series modeling. We propose a model-free pre-assessment framework that jointly quantifies intrinsic predictability via two complementary indicators: the spectral regularity score (capturing frequency-domain structure) and the maximum Lyapunov exponent (measuring chaotic dynamics). To our knowledge, this is the first work integrating linear spectral analysis with nonlinear dynamical systems theory for predictability assessment—shifting from conventional post-hoc, model-dependent evaluation to a theory-driven, priori feasibility criterion. Empirical validation on synthetic benchmarks and the real-world M5 retail sales dataset demonstrates statistically significant negative correlations (p < 0.01) between both metrics and prediction errors of diverse models—including ARIMA, LSTM, and N-BEATS—confirming their effectiveness in identifying highly predictable series and enabling principled allocation of forecasting resources.
Existing event prediction research primarily focuses on single-step next-event forecasting; long-horizon, multi-step joint prediction (of both time and type) remains unexplored. Method: We introduce HoTPP, the first long-horizon temporal point process (TPP) benchmark, covering critical domains such as finance and healthcare, and propose T-mAP—a theoretically grounded metric for systematically evaluating models’ long-term predictive capability. Contribution/Results: Empirical analysis reveals that mainstream marked temporal point process (MTPP) models consistently underperform simple baselines (e.g., Poisson or historical frequency) in long-horizon forecasting and suffer from mode collapse. We quantitatively identify autoregressive sampling and intensity-based loss functions as key bottlenecks limiting long-range performance. To foster reproducibility and advancement, we open-source a unified evaluation framework, implementations of state-of-the-art models, and comprehensive experimental results—paving the way from short-horizon “myopic” modeling toward genuine long-horizon event prediction.
Accurately estimating the maximal Lyapunov exponent from scalar time series is highly challenging in the absence of governing equations, tangent-space dynamics, and full state information. This work proposes the FEG-Pro framework, which leverages autocorrelation-guided sparse embedding and distance-weighted k-nearest neighbor multi-step prediction to analyze the finite-horizon slope of the logarithmic growth of geometric mean prediction errors. For the first time, error growth is treated as a structured profile, incorporating multidimensional diagnostic features such as curvature, residual roughness, monotonicity, and entropy of the error distribution. The method demonstrates strong performance on scalar observations from chaotic maps, the Mackey–Glass system, and the Lorenz-63 attractor, achieving close agreement with true Lyapunov exponents in near-linear regimes and retaining interpretable characteristics even under short-data conditions, thereby offering a novel paradigm for instability rate estimation and machine learning.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This work addresses the challenge of evaluating model outputs in scenarios where ground-truth outcomes are delayed, censored, or private, rendering conventional code-based deterministic evaluation methods ineffective for immediate validation. The authors propose RouteCast, a novel framework that enables autonomous generation of auditable provisional prediction scores through typed, staged route modeling, reference-based analogy, and deterministic transformations, thereby supporting traceable and decomposable assessment of strategic pathways. Evaluated on 21 retrospective cases, RouteCast achieves an AUC of 0.756—significantly outperforming blind-evaluated large language models (AUC = 0.678) and performing comparably to identity-revealed LLMs (AUC = 0.761)—demonstrating its effectiveness and feasibility in settings with delayed ground truth.
This work addresses the performance degradation of time series models in post-training quantization (PTQ), which arises from error propagation and amplification during quantization—particularly challenging in calibration-free or black-box settings where module sensitivity is hard to assess. To tackle this, the paper introduces discrete-time dynamical systems theory into quantization analysis for the first time. By modeling the inference process as a dynamical system, it proposes TQS, a quantizer-agnostic, prior-based sensitivity metric derived from trajectory sensitivity analysis, enabling calibration-free mixed-precision quantization budget allocation. The resulting TQS-PTQ framework significantly outperforms existing PTQ methods without relying on calibration data or second-order approximations, facilitating efficient low-bit deployment.
This work addresses the limitations of traditional time series methods, which are constrained by fixed forecasting horizons and struggle to support contextual reasoning, tool invocation, and structured decision-making in real-world scenarios. The authors propose AION, a framework that formalizes time series tasks as a triplet of task specification, workspace, and validation interface, integrating six core modules—agents, skills, rules, memory, evaluation, and protocols—to emphasize temporal grounding, knowledge-guided reasoning, and reliability assurance. By incorporating process traceability, multi-level auditing, and post-hoc experimental analysis, AION overcomes the constraints of static evaluation paradigms. In a Kaggle store sales forecasting case study, AION substantially outperforms direct modeling approaches, generating richer reasoning traces, intermediate artifacts, and audit steps, thereby demonstrating its effectiveness and superiority in handling complex, real-world time series tasks.