Score
Design and implement algorithms, modules, or analytic procedures that compute exponential moving averages to produce smoothed, exponentially weighted estimates over time-series or streaming values. This includes specifying and managing decay/forgetting parameters, online incremental update rules, bias correction and initialization, support for multivariate inputs and windows, and attention to numerical stability and efficiency to suppress transient noise and track longer-term trends.
This paper addresses time-series analysis under exponential-family distributions by proposing a unified, analytically tractable framework for filtering, prediction, and smoothing. The core method constructs a convex combination estimator of the exponentially weighted log-likelihood and the expected log-likelihood, yielding—*for the first time under standard exponential families*—exact, linear, recursive closed-form filters, predictors, and smoothers. By integrating exponential-family statistical modeling with recursive Bayesian inference, the approach achieves both computational efficiency (O(1) per-step update) and statistical consistency (asymptotic optimality). Theoretical analysis provides rigorous guarantees on consistency and convergence. Empirical evaluation on synthetic and real-world datasets demonstrates superior efficiency and robustness compared to existing approximate or non-recursive methods.
Traditional exponential moving average (EMA) fails to suppress asymptotic noise in stochastic dynamical system trajectories due to its fixed decay rate, resulting in averaged estimates lacking strong stochastic convergence. To address this, we propose the *p-EMA* method, which introduces a subharmonic decaying weight sequence—assigning gradually diminishing yet persistently non-negligible weight to recent observations. Under mild weak autocorrelation assumptions, p-EMA establishes the first exponentially weighted averaging framework with rigorous strong convergence guarantees. We prove that p-EMA achieves almost sure convergence and $L^2$ convergence, with asymptotic variance vanishing to zero—overcoming fundamental limitations of standard EMA. Furthermore, integrating p-EMA into SGD gradient estimation enhances optimization stability. Empirical results demonstrate over 40% faster error convergence under non-stationary noise compared to conventional EMA-based estimators.
This work addresses the lack of reliable uncertainty quantification in online trend estimation for nonstationary time series by proposing a general online bootstrap method applicable to trend estimators expressed as time-window-weighted sample means—such as exponential smoothing and moving averages. Built upon asymptotic theory, the method provides the first uniform-in-time coverage guarantees for trend inference under nonstationarity, enabling adaptive anomaly detection and A/B testing in streaming data settings. Empirical evaluations demonstrate that the framework achieves well-calibrated uncertainty estimates and scales effectively across diverse nonstationary scenarios, offering a practical and real-time solution for accurate uncertainty quantification in large-scale online time series analysis.
This work addresses the challenge of learning from small-sample, irregular, and non-stationary streaming multimodal time series—such as financial tick data, motion capture, and handwriting trajectories. Methodologically, it introduces a universal feature representation framework grounded in the path signature, leveraging iterative integrals and tensor algebra from control theory to extract robust, model-free, and sampling-invariant temporal features directly from raw trajectories; this effectively mitigates the exponential noise amplification induced by irregular sampling. Its key contribution is the first systematic signature-based framework explicitly designed for interpretable machine learning applications, bridging mathematical rigor with practical engineering deployment. Experiments demonstrate that the approach significantly enhances generalization and stability—even when paired with shallow classifiers—outperforming conventional time-series feature engineering. All code, reproducible experiments, and pedagogical Jupyter notebooks are publicly released.
Economic indicators are often released with significant lags, posing substantial challenges for nowcasting—particularly due to mixed-frequency observations, irregular sampling, missing data, and structural breaks (e.g., pandemic shocks), which undermine model stability. To address these issues, this paper introduces path signature regression into the nowcasting framework—the first such application. By embedding time series as continuous-time paths, the method natively accommodates asynchronous, sparse, and incomplete observations without requiring interpolation or temporal alignment. A linear regression model is then constructed in the signature feature space, and we theoretically establish its equivalence to Kalman filtering under mild conditions. Empirically, our approach achieves significantly lower forecasting errors than the New York Fed’s dynamic factor model for U.S. GDP growth nowcasting. Moreover, extended to daily-to-weekly fuel price prediction, it demonstrates strong cross-frequency generalization. This work establishes a novel paradigm for robust, real-time forecasting of heterogeneous, multi-source time series.
This study extends classical exponential smoothing to distributional time series—where observations are probability distributions on the real line—by introducing a principled and intuitive exponential smoothing framework in Wasserstein space. The smoothing parameter is consistently estimated by minimizing the Wasserstein distance, enabling effective capture and forecasting of the underlying distributional dynamics. Theoretical analysis establishes the consistency of this parameter estimator, while empirical evaluations on high-frequency financial returns and household electricity demand data demonstrate the model’s predictive accuracy and practical utility.
This work addresses the challenge of efficiently processing weighted data streams under privacy constraints in the time-decay streaming model, where traditional algorithms struggle with space complexity bottlenecks for fundamental tasks such as norm/moment estimation and frequency estimation. The paper introduces, for the first time, a systematic learning-augmented approach to this model by leveraging machine learning oracles that provide heavy-hitter information. Building on this insight, the authors design novel streaming algorithms that effectively solve norm and moment estimation, frequency estimation, and cascaded and rectangular moment estimation problems. Theoretical analysis establishes the space efficiency of the proposed methods, while experiments on both real-world and synthetic datasets demonstrate their practical efficiency and scalability, offering a new paradigm for statistical estimation over weighted data streams.
This study addresses the limitations of existing control charts for early monitoring of multi-stream binary processes, which suffer from inaccuracy due to reliance on asymptotic variance approximations and poor sensitivity to small shifts. To overcome these issues, the authors propose a Cumulative Standardized Binomial EWMA (CSB-EWMA) control chart that derives the exact time-varying variance of the EWMA statistic, enabling adaptive control limits without asymptotic assumptions and ensuring statistical rigor from the very first observation. As the first nonparametric, adaptive, and theoretically rigorous EWMA scheme tailored for multi-stream binary data, the CSB-EWMA achieves substantially improved early detection performance: under in-control average run lengths (ARL₀) of 370 or 500, it reduces the out-of-control ARL₁ to 3–7 for moderate shifts (δ = 0.2) while maintaining high stability for small shifts (coefficient of variation < 0.10).
This study addresses the efficient and numerically stable computation of unbiased sample covariance matrices in streaming and distributed settings. By unifying the algebraic, numerical, and statistical foundations of three classical algorithms—Gram, Welford, and Chan–Golub–LeVeque (CGL)—the work presents a novel theoretical framework that elucidates their differing numerical behaviors. A key innovation is the integration of conformal prediction to furnish distribution-free, finite-sample valid confidence intervals for individual elements of the covariance matrix. Experimental results demonstrate that Gram achieves the highest batch processing speed, Welford exhibits superior robustness to abrupt mean shifts, and CGL is best suited for distributed architectures; moreover, when combined with conformal prediction, all three algorithms attain nominal coverage guarantees.