Score
Self-normalized drift analysis is the ability to design and carry out probabilistic analyses that separate deterministic stabilizing drift from random fluctuations by projecting process increments and dividing them by a data-dependent normalization (for example a local variance or capacity factor). Practitioners build self-normalized martingale or projection-based arguments to bound hitting times, peak deviations, and variance-sensitive improvements in stochastic processes and randomized algorithms.
Conventional self-normalized CUSUM methods fail for mean change-point detection in locally stationary time series because they cannot disentangle long-run variance from stochastic fluctuations. Method: We propose a novel bivariate partial-sum-process-driven self-normalized CUSUM procedure that jointly constructs a mean-shift component and a variance-correction component, enabling asymptotically exact inference without estimating the long-run variance. Contribution/Results: Under local stationarity, the method achieves asymptotic α-level control and consistency. Its limiting distribution is derived via weak convergence theory and approximated by the integral of Brownian motion. Monte Carlo simulations demonstrate substantial finite-sample improvements over existing approaches. Empirical analyses on real financial and climate datasets confirm its robustness in detecting mean shifts.
This paper addresses the problem of detecting structural change points in tail-risk measures—such as Expected Shortfall (ES)—for weakly dependent financial time series. We propose a novel nonparametric change-point test based on self-normalization, the first to apply this technique to dynamic tail-risk monitoring. Crucially, it avoids estimating standard errors, thereby circumventing the challenging covariance estimation inherent in conventional methods under weak dependence. Under β-mixing conditions, we establish a functional central limit theorem for tail-risk estimators, enabling consistent multiple change-point inference. Empirically, the method robustly identifies abrupt shifts in market instability in S&P 500 index and U.S. Treasury yield data, demonstrating substantial improvements in both robustness and practical applicability for tail-risk surveillance.
This study addresses the limitations of traditional nonparametric methods in hypothesis testing for functional parameters, which often rely on bandwidth selection and consequently suffer from poor finite-sample performance. The authors propose a sample-splitting self-normalization (SS-SN) approach, extending the tuning-free self-normalization framework—previously unavailable for functional parameter testing—to a broad range of applications, including tests for marginal distributions, time reversibility, and spectral distribution change points. By integrating sample splitting with self-normalization, the method constructs a test statistic with a pivotal limiting distribution, and the paper derives its asymptotic null distribution as well as its power function under local alternatives. Numerical experiments demonstrate that SS-SN accurately controls type I error while achieving testing power comparable to or better than existing methods.
This work addresses the self-normalized concentration problem for multivariate vector-valued stochastic processes under non-sub-Gaussian (including heavy-tailed) settings—a gap left by existing scalar concentration theories, which do not extend naturally to high dimensions. We propose a unified analytical framework based on generalized sub-ψ tail conditions and establish, for the first time, time-uniform self-normalized concentration inequalities for vector-valued processes. Our analysis yields a sharp (constant-optimal) upper law of the iterated logarithm and develops multivariate empirical Bernstein inequalities. Methodologically, we integrate martingale theory, generalized cumulant generating function techniques, and vector-valued stochastic process analysis. The resulting confidence regions require no prior knowledge of variance, enabling robust and tight inference in linear regression, autoregressive modeling, and bounded-mean estimation. Empirically, our approach significantly strengthens convergence guarantees and estimation robustness under heavy-tailed noise.
This paper addresses three related hypothesis testing problems in functional time series: one-sample testing, two-sample comparison, and multiple change-point detection. We propose a unified, robust nonparametric testing framework applicable to arbitrary sampling designs—from sparse to dense—and to contaminated observations with measurement error, without requiring estimation of nuisance parameters such as long-run covariance or error variance. Leveraging B-spline-based functional estimation and self-normalization, combined with high-dimensional Gaussian approximation theory, our approach overcomes the challenges of insufficient tightness and joint weak convergence of nonparametric test statistics, and rigorously characterizes the sparse–dense phase transition boundary. The theoretical guarantees hold under a diverging-dimension regime. Extensive numerical experiments and real-data applications—including AU.SHF implied volatility and urban traffic flow—demonstrate excellent finite-sample performance and practical utility.
研究解决了机器学习系统在线修正时监控器误报问题,通过Huber-style方法减少误报并提高有效性。
本文研究了基于顺序周期图的谱密度积分估计量的自归一化方法,解决了线性和非线性泛函中未知谱量的问题。
本文针对固定宽度顺序停止规则在无限方差或长程依赖情况下的失效问题,提出了一种基于联合泛函极限定理的方法,并引入了序列子抽样过程来解决。
This work proposes a unified and interpretable Gaussian Mixture Model (GMM) framework for simultaneously performing anomaly detection and distribution drift identification in data streams. The approach automatically selects the number of mixture components via the Bayesian Information Criterion, initializes components using k-means, and sets anomaly thresholds by combining negative log-likelihood with Extreme Value Theory. Drift intensity is quantified by the proportion of data not covered by existing components. Each alert is accompanied by an intuitive explanation: anomalies correspond to points significantly deviating from known mechanisms (3–10σ from the nearest cluster), while drift reflects an increasing proportion of observations governed by previously unseen mechanisms. Evaluated on seven benchmarks, the method matches state-of-the-art performance in anomaly detection and achieves drift detection accuracy comparable to Maximum Mean Discrepancy (MMD) when novel mechanisms emerge, with all outputs remaining inherently interpretable.
This work proposes a self-balancing sequential sampling method tailored for applications such as auditing, scheduling, and sampling, where both unpredictability and distributional convergence are critical. By adaptively adjusting selection probabilities, the method achieves optimal $O(n^{-1})$ convergence of the empirical distribution to the target distribution while preserving maximal sample unpredictability—substantially improving upon the $O(n^{-1/2})$ rate of conventional i.i.d. sampling. Theoretical analysis reveals that this approach uniquely solves an entropy-regularized optimization problem and exhibits a diffusion limit linked to the Ornstein-Uhlenbeck process. Empirical results confirm its effectiveness in minimizing repeated selections and coverage gaps, all while rigorously maintaining control over unpredictability.