Score
Designs and implements statistical procedures and algorithms to locate, date, and quantify changes in the underlying data‑generating process of ordered data (typically time series or regression contexts); this includes constructing and applying tests for the presence of one or more structural breaks, estimating breakpoints and segment‑specific parameters, and producing inference and diagnostics that are valid in the presence of single or multiple structural changes.
This work addresses the challenge of accurately attributing detected change points in multivariate time series to specific subsets of variables. The authors propose a post-hoc, nonparametric testing framework that, after an offline change point has been identified, determines whether the change occurs in one of two pre-specified coordinate blocks or in both. Built upon two-sample nonparametric hypothesis testing, the method offers rigorous theoretical guarantees for Type I error control. Empirical evaluations on both synthetic and real-world datasets demonstrate that the proposed approach achieves high attribution accuracy and strong robustness in identifying the components responsible for the change.
This paper addresses structural break detection in the cross-sectional mean of panel data. We propose a novel weighted least squares change-point test that constructs a cross-sectional mean sequence and estimates nuisance parameters to formulate a test statistic independent of bandwidth selection and long-run variance estimation; its limiting distribution is analytically tractable and robust under both weak and strong cross-sectional dependence. Theoretically, the method is proven to be consistent and asymptotically efficient. Monte Carlo simulations demonstrate excellent finite-sample size control and power. The key contribution lies in establishing, for the first time, a unified asymptotic inference framework that requires no bandwidth tuning and avoids covariance kernel estimation—enabling flexible weight design and substantially enhancing adaptability and practicality for complex cross-sectional dependence structures.
Existing methods struggle to effectively detect structural breaks in dynamical systems driven by nonlinear, nonstationary trajectories arising from external interventions or environmental shifts. This work proposes a unified framework that, for the first time, jointly models residual discrepancies and normalized parameter drifts to construct a test statistic. By integrating a multi-scale seeded narrowest-over-threshold algorithm, order-preserving segmentation, and symmetric contrastive calibration, the method achieves precise localization of structural changes in ordinary differential equation–driven systems. It simultaneously accounts for model fit and evidence of parameter variation, demonstrating robustness under both stable and divergent trajectories while attaining near-minimax localization accuracy and effective false discovery rate (FDR) control. Experiments show significant improvements over state-of-the-art approaches in detection accuracy and FDR management, with successful applications to modeling COVID-19 transmission dynamics and global temperature trends.
This paper addresses the joint estimation of structural break points—including their number, locations, and associated regression coefficients—in time series regression. We propose an ℓ₀-regularized mixed-integer quadratic programming (MIQP) framework that exactly reformulates the ℓ₀ penalty into a globally solvable mixed-integer optimization (MIO) model. Our method supports hard constraints on the number of breaks and enjoys both statistical consistency and computationally verifiable global optimality. Theoretical analysis demonstrates significantly improved break localization accuracy over mainstream approaches such as LASSO, particularly in multi-break settings. Empirical evaluations confirm the framework’s robustness and practical utility on economic and business time series. The core contribution lies in unifying exact optimization, statistical interpretability, and computational tractability for structural change modeling—thereby bridging a critical gap between theoretical rigor and scalable implementation in change-point regression.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
This study addresses the challenge of simultaneously identifying structural break points and sparse active predictors in high-dimensional regression settings. The authors propose a three-stage procedure: first, active variables are screened via Sure Independence Canonical Screening (SICS); second, potential break points are estimated using Ratio-Controlled Regression Screening (RCRS); and third, an information criterion is employed to eliminate redundant variables and spurious breaks. The method accommodates a growing number of break points with sample size and is compatible with both stationary and cointegrated sparse predictors, enabling consistent selection and estimation of true break locations and active variables. Simulation studies and empirical analyses demonstrate that the approach achieves high accuracy and robustness in detecting both relevant predictors and structural breaks.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.
This study addresses the limitation of traditional statistical process control, which emphasizes detecting historical shifts while neglecting the acceptability of the current process state. To overcome this, the authors propose a Bayesian sequential monitoring framework tailored for recoverable processes subject to parameter drift. By recursively computing the posterior probability that the process is in-control at the current time, the method shifts the monitoring focus toward real-time state assessment. The framework integrates time-to-failure modeling, Gaussian and binomial tracking, and multivariate data analysis within a unified Bayesian formulation. Demonstrated through simulation studies and an application to white wine quality data, the approach effectively identifies the current operational status of dynamic recoverable processes, significantly enhancing both monitoring accuracy and practical applicability.