Score
Design and implement multi-stage statistical filtering pipelines that sequentially apply different tests or thresholds to identify, remove, or downweight anomalous or noisy items while preserving legitimate variance; this includes choosing stage ordering, aggregation rules, and per-stage criteria. Build analyses and validations of the cascade’s behavior—e.g., its impact on optimization stability, early oscillations, and false-positive/false-negative tradeoffs—and tune stage thresholds to minimize harm to benign variability.
This work addresses the problem of multiple hypothesis testing for multivariate Gaussian means under arbitrary covariance dependence structures. The authors propose a novel approach that integrates maximum residual descent (MRD) with a multi-stage calibration scheme. By introducing a new representation of residual statistics based on a single active precision matrix, the method achieves covariance-adaptive residualization, substantially reducing computational complexity. It replaces model-dependent thresholds with a simple multi-stage calibration rule, combining generalized stepwise critical values and precision matrix reconstruction techniques. The proposed procedure significantly lowers the normalized misclassification risk across diverse dependence structures, achieving error discovery rate control close to the nominal level, extremely low missed detection rates, near-perfect statistical power, and accurate estimation of the number of true signals.
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.
Traditional static thresholds fail in nonstationary time-series anomaly detection due to concept drift, regime shifts, and multiscale dynamics. To address this, we propose two adaptive thresholding frameworks. Our method integrates statistical online learning with sequential segmentation, employing a context-sensitive dynamic threshold adjustment strategy. Key contributions include: (1) a piecewise confidence sequence that enables online modeling of local statistical properties; and (2) a multiscale adaptive confidence segment mechanism that rigorously controls the false positive rate under distributional evolution—backed by theoretical guarantees and offering full interpretability. Evaluated on the Wafer Manufacturing benchmark dataset, our approach achieves significantly higher F1 scores than percentile- and rolling-percentile-based baselines, demonstrating superior robustness and practical efficacy in real-world nonstationary settings.
In sequential anytime testing, combining e-process evidence across filtrations poses a fundamental challenge: an e-process valid under a coarse filtration may fail to retain validity under a finer one, even for the same null hypothesis. Method: We propose the “adjusted combination” framework, introducing the class of *adjuster* functions—characterizing necessary and sufficient conditions for cross-filtration evidence boosting. We prove that an adjuster is essential to restore anytime validity under the original filtration and quantify its logarithmic cost. Our approach unifies e-process theory, generalized test martingales, and filtration coarsening/refinement techniques. Results: We establish a complete characterization theorem for adjusters and validate the framework on real financial data for randomness testing. The work provides novel theoretical tools and practical methodology for sequential independence testing and predictive model evaluation.
To address the limitations of existing scan statistics for sparse anomaly detection under multi-source heterogeneous coordinate systems—namely, their reliance on strong distributional assumptions and difficulty in calibration—this paper proposes a rank-based high-criticism scoring method. The approach leverages only the relative ordering among independent observations, avoiding parametric modeling entirely, and introduces a novel rank-based high-criticism framework to nonparametrically characterize detectability conditions. We theoretically establish that detection power is uniquely determined by the probability that an anomalous observation exceeds a typical one, and prove asymptotic optimality under both exponential families and convolution models. The method achieves strong robustness and theoretical interpretability: it successfully identifies process anomalies in pharmaceutical quality control data, and simulations demonstrate performance approaching that of the oracle test.
This study addresses the limitations of existing control charts for early monitoring of multi-stream binary processes, which suffer from inaccuracy due to reliance on asymptotic variance approximations and poor sensitivity to small shifts. To overcome these issues, the authors propose a Cumulative Standardized Binomial EWMA (CSB-EWMA) control chart that derives the exact time-varying variance of the EWMA statistic, enabling adaptive control limits without asymptotic assumptions and ensuring statistical rigor from the very first observation. As the first nonparametric, adaptive, and theoretically rigorous EWMA scheme tailored for multi-stream binary data, the CSB-EWMA achieves substantially improved early detection performance: under in-control average run lengths (ARL₀) of 370 or 500, it reduces the out-of-control ARL₁ to 3–7 for moderate shifts (δ = 0.2) while maintaining high stability for small shifts (coefficient of variation < 0.10).
This study addresses the critical yet underexamined role of data filtering in clinical machine learning, which alters statistical structures and directly impacts task complexity and model performance. Despite these effects, existing research frequently treats filtering as routine preprocessing with insufficient transparency. This work reconceptualizes data filtering as a core component of the scientific method, advocating its integration into the broader research paradigm rather than its treatment as a mere technical step. To this end, we develop a transparent and interpretable clinical data preprocessing pipeline and release the corresponding code as open source. Our analysis elucidates the mechanisms through which filtering decisions critically influence data distributions and downstream model efficacy. Ultimately, this research provides a novel framework for enhancing methodological rigor and reproducibility in clinical artificial intelligence studies.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.
This study addresses the sensitivity to noise schedules and the lack of theoretical justification in multi-step sampling for consistency models. By analyzing the composition of noising and denoising operators, it establishes a non-asymptotic convergence theory under explicitly verifiable stability assumptions. Methodologically, the analysis decouples initialization error contraction from approximation error accumulation, revealing that large early-stage noise drives contraction while small late-stage noise controls residual bias, with explicit constants derived for strongly log-concave targets. Experiments confirm that the theoretically predicted contraction and approximation profiles are reliably measurable. This work provides both rigorous theoretical guidance and a practical framework for designing multi-step consistency samplers.
本文提出了一种非递归滤波器,通过最小化折现的凸组合损失来处理时间序列,并应用于金融资产交易数据中以消除微观结构噪声对波动率估计的影响。