When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of independence assumptions and inflated confidence in sliding-window time series classification caused by overlapping test windows. It proposes a statistical auditing framework demonstrating that subject-disjoint evaluation cannot eliminate window-level dependencies. Performance claims are mapped to aggregation rules and dependency-robust inference, distinguishing prediction counts from independent evidence. The authors introduce a variance-equivalent analysis based on effective information content, employing Bartlett-HAC inference, frozen-model recalculation, and multiple testing corrections for error calibration. This work substantially reduces Type I error rates and quantifies variance inflation under high overlap, establishing a reproducible paradigm for auditing statistical confidence in time series classification.
📝 Abstract
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.
Problem

Research questions and friction points this paper is trying to address.

sliding-window classification
statistical confidence
time-series
data dependence
inference audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sliding-window classification
Dependence-robust inference
Statistical audit
Variance-equivalent information
Time-series evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xinze Shi
Beijing University of Posts and Telecommunications, Beijing, China
Litian Zhang
Litian Zhang
Beihang University
B
Binrui Shi
Beijing University of Posts and Telecommunications, Beijing, China