The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that real-time monitoring in side-channel evaluation often leads to uncontrolled false positive rates under fixed-threshold testing. To overcome this, it proposes a betting-based testing framework utilizing SKIT-type swap e-processes for electromagnetic traces of ML-KEM. Under an explicitly conditionally symmetric null hypothesis, the method achieves anytime-valid leakage detection with strict, uniform control of the false positive probability across the entire time domain. Although the approach requires 1.7–2.4 times more data than conventional methods, it supports early stopping—consuming only 2–8% of the measurement budget—and significantly mitigates the risk of false positives induced by repeated peeking. Consequently, this work provides a reliable statistical framework for dynamic side-channel monitoring.
📝 Abstract
Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this monitored decision rule. We study anytime-valid leakage detection based on testing by betting: SKIT-type swap-pair e-processes whose false-alarm probability is controlled uniformly over time under an explicit conditional symmetry null. In matched comparisons that share the frozen witness, rows and payoff, first-crossing detection needed 1.68-2.00$\times$ the traces of a fixed-horizon randomization test at 80% detection on synthetic streams, and 1.68-2.38$\times$ on degraded recordings from an open ML-KEM electromagnetic dataset with the primary Ridge witness at $\alpha=0.05$. With the same primary witness and level, on undegraded reference and pqm4 recordings the monitored procedure stopped early: its median stopping point was 62-72 and 146-316 evaluation traces, i.e. 2-8% of a conservative 4096-trace budget. Under exact designed nulls on the recorded backgrounds, repeated-look $|t|>4.5$ screening over all 13000-20000 samples raised a false alarm in 2.7-12.9% of replicates, against 0.0-4.7% for terminal-only screening and no rejection by a sample-wise e-Bonferroni process, which in a prespecified follow-up detected natural-label associations in 4 of 4 backgrounds after 840-3288 traces. All recordings come from one device, and natural-label results are descriptive; we state the assumptions each claim requires.
Problem

Research questions and friction points this paper is trying to address.

side-channel leakage detection
anytime-valid testing
ML-KEM
electromagnetic traces
false alarm control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anytime-valid inference
Testing by betting
Side-channel leakage detection
ML-KEM
e-processes
🔎 Similar Papers
G
Georgios Feretzakis
securagen.ai
A
Alexandros Papaspyridis
securagen.ai