🤖 AI Summary
This paper addresses the lack of rigorous statistical guarantees in novelty detection on path space. Methodologically, it introduces the first nonparametric hypothesis testing framework based on signature statistics: it constructs a smooth CVaR surrogate objective using the shuffle product identity of path signatures and leverages transport cost inequalities to control Type-I error for non-Gaussian processes—including laws of rough differential equations (RDEs). A novel SVM-based algorithm optimizes this objective, enabling computable estimates of quantiles and p-values. Contributions include: (i) the first incorporation of transport inequalities into path-space novelty detection; (ii) exact false positive rate control without Gaussianity assumptions; (iii) theoretical lower bounds on Type-II error and a general power bound under absolutely continuous alternative hypotheses. Empirical validation on synthetic anomalous diffusion data and real molecular biology datasets confirms statistical power and robustness.
📝 Abstract
We frame novelty detection on path space as a hypothesis testing problem with signature-based test statistics. Using transportation-cost inequalities of Gasteratos and Jacquier (2023), we obtain tail bounds for false positive rates that extend beyond Gaussian measures to laws of RDE solutions with smooth bounded vector fields, yielding estimates of quantiles and p-values. Exploiting the shuffle product, we derive exact formulae for smooth surrogates of conditional value-at-risk (CVaR) in terms of expected signatures, leading to new one-class SVM algorithms optimising smooth CVaR objectives. We then establish lower bounds on type-$mathrm{II}$ error for alternatives with finite first moment, giving general power bounds when the reference measure and the alternative are absolutely continuous with respect to each other. Finally, we evaluate numerically the type-$mathrm{I}$ error and statistical power of signature-based test statistic, using synthetic anomalous diffusion data and real-world molecular biology data.