Score
Design, build, or analyze algorithms and model components that shape and calibrate a decision boundary around "normal" (in-distribution) data so that normal samples are compactly enclosed and well separated from outliers; this includes methods that project normals onto hyperspherical surfaces, pull normals into compact clusters, maximize manifold coverage, and set or adjust thresholds for reconstruction-based baselines and other boundary geometries.
本文解决了机器学习系统中因数据聚类导致阈值设定不准确的问题,通过提出一种新的有效样本量计算方法来修正阈值设定。
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
Existing anomaly detection methods typically assume that normal data occupy a non-zero volume in the ambient space, overlooking their intrinsic geometric structure as lying on a low-dimensional manifold, which limits performance. This work proposes a novel manifold projection paradigm: learning a projection operator that maps inputs onto the manifold of normal samples and using the projection residual as the anomaly criterion. By avoiding explicit modeling of the degenerate data distribution, the approach prevents misclassifying rare yet normal instances and provides a unified explanation for both the effectiveness and failure modes of reconstruction-based methods. Extensive experiments demonstrate that the proposed framework significantly outperforms conventional boundary-learning approaches and achieves state-of-the-art results across multiple benchmarks compared to existing reconstruction-based models.
This paper addresses the challenge of anomaly detection in complex functional and distributional data—particularly under low anomaly prevalence, where detection is inherently difficult. We propose an invariant coordinate selection (ICS) framework that is coordinate-free, enabling its application beyond traditional multivariate vector data. Our method provides, for the first time, a coordinate-free definition of ICS in abstract Euclidean spaces. It leverages Bayes–Hilbert space embedding and maximum penalized likelihood spline smoothing to construct compositional spline-based density representations, thereby achieving robust dimensionality reduction and anomaly identification for distributional data. Evaluated on daily maximum temperature distribution sequences from northern Vietnam (1987–2016), our approach significantly outperforms existing methods, successfully detecting critical climate anomalies—including previously overlooked extreme events—demonstrating both theoretical generalizability and practical efficacy in real-world functional data analysis.
Accurate boundary detection in high-dimensional data remains a central challenge in unsupervised learning, particularly in the presence of non-linear structures and heterogeneous densities. In this work, we introduce Mean Curvature Boundary Points (MCBP), a novel geometric framework grounded in Geometric Machine Learning that departs from traditional density-based approaches by explicitly modeling the intrinsic curvature of the data manifold. The method relies on a discrete approximation of the shape operator, estimated from local k-nearest neighbor patches, to compute pointwise mean curvature without requiring explicit manifold parametrization. The key insight of MCBP is to use mean curvature as a principled descriptor of boundary structure: high-curvature regions naturally correspond to transitions between clusters, geometric irregularities, and low-density interfaces. This yields a unified geometric interpretation of boundary, outlier, and transition points. We further introduce an adaptive percentile-based thresholding scheme that enables multiscale boundary extraction without relying on ad hoc density parameters. Beyond detection, we propose a curvature-driven data decomposition that separates samples into smooth (low-curvature) and boundary (high-curvature) subsets, effectively acting as a non-linear geometric filtering mechanism. This representation enhances cluster separability and improves the robustness of downstream unsupervised algorithms. Extensive experiments on synthetic and real-world datasets demonstrate that MCBP consistently improves clustering performance, particularly in complex and high-dimensional scenarios. These results position MCBP as a concrete contribution to Geometric Machine Learning, highlighting the potential of curvature-aware analysis as a unifying paradigm bridging differential geometry and data-driven modeling.
This study addresses the failure of exchangeability assumptions in spatially dependent data, which leads to inadequate coverage and instability in conformal prediction intervals. To overcome this, we propose a sequential conditioning framework that calibrates residuals through sequential whitening to eliminate spatial heterogeneity. This approach is extended to large-scale networks via nearest-neighbor approximation, with theoretical bounds established for coverage loss under covariance misspecification. The method achieves distribution-free, finite-sample exact coverage and asymptotic oracle efficiency under arbitrary spatial designs. Empirical evaluations demonstrate that the proposed framework yields narrower and more stable prediction intervals. In a PM2.5 application, it effectively identifies regions at risk of undercoverage, significantly outperforming existing global and local methods.
研究通过几何特性预测异常检测性能,提出伪异常探针以改善无异常数据时的模型选择。
本文探讨了在场景优化和无分布认证中,通过确定性边界机制及随机观察边界大小来精确计算违反风险的方法。
本文提出了一种针对具有任意边界形状的边界不连续设计(BDD)的操纵测试方法,通过k-均值聚类和二项平衡测试来检测处理组与对照组的均衡性。
本文针对聚类中的不确定性问题,提出基于流形假设的聚簇方法(MBC),通过几何与样本量度结合定义不确定性区间,量化数据固有的模糊性。