Score
Designs, implements, and evaluates procedures that measure, select, and optimize the bandwidth (scale) parameter of similarity kernels or spectral operators, including data‑adaptive and spectrally tuned methods. These procedures compute bandwidth measures, tune kernel scales to match intrinsic data structure, and prevent degeneracies (for example uniform membership collapse) to stabilize spectral embeddings, clustering, and related analyses.
该研究通过自适应调整高斯带宽,使核函数的谱复杂性与数据流形的内在复杂性相匹配,从而提高半监督学习任务中的分类和标签传播准确性。
This work addresses the limitations of traditional fuzzy clustering, which often yields uniform solutions due to its assumption of equal feature importance and high sensitivity to fuzzification parameters, thereby struggling to uncover complex overlapping cluster structures. To overcome these issues, the authors propose a Kernel Fuzzy Relational Clustering (KFRC) framework that innovatively integrates spectral analysis with a two-stage adaptive bandwidth selection mechanism to enable unsupervised kernel metric learning. Additionally, a novel fuzzification function is introduced, and the theoretical conditions under which clustering collapse occurs are rigorously characterized to systematically avoid uniform solutions. Experimental results demonstrate that KFRC consistently recovers intricate overlapping structures across multiple synthetic and real-world datasets, always producing pure fuzzy partitions and significantly enhancing both representational capacity and robustness.
To address the high trial-and-error cost and computational overhead in selecting dimensionality reduction techniques and tuning their hyperparameters, this paper proposes a dataset-adaptive, structural-complexity-driven optimization framework. Our method introduces, for the first time, a formal definition and quantitative measure of intrinsic data structural complexity, derived from projection error analysis and manifold geometric modeling—enabling *a priori* assessment of dimensionality reduction efficacy. This complexity metric guides automated algorithm selection (e.g., PCA, t-SNE, UMAP) and hyperparameter configuration, eliminating futile trials. Experiments across multiple benchmark datasets demonstrate that the proposed metric accurately approximates ground-truth data complexity; it reduces hyperparameter search time by 72% on average while preserving the fidelity of reduced-dimensional representations.
To address insufficient uncertainty quantification in data bias correction, this paper proposes Selective Bandwidth Kernel Density Estimation (SB-KDE), a multivariate KDE method. SB-KDE introduces a novel learnable selective KDE factor that jointly controls the scale and shape of multidimensional kernel functions. An adaptive bandwidth selection mechanism is incorporated, jointly optimizing bandwidth via two complementary criteria: Mean Conditional Squared Error (MCSE), which prioritizes correction accuracy by minimizing RMSE, and Least-Squares Cross-Validation (LSCV), which ensures overall probability density function (PDF) fidelity. Uncertainty in bias correction is explicitly modeled through the conditional PDF’s expectation and its associated confidence intervals. Extensive experiments on both synthetic and real-world datasets demonstrate that SB-KDE significantly outperforms conventional non-selective KDE methods, achieving simultaneous improvements in correction accuracy and uncertainty characterization.
Hyperspectral image (HSI) semantic segmentation suffers from spectral redundancy and high computational cost, while existing band selection methods often rely on preprocessing and are decoupled from downstream tasks. To address this, we propose Embedded Hyperspectral Band Selection (EHBS), the first end-to-end trainable framework that integrates band selection directly into semantic segmentation training. EHBS introduces a differentiable stochastic band gating mechanism for dynamic spectral filtering, employs a differentiable ℓ₀-norm regularization to precisely control band sparsity, and incorporates a parameter-free dynamic optimizer (DoG) that adaptively adjusts learning rates. Evaluated on two mainstream HSI segmentation benchmarks, EHBS achieves state-of-the-art performance—delivering higher accuracy and significantly reduced model complexity—while demonstrating strong generalization to grouped feature selection tasks.
This work addresses the challenges of bandwidth selection in kernel density estimation and its poor performance on multimodal, spiky, or discrete data by proposing a spectral method operating in the characteristic function domain. Treating bandwidth as a spectral truncation point, the approach integrates a data-driven noise-floor estimator, an adaptive Wiener tapering kernel, and a reconstruction architecture that combines a smooth basis with a band-limited residual. This framework unifies spectral decomposition, Wiener filtering, and Gaussian mixture modeling to enable automatic bandwidth selection and high-fidelity density estimation. Evaluated on the Marron–Wand benchmark suite (n = 5000), the method achieves the lowest integrated mean squared error and demonstrates superior robustness and reproducibility across six diverse real-world and synthetic datasets.
This work addresses the limitations of traditional conformal prediction, which relies on data exchangeability and struggles with non-exchangeable time series exhibiting seasonality, periodicity, or time-varying structures. The authors propose a spectral adaptive conformal prediction method that constructs weighted quantiles based on local spectral similarity and incorporates an online miscoverage rate calibration mechanism. This approach preserves finite-sample coverage guarantees while effectively capturing dynamic uncertainty in structured non-exchangeable data. By integrating spectral analysis, weighted conformal prediction, and effective sample size diagnostics, the method demonstrates superior performance over fixed spectral weighting strategies across simulations and three real-world U.S. datasets, confirming its robustness and reliability in handling complex temporal dependencies.
This work addresses the challenge that single-bandwidth kernel spectral clustering struggles to capture the complex multiscale distance structures inherent in high-dimensional data. To overcome this limitation, the authors propose a multi-kernel spectral clustering method that adaptively selects bandwidths via empirical quantiles and fuses multiple kernel functions to construct a low-rank approximate similarity matrix. They establish a row-wise ℓ₂,∞ perturbation bound for the leading eigenvectors of the normalized Laplacian, enabling fine-grained control over observation-level spectral embeddings. Under conditions of sufficient eigenvalue gap and inter-cluster separation, the proposed approach—combined with approximate K-means—achieves exact cluster recovery with high probability, significantly improving clustering performance in high-dimensional mixture models with multiscale structure.
This work addresses the inaccuracy and instability of eigenfunction estimation in kernelized diffusion maps arising from the difficulty of kernel selection. To overcome this, two adaptive kernel selection strategies are proposed: first, a variational outer-loop optimization that jointly tunes continuous kernel parameters by maximizing eigenvalues, enforcing subspace orthogonality, and applying RKHS regularization; second, an unsupervised cross-validation procedure that combines eigenvalue-based criteria with random Fourier features to enable scalable selection of both kernel families and bandwidths. Theoretically, the study establishes Lipschitz dependence of the KDM operator on kernel weights, continuity of spectral projections, and residual control, and proves exponential consistency of the cross-validation selector over finite kernel dictionaries. Experiments demonstrate that the proposed methods substantially improve the accuracy, stability, and computational scalability of eigenfunction estimation.
This study addresses the challenge of data-driven bandwidth selection for spectral density estimation near zero frequency by proposing a local-to-zero cross-validated log-likelihood criterion, denoted CVLL_c, which selects the bandwidth by summing over the frequency interval $[1, n^c]$ with $0 < c < 1$. It is shown for the first time that when $4/5 < c < 1$, the expectation of the key term in CVLL_c converges to the asymptotic mean squared error of the zero-frequency spectral estimator, thereby overcoming limitations inherent in conventional global CVLL approaches. The theoretical foundation for applying this criterion to heteroskedasticity- and autocorrelation-consistent (HAC) standard error estimation is established through Fourier frequency truncation, Taylor expansion, and asymptotic analysis.