spectral similarity weighting

Designs and implements weighting schemes based on spectral (frequency-domain) similarity metrics, including methods to compute local spectral similarity scores and perform spectral comparisons. Uses those similarity-derived weights to prioritize or reweight calibration residuals and examples—e.g., for forming weighted conformal quantiles or selecting calibration examples most relevant to a given test point.

spectralsimilarityweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional conformal prediction, which relies on data exchangeability and struggles with non-exchangeable time series exhibiting seasonality, periodicity, or time-varying structures. The authors propose a spectral adaptive conformal prediction method that constructs weighted quantiles based on local spectral similarity and incorporates an online miscoverage rate calibration mechanism. This approach preserves finite-sample coverage guarantees while effectively capturing dynamic uncertainty in structured non-exchangeable data. By integrating spectral analysis, weighted conformal prediction, and effective sample size diagnostics, the method demonstrates superior performance over fixed spectral weighting strategies across simulations and three real-world U.S. datasets, confirming its robustness and reliability in handling complex temporal dependencies.

conformal predictionnon-exchangeable dataprediction intervals

This work addresses the challenge of conformal prediction under distributional shifts and structural changes in non-stationary streaming data, where traditional methods relying on exchangeability assumptions often fail. To overcome this limitation, the authors propose the DASC framework, which uniquely integrates local spectral similarity with an optimal transport–based drift score to dynamically weight residuals and adaptively adjust both the calibration pool and the target miscoverage level. This approach enables adaptive uncertainty quantification for non-exchangeable streaming data and introduces an online effective sample size diagnostic to assess predictive robustness. Experimental results on synthetic and real-world datasets—including electricity load, weather, and financial time series—demonstrate that DASC consistently achieves nominal or conservative coverage while reducing average prediction interval width by 28%–42% compared to existing methods.

conformal predictiondistributional driftnon-exchangeable data

WeSpeR: Population spectrum retrieval and spectral density estimation of weighted sample covariance

Oct 18, 2024
BO
Benoit Oriol
🏛️ Université Paris-Dauphine | Société Générale Corporate and Investment Banking

This paper addresses the challenge of estimating the spectral distribution of high-dimensional weighted sample covariance matrices. We propose WeSpeR, a novel algorithm that (i) rigorously establishes, for the first time, that the limiting spectral distribution admits a regular continuous density; (ii) introduces the first end-to-end joint framework simultaneously performing continuous spectral density estimation, precise support set bounding, and population-level spectral inversion; and (iii) integrates asymptotic random matrix theory analysis, numerical Stieltjes transform inversion, adaptive support localization, and grid optimization. Experiments demonstrate that WeSpeR significantly improves spectral density estimation accuracy, faithfully recovers the true population eigenvalue distribution, and exhibits both theoretical convergence guarantees and strong robustness against model misspecification and noise. By unifying theoretical analysis and practical estimation, WeSpeR establishes a new paradigm for high-dimensional weighted covariance modeling.

Computing non-linear shrinkage for weighted covarianceDeriving efficient algorithm for large-scale dataSpeeding up high-dimensional covariance estimation

Variance-Adjusted Cosine Distance as Similarity Metric

Feb 04, 2025
SS
Satyajeet Sahoo
🏛️ IIT Kharagpur

Traditional cosine similarity assumes data reside in a Euclidean space and ignores the variance and covariance of random variables, leading to inaccurate similarity estimates when features exhibit correlation and heteroscedasticity. To address this, we propose Variance–Covariance-Corrected Cosine distance (VC-Cosine), the first method to explicitly incorporate second-order statistical structure—i.e., feature variances and covariances—into cosine distance computation. VC-Cosine whitens the feature space via the empirical covariance matrix, thereby adaptively reweighting dimensions according to their correlations and variabilities, and aligning the inner-product geometry with the true underlying data distribution rather than relying on isotropic assumptions. Experiments on the Wisconsin Breast Cancer dataset demonstrate that, when integrated with a k-nearest neighbors classifier, VC-Cosine achieves 100% test accuracy—substantially outperforming standard cosine similarity and other state-of-the-art similarity measures.

Impact of variance and correlation on cosine distance accuracyLimitations of traditional cosine similarity in random variable spaceProposal of a variance-adjusted cosine similarity for improved performance

This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.

Developing visualization for calibration and generalization errorProving relationship between full and confidence calibration errorReassessing calibration metrics in machine learning

Latest Papers

What's happening recently
View more

This study addresses the performance bottleneck in personalized prediction caused by inaccurate patient similarity measures. We propose a supervised weighted cosine similarity method based on relaxed adaptive group Lasso, which leverages supervised learning to adaptively estimate feature weights, thereby overcoming the limitations of traditional unsupervised similarity computations. This approach significantly enhances personalized modeling for binary classification data. Experiments on ICU datasets demonstrate that the proposed model effectively improves discriminative power and overall predictive accuracy, as evidenced by improved Brier scores. Although calibration exhibits a marginal decline, this work establishes a novel paradigm for precise clinical prediction by integrating supervision into similarity metric learning.

Binary Response DataPatient SimilarityPersonalized Predictive Modelling

This work addresses the unclear failure mechanisms of classical spectral descriptors—such as the Heat Kernel Signature (HKS) and Wave Kernel Signature (WKS)—in non-rigid 3D shape retrieval, particularly the lack of systematic analysis regarding the contribution of different scale components. The authors propose a frequency-scale saliency framework that, for the first time, establishes a quantitative relationship between scale intervals in spectral descriptors and retrieval performance, revealing that short-scale components predominantly drive performance while long-scale components are detrimental. By characterizing category-level scale dependencies through spectral category fingerprints, they introduce a saliency-weighted strategy to optimize retrieval. Evaluated on the SHREC'11 benchmark, the method improves mean average precision (mAP) by 0.156 for challenging categories, with gains validated as stable and reliable through cross-validation and randomized controlled experiments.

3D shape retrievalnon-rigid shapesretrieval failure

This work proposes a novel data quality diagnostic method for detecting label noise in training datasets by analyzing the tail index (α) of the eigenvalue distribution of weight matrices in neural network bottleneck layers. The study establishes, for the first time, a strong correlation between the tail index and the level of label noise, positioning α as a dedicated metric for data quality assessment rather than generalization analysis. Leveraging random matrix theory, spectral analysis, and tail index estimation, the approach achieves an R² of 0.984 in predicting test accuracy across 21 noise levels—significantly outperforming conventional methods. Furthermore, on the CIFAR-10N benchmark, it successfully identifies 9% of artificially introduced annotation errors with only a 3% false positive rate.

data qualityeigenvalue tail indexlabel noise

This study addresses the problem of quantifying the goodness-of-fit of moving average MA(q) models to the spectral density of stationary processes by proposing a spectral-domain coefficient of determination. This coefficient extends, for the first time, the classical notion of the coefficient of determination into the framework of spectral analysis to measure how closely an MA(q) model approximates the true spectral density. Constructed via periodogram-based estimation, the proposed coefficient is shown to possess asymptotic normality under rigorous derivation, enabling the development of both a model order selection criterion and a goodness-of-fit test specifically tailored for MA(q) models. The approach adaptively identifies the minimal order q that achieves a pre-specified accuracy level, offering a method that is theoretically sound and practically useful.

coefficient of determinationMA(q) modelmodel order selection

This study addresses the exponential growth in calibration costs caused by likelihood ratios in weighted conformal prediction under covariate shift. To mitigate this issue, we propose a sketched calibration method that computes weights via compressed covariates, thereby reducing sensitivity to irrelevant shift directions and effectively controlling calibration costs. Theoretically, we prove that the compression operation does not increase shift-dependent calibration overhead and quantify response distribution leakage as a key metric for coverage guarantees. Empirically, our experiments demonstrate that the proposed approach reduces the proportion of infinite prediction sets from 37.8% to 8.0%, significantly improving predictive efficiency while maintaining the target coverage rate.

calibration costchi-square divergenceconformal prediction

Hot Scholars

SP

Shrideep Pallickara

Professor of Computer Science, Colorado State University
Distributed SystemsCyberinfrastructureSpatial Data ScienceClouds
TB

Tanjim Bin Faruk

Ph.D. Student, Colorado State University
Machine LearningDeep LearningVision TransformerRemote Sensing
SL

Sangmi Lee Pallickara

Professor of Computer Science, Colorado State University
Big DataScientific Data ManagementData storageAnalytics
WD

Wenjie Du

Univeristy of science and technology of China