non-parametric tail estimation

Designs and implements non‑parametric estimators of distribution tails and extreme-value characteristics, constructing data‑driven extreme value distributions (DDEVd) and estimating long‑return levels beyond the sample without imposing parametric tail forms. This work includes reconstructing the base distribution with kernel smoothing, selecting optimal kernel bandwidths, and aggregating observations metastatistically to improve tail inference.

non-parametrictailestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenging problem of bandwidth selection in extreme value distribution estimation by proposing a data-driven kernel-based estimator. For the first time, a complete analytical expression of the mean integrated squared error (MISE) for this estimator is derived, establishing a rigorous theoretical foundation and stability conditions for optimal bandwidth selection. By integrating extreme value theory with kernel density estimation, the proposed method enables adaptive and optimal bandwidth choice, significantly enhancing the accuracy of extreme value distribution estimation while ensuring numerical stability.

bandwidth selectiondata driven estimationextreme value distribution

This study addresses the limitations of classical extreme value theory, which relies on asymptotic assumptions and struggles to accurately characterize extreme risks under data sparsity or non-stationarity. The authors propose DDEVD, a nonparametric approach that reconstructs the underlying distribution via kernel density estimation and incorporates superstatistical aggregation, thereby avoiding parametric assumptions about the tail. They derive theoretical conditions for optimal bandwidth selection and stability, establishing the explicit criterion $m < C n^{1+\gamma/2}$ that links extrapolation reliability directly to the extreme value index for the first time. Applied to sub-hourly Alpine precipitation data, DDEVD reliably reproduces 100-year return events using only ten years of observations (calibration ratio: 0.96). In metal micrograph analysis, it improves estimation accuracy for safety-critical grain-size extremes by 58% over the log-normal model.

data-sparseExtreme Value Theorynon-stationary

Perturbation-based Inference for Extreme Value Index

Dec 09, 2025
YT
Yiwei Tang
🏛️ Fudan University | Rice University

To address unreliable extreme value index (EVI) inference caused by scarcity of tail data, this paper proposes a perturbation-based synthetic exceedance generation method: controlled noise is injected into exceedances above a high threshold, followed by generalized Pareto distribution (GPD) modeling and construction of a consistent pivotal statistic. Innovatively, the perturbation mechanism is integrated with differential privacy guarantees; when GPD approximation error is substantial, a refined perturbation strategy is further introduced to enhance robustness. Experiments demonstrate that the proposed method significantly outperforms existing EVI inference approaches in terms of confidence interval coverage, width control, and resilience to model misspecification. It establishes a novel paradigm for reliable extreme-value analysis under sparse tail regimes.

Constructs confidence intervals using perturbed exceedancesEnsures differential privacy in tail inferenceEstimates extreme value index with data scarcity

This work addresses the challenge of traditional extrapolation methods failing in extreme regions due to data scarcity in the tails—a common issue in machine learning. To overcome this limitation, the authors propose a unified extreme-value extrapolation framework that integrates extreme value theory with statistical learning. Built upon asymptotic representations of univariate and multivariate tail distributions, the framework combines extreme value index estimation, tail distribution modeling, and dependence structure analysis. It is applicable to both supervised and unsupervised settings and accommodates both asymptotically dependent and independent data. Empirical evaluations demonstrate that the method substantially outperforms existing approaches in tasks such as extreme quantile regression, anomaly detection, and generative AI, yielding improved accuracy and robustness in predicting rare and extreme events.

anomaly detectionextrapolationextreme value theory

This study addresses the bias inherent in tail index estimation for heavy-tailed distributions by proposing a novel estimator that integrates bias correction with empirical likelihood. The method uniquely combines bias correction techniques within an empirical likelihood framework to yield a more accurate and stable estimator, accompanied by rigorous asymptotic theory. Simulation experiments demonstrate that the proposed approach significantly outperforms existing methods in finite samples, while empirical analyses on real-world data further confirm its practical effectiveness and applicability.

bias correctionempirical likelihoodextreme value analysis

Latest Papers

What's happening recently
View more

This work addresses the challenge of simulating multivariate extreme events and estimating rare-event probabilities under both heavy-tailed and light-tailed distributions by proposing Self-Similar Generative Estimation (SS-GEN). The method uniquely integrates the asymptotic tail structure from extreme value theory into a deep generative modeling framework, leveraging self-similar decomposition to decouple the tail distribution into an explicit radial component and a nonparametric angular component. This transformation recasts tail modeling as a standard generative task on a compact domain, eliminating the need for specialized architectures or parametric tail assumptions. Theoretical analysis shows that SS-GEN achieves vanishing relative error for regularly varying distributions and vanishing log-relative error for Weibull-type tails. Empirical results confirm its ability to accurately generate representative extreme samples and reliably estimate rare-event probabilities.

extreme value theorygenerative modelingmultivariate extremes

This study addresses the limitation of the classical Weitzman overlap coefficient, which is restricted to pairwise probability distributions, by extending it for the first time to the case of k (k ≥ 2) independent distributions. The authors reformulate the generalized overlap coefficient as the expectation of a specific function and develop a nonparametric estimator that circumvents the need for closed-form density expressions by integrating kernel density estimation with the method of moments. The resulting framework offers a flexible and broadly applicable measure of overlap among multiple distributions. Extensive Monte Carlo simulations demonstrate that the proposed estimator exhibits robust performance across diverse distributional settings, combining strong theoretical validity with practical utility, thereby providing a valuable tool for multivariate overlap analysis.

distribution overlapk distributionskernel density estimation

This work proposes a semiparametric density estimation approach that addresses the inefficiency of traditional kernel density estimation in high-dimensional settings or when the underlying distribution deviates substantially from assumed parametric forms. The method multiplies a parametric initial guess—such as a normal distribution—by a nonparametric kernel-based correction factor, thereby preserving robustness against departures from the parametric family while significantly enhancing local estimation accuracy. By integrating parametric priors with nonparametric adjustments, the framework incorporates a tailored bandwidth selection strategy and naturally extends to nonparametric regression. Theoretical analysis and extensive simulations demonstrate that, even when the true density markedly departs from normality, the proposed estimator consistently outperforms conventional kernel density estimators across various Gaussian mixture models, particularly excelling in high-dimensional scenarios.

density estimationkernel estimatornonparametric

This work addresses the lack of robustness in classical moment-based methods when applied to heavy-tailed or contaminated data. Building upon a distribution–kernel duality representation, the authors propose a weak moment estimation framework that enables parameter inference through weak moment matching and regularized density reconstruction—without requiring explicit density estimation. The approach inherently satisfies Hampel’s local robustness criterion, and its influence function admits a closed-form expression. By employing Schwartz kernels together with Tikhonov regularization, the method circumvents the need for Huber-style tuning or post hoc truncation. Empirical results demonstrate that weak moment estimation substantially outperforms conventional robust techniques in Cauchy and t₃ models, and achieves optimal parametric convergence rates in bivariate t₃ scale estimation—even in settings where maximum likelihood estimation fails.

contaminated dataheavy-tailed distributionsrobust estimation

This study addresses the challenge of simultaneously modeling time-varying tail dependence structures and heteroskedastic marginal distributions in independent but non-identically distributed random vectors. The authors propose a nonparametric method to estimate the integrated tail copula and establish its asymptotic theory. Notably, this work is the first to investigate time-varying tail dependence under heteroskedastic margins, demonstrating that heteroskedasticity does not affect the limiting distribution of the integrated tail copula estimator. This insight enables the construction of an effective statistical test for assessing whether the tail copula remains constant over time. Theoretical analysis confirms the asymptotic efficiency of the proposed estimator, and simulation studies corroborate its strong finite-sample performance and high statistical power.

heteroscedastic extremesmultivariate extreme valuenonidentically distributed

Hot Scholars

JR

Jordan Richards

Lecturer of Statistics, University of Edinburgh
Extreme value theorySpatial statisticsEnvironmental scienceStatistical deep learning
KC

Kim Christensen

Imperial College London
Complexity & Networks ScienceStatitical Physics
JS

Johan Segers

ISBA, LIDAM, Université catholique de Louvain
mathematical statisticsapplied probability
SR

Sayak Ray Chowdhury

Assistant Professor of Computer Science and Engineering, Indian Institute of Technology, Kanpur
Machine LearningMulti-armed BanditsReinforcement LearningDifferential Privacy