Score
Theoretical and inferential tools for modeling tail behavior and rare events, deriving conditions that relate extrapolation ranges to sample size and tail index, and developing robust inference procedures for sparse exceedances and discrete supports.
Quantifying tail risk (e.g., financial crashes, extreme weather) under serially dependent data remains challenging due to the inadequacy of standard independence-based tail modeling. Method: This paper proposes a Bayesian inference framework based on the Generalized Pareto Distribution (GPD) that jointly models static marginal and dynamic conditional tail behavior of threshold exceedances, integrating beta-mixture dependence structures, heteroscedastic regression, and empirical log-likelihood theory—thereby relaxing the restrictive independence assumption. Contribution/Results: It establishes, for the first time under serial dependence, asymptotically honest Bayesian credible regions for tail parameters and derives theoretical conditions for prior admissibility. Simulation studies demonstrate substantial improvements over classical methods across ARMA, GARCH, and Markov copula models. Empirical applications to U.S. interest rates and Swiss electricity demand confirm high accuracy, robustness, and practical applicability in real-world tail risk assessment.
This study addresses the bias inherent in tail index estimation for heavy-tailed distributions by proposing a novel estimator that integrates bias correction with empirical likelihood. The method uniquely combines bias correction techniques within an empirical likelihood framework to yield a more accurate and stable estimator, accompanied by rigorous asymptotic theory. Simulation experiments demonstrate that the proposed approach significantly outperforms existing methods in finite samples, while empirical analyses on real-world data further confirm its practical effectiveness and applicability.
This paper identifies the systematic failure of classical statistical methods—including mean-based inference, principal component analysis (PCA), and asymptotic normality assumptions—under heavy-tailed distributions, particularly in medium-sample-size (medium-*n*) real-world settings where they are routinely misapplied. Methodologically, it challenges the uncritical adoption of Gaussian and stable-distribution assumptions and introduces the “Median Law” theoretical framework, which formalizes fundamental limitations under heavy tails: unreliable sample means, distorted empirical distributions, and degenerate principal components. The approach integrates extreme value theory, generalized stable distribution modeling, robust parametric estimation, and pre-asymptotic analysis. Empirical validation draws on counterexamples from finance, economics, and psychology, supplemented by cross-disciplinary case studies. The core contribution is a foundational rethinking of uncertainty quantification and causal inference: it demonstrates that many canonical “cognitive biases” are, in fact, rational inferences under heavy-tailed probability structures—thereby advocating a paradigm shift in statistical practice from idealized asymptotics to empirically grounded probabilistic modeling.
Existing probabilistic forecasting evaluation methods lack the ability to characterize tail calibration—critical for high-impact extreme events, whose reliability is increasingly vital for risk-informed decision-making. Method: This paper introduces, for the first time, a general definition of tail calibration, rigorously connecting it to classical probabilistic calibration theory and integrating the Peaks-over-Threshold (POT) framework from extreme value theory. We develop an operational diagnostic framework by unifying probabilistic calibration theory, extreme-value statistics, diagnostic statistical tests, and empirical analysis. Contribution/Results: Applied to European precipitation forecasts, our framework significantly improves the quantification of predictive credibility for high-impact, rare events. It enables rigorous assessment of tail behavior in probabilistic forecasts and establishes a novel paradigm for extreme-event risk assessment and decision support.
This work addresses the challenge of traditional extrapolation methods failing in extreme regions due to data scarcity in the tails—a common issue in machine learning. To overcome this limitation, the authors propose a unified extreme-value extrapolation framework that integrates extreme value theory with statistical learning. Built upon asymptotic representations of univariate and multivariate tail distributions, the framework combines extreme value index estimation, tail distribution modeling, and dependence structure analysis. It is applicable to both supervised and unsupervised settings and accommodates both asymptotically dependent and independent data. Empirical evaluations demonstrate that the method substantially outperforms existing approaches in tasks such as extreme quantile regression, anomaly detection, and generative AI, yielding improved accuracy and robustness in predicting rare and extreme events.
This study addresses the challenges of modeling extremes in data containing zero-valued interior points, where conventional methods struggle to accurately estimate thresholds and tail characteristics. To overcome these limitations, the paper proposes the first unified mixture model for extreme value analysis that simultaneously captures the interior distribution, tail behavior, and their relative proportions. Parameter inference is performed via maximum likelihood estimation, and model validation is comprehensively assessed using mean excess plots, parameter stability plots, and Pickands plots. Extensive simulations and real-data experiments demonstrate that the proposed approach significantly outperforms existing methods in threshold selection, tail parameter estimation, and overall model stability, effectively mitigating the inadequacies of traditional models in representing interior point structures.
Existing extreme-value models struggle to flexibly characterize both marginal tail behavior and joint dependence structures of bivariate threshold exceedances under asymptotic independence. This work proposes a novel bivariate subasymptotic parametric model that, while ensuring convergence of the margins to the generalized Pareto distribution, naturally captures the evolution of extremal dependence with varying thresholds through its scale parameters. The model encompasses the standard multivariate generalized Pareto distribution as a limiting special case and accommodates a broad spectrum of tail dependence patterns. Inference is carried out via a likelihood-free neural Bayesian approach with tailored priors, enabling direct computation and interpretation of failure probabilities. Extensive simulations and an analysis of Belgian rainfall extremes demonstrate the model’s flexibility and the effectiveness of the proposed inference framework.
This study addresses the lack of interpretable, scalable diagnostic tools for extreme value regression models that can identify regions in covariate space where local fit is poor. The authors propose two visualization-based diagnostics—standardized tail plots and normalized residual plots—leveraging the asymptotic distribution of normalized exceedance probabilities to construct sample-size-invariant uncertainty bounds. This enables consistent assessment of both global and local goodness-of-fit. Notably, the approach provides the first framework for local diagnostics in low-dimensional or non-Euclidean covariate domains, supports model comparison across varying sample sizes, and facilitates large-scale model screening. In two real-world applications, the method successfully evaluated thousands of candidate models, yielding actionable modeling recommendations that substantially enhance the reliability and practical utility of extreme value regression models.
This work proposes a unified framework for anomaly detection based on surprisal, addressing the limitations of traditional methods that rely on ad hoc rules or strong modeling assumptions and often fail to identify “inlier” anomalies in low-density regions. The approach defines anomaly scores as the surprisal of observations under a possibly misspecified model and quantifies anomaly severity via upper-tail probabilities. By reducing high-dimensional anomaly detection to a univariate tail estimation problem, the method combines the empirical distribution with the Generalized Pareto Distribution (GPD) to estimate tail probabilities and leverages the Dvoretzky–Kiefer–Wolfowitz inequality to provide finite-sample confidence guarantees. Experiments demonstrate that the method robustly detects both tail and inlier anomalies even under substantial model misspecification, achieving strong performance on synthetic data as well as real-world datasets including French mortality records and cricket test matches.