Score
Designs and constructs diagnostic plots of normalised or standardised residuals (including specialised tail/exceedance plots) that visualise residual patterns and tail behaviour and display uncertainty bounds that are roughly insensitive to sample size. Uses these plots to detect and quantify regional or global model misfit and to compare fit across non‑uniform covariate domains.
This study addresses the lack of interpretable, scalable diagnostic tools for extreme value regression models that can identify regions in covariate space where local fit is poor. The authors propose two visualization-based diagnostics—standardized tail plots and normalized residual plots—leveraging the asymptotic distribution of normalized exceedance probabilities to construct sample-size-invariant uncertainty bounds. This enables consistent assessment of both global and local goodness-of-fit. Notably, the approach provides the first framework for local diagnostics in low-dimensional or non-Euclidean covariate domains, supports model comparison across varying sample sizes, and facilitates large-scale model screening. In two real-world applications, the method successfully evaluated thousands of candidate models, yielding actionable modeling recommendations that substantially enhance the reliability and practical utility of extreme value regression models.
This study addresses the limitations of traditional residual plot diagnostics—namely, their reliance on subjective human interpretation, low efficiency, and poor scalability—by introducing computer vision techniques for the first time to automate the assessment of residual plots in linear models. The authors develop the R package autovi and an accompanying Shiny-based interactive web application, autovi.web. Their approach leverages deep visual models to quantify the strength of structural signals in residual plots and produces interpretable diagnostic metrics. This methodology significantly enhances the consistency and efficiency of model fit evaluation, offering statisticians and data analysts a robust, objective, and scalable tool for automated diagnostic assessment in statistical modeling.
This study addresses the lack of intuitive and interpretable graphical diagnostic tools for parametric models in survival analysis by proposing a local normalized hazard comparison curve. The method constructs a test process that, under correct model specification, is approximately standard normal at each time point by locally normalizing the difference between nonparametric and parametric hazard rate estimates. It is applicable to a range of models, including exponential, Weibull, Gompertz, Gamma, and parametric Cox regression, substantially enhancing the interpretability and practical utility of goodness-of-fit assessment. Leveraging stochastic process theory, the authors develop an S-Plus implementation that demonstrates strong empirical power in both simulation studies and real-data applications, and the approach naturally extends to discrete-time settings.
This study addresses the challenges of modeling extremes in data containing zero-valued interior points, where conventional methods struggle to accurately estimate thresholds and tail characteristics. To overcome these limitations, the paper proposes the first unified mixture model for extreme value analysis that simultaneously captures the interior distribution, tail behavior, and their relative proportions. Parameter inference is performed via maximum likelihood estimation, and model validation is comprehensively assessed using mean excess plots, parameter stability plots, and Pickands plots. Extensive simulations and real-data experiments demonstrate that the proposed approach significantly outperforms existing methods in threshold selection, tail parameter estimation, and overall model stability, effectively mitigating the inadequacies of traditional models in representing interior point structures.
Existing probabilistic forecasting evaluation methods lack the ability to characterize tail calibration—critical for high-impact extreme events, whose reliability is increasingly vital for risk-informed decision-making. Method: This paper introduces, for the first time, a general definition of tail calibration, rigorously connecting it to classical probabilistic calibration theory and integrating the Peaks-over-Threshold (POT) framework from extreme value theory. We develop an operational diagnostic framework by unifying probabilistic calibration theory, extreme-value statistics, diagnostic statistical tests, and empirical analysis. Contribution/Results: Applied to European precipitation forecasts, our framework significantly improves the quantification of predictive credibility for high-impact, rare events. It enables rigorous assessment of tail behavior in probabilistic forecasts and establishes a novel paradigm for extreme-event risk assessment and decision support.
Existing visualization methods lack systematic support for random variables, making it challenging to balance statistical rigor and graphical flexibility in exploratory data analysis. This work proposes a mathematical framework that treats visualizations as continuous functions, leveraging the Continuous Mapping Theorem to handle random-variable inputs. By decomposing visualization functions, the approach identifies and reconstructs components incompatible with uncertainty-aware data. As a core contribution, the study introduces the first complete integration of uncertainty into the grammar of graphics, ensuring that resulting visualizations inherit the convergence properties of their underlying random variables. To operationalize this framework, the authors develop ggdibbler, an R package extending ggplot2 that enables direct plotting of random variables, thereby enhancing analytical flexibility while preserving statistical validity.
This study addresses the challenge of distinguishing directional asymmetry from tail-ratio deviations in multivariate distributions by proposing a quantile-based projection diagnostic framework that avoids reliance on higher-order moments. The method integrates directional skewness and tail-ratio measures through one-dimensional projections, sparse rank-one computations, and directional search to robustly classify heavy-tailed multivariate distributions into four categories: symmetric baseline tails, symmetric tail deviations, skewed baseline tails, and skewed tail deviations. Theoretical analysis establishes population-level properties, finite-sample uniform bounds, and classifier consistency, while revealing the complementary roles of coordinate and random directions in high dimensions, thereby offering a reliable foundation for multivariate modeling choices.
本文提出一种鲁棒建模框架,通过最大似然估计参数,解决含内部点数据中极端值建模与尾部估计的问题。
This study addresses the instability and poor interpretability of parameter estimates in extreme value distributions under small-sample settings, which undermine the reliability of risk assessments. For the first time, the authors systematically apply the orthogonal parametrization theory of Cox and Reid (1987) to the generalized extreme value, generalized Pareto, and Gumbel distributions, proposing an orthogonal reparameterization approach that aligns model parameters more closely with practical modeling objectives. Simulation experiments demonstrate that this method substantially enhances both the stability and interpretability of parameter estimates in small samples, thereby providing a more robust statistical foundation for extreme value modeling.
This study addresses the challenges of small-sample linear regression, where violations of residual normality, homoscedasticity, and independence often lead to biased estimates and poor generalization. To mitigate these issues, the authors propose GP-MEVT, a novel method that uniquely integrates Gaussian processes with extreme value theory for data augmentation. This approach expands the predictor space while preserving the underlying linear structure and incorporates controlled noise through residual variability. Empirical results demonstrate that GP-MEVT substantially improves adherence to model assumptions—achieving a 67.1% satisfaction rate, markedly higher than the 17.3% and 21.2% obtained by conventional bootstrap methods—yields parameter estimates closer to true values, and achieves superior root mean squared error (RMSE) performance, thereby significantly enhancing regression accuracy in small-sample settings.