π€ AI Summary
This study addresses the challenge of detecting strongly influential outliers in clustered data within mixed-effects models, where existing methods are often constrained by model assumptions and lack generalizable diagnostic tools. To overcome these limitations, this work proposes a model-agnostic, point-level anomaly detection framework that integrates SHAP values with residuals to construct comprehensive features. It leverages Normalizing Flows to map complex distributions into a standard space, enabling precise outlier identification across diverse base learners, including linear models, random forests, and gradient boosting trees. The primary contributions include providing goodness-of-fit diagnostics through flow-based modeling, empirically validating the methodβs effectiveness and generalizability across multiple architectures, and systematically analyzing its advantages and limitations.
π Abstract
Influential Outlier Detection is developed for mixed-effects models on clustered data. The Influential Outlier Metric is defined as a combination of SHapley Additive exPlanantion (SHAP) values and model residuals, both of which undergo a change of measure transformation. Building on previous work showcasing the suitability of using Normalizing flows to map arbitrary distributions to a flexible base distribution for statistical inference, the Normalizing Flows are constructed to allows contextual information and also provide a goodness of fit diagnostic for model evaluation. The use of SHAP values in the construction moves away from model specific tools and instead provides point-wise model agnostic influential outlier. The advantages and limitations of this approach are examined in several models including the linear model, the random forest, and gradient-boosted trees.