Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of x-hacking in AutoML, where the coexistence of multiple near-optimal models may lead users to select models based on interpretability rather than predictive performance. To investigate this, we propose the ARSA ML framework, which leverages Rashomon effect analysis and binary classification benchmarks to quantify Rashomon set structures and predictive multiplicity, systematically evaluating the explanation stability and x-hacking susceptibility of mainstream AutoML tools. This work is the first to reveal fundamental differences in Rashomon set structures across AutoML frameworks, addressing a critical gap in existing tools’ inability to identify x-hacking risks. Experiments demonstrate that H2O exhibits significantly greater structural asymmetry than AutoGluon, resulting in higher predictive divergence and explanation instability, thereby exposing users to more severe x-hacking vulnerabilities.
📝 Abstract
The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification datasets, and conduct a post-hoc x-hacking analysis revealing a consistent structural asymmetry: AutoGluon produces larger, diverse sets with stable explanations, while H2O generates compact sets with markedly higher prediction divergence and explanation instability -- making H2O users considerably more exposed to x-hacking. This gap persists across all evaluated metrics and epsilon thresholds, pointing to a fundamental difference in each framework's model-building strategy. ARSA ML is available at https://pypi.org/project/arsa-ml/ .
Problem

Research questions and friction points this paper is trying to address.

AutoML
Rashomon effect
predictive multiplicity
x-hacking
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rashomon sets
AutoML
x-hacking
predictive multiplicity
ARSA ML
K
Katarzyna Woźnica
Warsaw University of Technology; Systems Research Institute, Polish Academy of Sciences
K
Katarzyna Rogalska
Warsaw University of Technology
Z
Zuzanna Sieńko
Warsaw University of Technology
Mustafa Cavus
Mustafa Cavus
Eskisehir Technical University, Department of Statistics
Statistical Machine LearningExplainable Artificial IntelligenceDesign of Experiment