🤖 AI Summary
This work proposes a context-aware multi-model optimization approach to overcome the limitations of relying on a single optimal model. By systematically identifying multiple models that exhibit substantially different feature selections yet achieve comparable predictive performance, the method preserves overall accuracy while enhancing interpretability. Applied to the METABRIC gene expression dataset, the approach successfully generates a diverse ensemble of high-performing models, uncovering multiple plausible biological interpretation pathways underlying the data. Compared to baseline methods, the resulting models demonstrate superior trade-offs between feature dissimilarity and performance consistency, thereby significantly improving both model interpretability and scientific insight.
📝 Abstract
In this paper, we propose an approach to finding sets of similar-performing models (in terms of loss/accuracy measurements) with highly different context-aware characteristics. Through experiments on the METABRIC dataset, we show that the proposed method finds multiple models with highly different gene expressions than those found by the control methodology without performance penalties. We argue that the proposed methodology is important whenever one aims to analyze any global characteristic of a model to extract insight into the underlying phenomenon being studied.