🤖 AI Summary
This work addresses the challenge of interpreting black-box machine learning models by proposing a generative diagnostic framework that reveals model data preferences and decision mechanisms through controllable synthetic samples. Methodologically, it formally defines and unifies three types of “model-query samples”—high-risk, parameter-sensitive, and model-comparative—via gradient-guided optimization, latent-space inversion, and constrained generative modeling, augmented by loss-sensitivity analysis for fine-grained semantic control. Extensive experiments across diverse architectures (CNNs, Transformers) and modalities (image, tabular data) demonstrate that the framework effectively characterizes decision boundaries, pinpoints vulnerability regions, and quantifies inter-model discrepancies. It significantly enhances the capability to verify model behavior interpretability, offering a general-purpose diagnostic tool applicable across tasks and model architectures.
📝 Abstract
There is a growing need for investigating how machine learning models operate. With this work, we aim to understand trained machine learning models by questioning their data preferences. We propose a mathematical framework that allows us to probe trained models and identify their preferred samples in various scenarios including prediction-risky, parameter-sensitive, or model-contrastive samples. To showcase our framework, we pose these queries to a range of models trained on a range of classification and regression tasks, and receive answers in the form of generated data.