Generating Samples to Question Trained Models

📅 2025-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of interpreting black-box machine learning models by proposing a generative diagnostic framework that reveals model data preferences and decision mechanisms through controllable synthetic samples. Methodologically, it formally defines and unifies three types of “model-query samples”—high-risk, parameter-sensitive, and model-comparative—via gradient-guided optimization, latent-space inversion, and constrained generative modeling, augmented by loss-sensitivity analysis for fine-grained semantic control. Extensive experiments across diverse architectures (CNNs, Transformers) and modalities (image, tabular data) demonstrate that the framework effectively characterizes decision boundaries, pinpoints vulnerability regions, and quantifies inter-model discrepancies. It significantly enhances the capability to verify model behavior interpretability, offering a general-purpose diagnostic tool applicable across tasks and model architectures.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Interpretability, Explainability, and TransparencyKnowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
There is a growing need for investigating how machine learning models operate. With this work, we aim to understand trained machine learning models by questioning their data preferences. We propose a mathematical framework that allows us to probe trained models and identify their preferred samples in various scenarios including prediction-risky, parameter-sensitive, or model-contrastive samples. To showcase our framework, we pose these queries to a range of models trained on a range of classification and regression tasks, and receive answers in the form of generated data.
Problem

Research questions and friction points this paper is trying to address.

Probe trained machine learning models
Identify preferred data samples
Generate data for model questioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

mathematical framework probes models
identifies preferred data samples
generates data for analysis
🔎 Similar Papers
No similar papers found.
E
E. Mehmet Kıral
Keio University
N
Nurcsen Aydin
University of Warwick
S
S. Ilker Birbil
University of Amsterdam