Score
Designs and implements diagnostic analyses and evaluation pipelines for classifiers that operate on or are derived from diffusion models, including quantitative benchmarking and bias assessments. Builds visualization and analytic artifacts—e.g., reconstruction‑error heatmaps, cross‑attention / U‑Net attention maps, and standardized benchmark evaluations—to identify reliance on foreground vs. background signals, failure modes, and dataset or model biases.
To address the challenges of artifact localization and interpretable repair in text-to-image diffusion models, this paper proposes a two-stage “diagnose-then-treat” optimization framework. In the first stage, a pixel-level artifact detector is constructed to enable fine-grained, localization-aware defect identification. In the second stage, the detection confidence map is integrated into the diffusion reverse process via gradient modulation and pixel-wise weighted loss to guide precise artifact correction. Our key contributions include: (i) the first introduction of localization-aware diagnostic modeling into diffusion optimization; (ii) construction of a million-scale defective image dataset with a human-in-the-loop annotation protocol. Experiments across multiple mainstream diffusion models show an average 42.7% reduction in artifact rate, a 3.2 improvement in FID, and an mAP@0.5 of 68.9—demonstrating both strong visual interpretability and restoration efficacy.
This work addresses the opacity and poorly understood bias characteristics of diffusion-based classifiers in zero-shot classification. The authors introduce ASOB-Bench, the first benchmark designed to systematically evaluate decision biases along three dimensions: attribute binding, size-order bias, and background dependence. By integrating reconstruction error heatmaps and U-Net cross-attention visualizations, they uncover the underlying mechanisms driving these biases. Experimental results reveal that while diffusion classifiers exhibit fewer attribute mismatches compared to OpenCLIP, they suffer from more pronounced size-order bias and background dependence, leading to a significant drop in accuracy on ImageNet-B. These findings highlight distinct bias profiles relative to vision-language models and expose potential failure modes inherent to generative classification paradigms.
Multimodal large language models (MLLMs) benchmarks suffer from pervasive non-visual shortcut learning—models achieve high scores by exploiting textual biases, linguistic priors, or superficial statistical patterns, severely compromising the validity of visual understanding evaluation. Method: We propose a “test-set stress testing” and “iterative bias pruning” framework that leverages LLMs to actively detect and quantify textual biases in benchmarks. Using k-fold cross-validation, we fine-tune an LLM and integrate it with random forests and handcrafted features to score and prune biased samples. Contribution/Results: Our method systematically identifies and eliminates non-visually solvable instances across four mainstream benchmarks, yielding the debiased benchmark VSI-Bench-Debiased. It exhibits significantly reduced non-visual solvability, widened performance gaps on visually blind tasks, and robustly advances a vision-centric, reliable paradigm for multimodal evaluation.
While generated images appear photorealistic to human observers, it remains unclear whether they are equally indistinguishable to neural network classifiers—revealing a potential perception gap between human vision and model discrimination. Method: We propose a distribution-level discriminability analysis framework, conducting controlled comparative experiments across multiple diffusion architectures (DiT, EDM2, U-ViT), complemented by feature attribution and classifier guidance techniques. Contributions/Results: (1) State-of-the-art diffusion models still exhibit significant classifier-detectable artifacts; cross-architecture samples are readily distinguishable, whereas same-family models of varying scales remain largely confusable. (2) We pioneer the use of off-the-shelf classifiers as diagnostic tools for generative models, uncovering a “model self-cannibalization imbalance” phenomenon wherein generators overfit to classifier-specific biases. (3) Classifier guidance is empirically shown to meaningfully enhance perceptual realism and provides an interpretable theoretical foundation for classifier-aware data augmentation.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
Existing model explanation methods often suffer from attribution bias or even erroneous interpretations due to inadequate consideration of baseline selection. This work reformulates the model explanation task by unifying gradient-based methods, Integrated Gradients (IG), and Taylor expansion approaches, thereby systematically revealing— for the first time—the pivotal role of the baseline in attribution. Building on this insight, the authors propose an evaluation framework grounded in attribution error and develop a general-purpose explanation method with a well-defined, principled baseline that supports feature attribution at arbitrary network layers. The refined IG variant significantly improves explanation accuracy across multiple benchmarks, and attributions derived from different layers coherently reflect the hierarchical nature of feature extraction in deep networks.