Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing data attribution methods for diffusion models, which overlook the dynamic evolution of semantics during generation. To this end, we propose CADT, a framework that constructs counterfactual trajectories to extend static scalar attribution into dynamic response analysis, thereby revealing the evolutionary mechanisms of training samples throughout the denoising process. Furthermore, this work pioneers a dynamic trajectory-based concept attribution paradigm that leverages covariance-aware kernel calibration to align query and training representations, enabling a deeper analytical shift from identifying "who influences" to understanding "how influence occurs." Extensive experiments demonstrate that CADT significantly outperforms existing baseline methods across hierarchical, compositional, and stylistic attribution tasks on multiple public datasets.
📝 Abstract
Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiting for diffusion models, where semantic factors emerge through evolving representation dynamics during denoising. We therefore reformulate diffusion data attribution as attributing factor-induced internal response trajectories. In this paper, we propose a novel Concept Attribution method through Dynamic Trajectories(CADT). We argue that attribution should therefore ask not only \emph{which} examples matter, but also \emph{how} their influence unfolds during generation. Specifically, we construct matched counterfactual pairs at identical noisy states to isolate factor-specific representation displacements, and model their directional and magnitude evolution across denoising as dynamic attribution signatures. For each training example and generated query, CADT extracts stage-wise feature vectors and integrates them along the denoising process to form a trajectory descriptor. Applying the same construction across the training set yields a bank of factor-specific trajectory descriptors. The covariance statistics of this bank are then used to construct . CADT uses this covariance-aware positive-semidefinite kernel to calibrate the query and training representations, and compares the calibrated query trajectory with each training trajectory to produce the final training-sample attribution scores. Experiments on multiple public datasets show consistent improvements over existing diffusion attribution baselines across hierarchical, compositional, and style attribution.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Data Attribution
Counterfactual Trajectories
Representation Dynamics
Concept Attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Models
Data Attribution
Counterfactual Trajectories
Dynamic Representation
Covariance-aware Kernel