🤖 AI Summary
This study investigates whether denoising models genuinely encode human visual illusions and their underlying mechanisms within internal representations. By analyzing internal activations across multiple architectures—combined with feature visualization, channel ablation, psychophysical modeling (FLODOG), and parametric illusion-strength experiments—the work uncovers, for the first time, a perception-like “phantom” representation that is decoupled from model output. Specifically, certain channels in intermediate layers exhibit high sensitivity to brightness illusions: their activation magnitudes correlate strongly with human perceptual judgments (Spearman ρ ≥ 0.70) and vary monotonically with illusion strength, yet they exert no influence on the final reconstructed pixels. These findings provide causal evidence for human-like perceptual mechanisms embedded within deep neural networks.
📝 Abstract
Deep neural networks trained on natural images are shown to produce outputs consistent with human observers for brightness illusions. While this phenomenon has been documented across architectures, all evidence, to date, is measured at the output level: restored pixels, decoded trajectories, or classification decisions. Whether these models actually represent illusions internally, and if so where and how, remains unknown. We show that denoising models develop illusion-sensitive representations at specific internal layers, across varied architectures. Specifically, we identify the layers and channels that discriminate illusory from physically matched control regions. We show that the denoising objective is a more important driver of the effect than the architecture. On domain-appropriate stimuli, these activations track a validated psychophysical model of human brightness perception (FLODOG; Spearman $ρ\geq 0.70$) and scale monotonically with parametric illusion strength. Leveraging these findings, we provide causal evidence via channel ablation showing that illusion-sensitive channels specifically and substantially affect the internal signal. Yet injecting these representations into the generation pipeline produces no measurable pixel shift across all tested architectures; we term such representations perceptual phantoms: active in internal processing yet invisible to any output-based evaluation. While related internal-output dissociations have been characterized in language models, this is the first such characterization for perceptual representations in denoising vision models.