🤖 AI Summary
This work addresses the interpretability and safety of deep neural network representations. We propose a novel geometric modeling paradigm: treating neural representations as coordinate systems embedded in the data distribution, abstracting generic data distributions as random lattice structures, and—crucially—introducing percolation theory for the first time to analyze their structural properties. Building on this foundation, we establish a unified classification framework for three representation types—contextual, compositional, and surface features—qualitatively integrating multiple mechanistic interpretability findings. By unifying geometric representation modeling, random lattice theory, and percolation analysis, we derive mathematically verifiable links between data distribution structure and neural representation type. This yields the first theoretically rigorous and empirically testable mathematical framework for representation disentanglement, and charts a principled path toward safe, reliable AI representation learning. (149 words)
📝 Abstract
Decomposing a deep neural network's learned representations into interpretable features could greatly enhance its safety and reliability. To better understand features, we adopt a geometric perspective, viewing them as a learned coordinate system for mapping an embedded data distribution. We motivate a model of a generic data distribution as a random lattice and analyze its properties using percolation theory. Learned features are categorized into context, component, and surface features. The model is qualitatively consistent with recent findings in mechanistic interpretability and suggests directions for future research.