π€ AI Summary
This study addresses the limitation of conventional research that reduces neural network interference to geometric overlap while neglecting code statistics and actual interactions. We propose the concept of "effective interference," which integrates feature geometry with coding statistics to distinguish constructive from destructive interference and quantify interaction strength. Based on sparse autoencoders and a local fixed-support assumption, this work reveals that constrained architectures shape interference patterns through four mechanisms, including orthogonalization and bias compensation. Our findings demonstrate that architectural constraints selectively reduce co-activation overlap while preserving beneficial cross-contributions. These results establish that interference inherently depends on usage patterns and network architecture, thereby overcoming the theoretical limitations of purely geometric perspectives.
π Abstract
Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.