Toy Combinatorial Interpretability Models Reveal Lottery Tickets in Early Feature Space

📅 2026-05-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the intrinsic structural properties preserved by winning subnetworks in the lottery ticket hypothesis within feature space. By constructing an interpretable compositional toy model and integrating feature-space distance metrics, lightweight probing techniques, structured SGD training, and sparse retraining, the study reveals that winning tickets do not rely on specific weight values or neuron identities. Instead, they correspond to a family of compatible encoding positions that are proximate to their final representations at initialization and exhibit low mutual interference. The proposed lightweight probe significantly outperforms conventional weight-based lottery identification methods in both accuracy and encoding recovery fidelity, thereby demonstrating that the geometric structure of feature space plays a dominant role in the lottery ticket mechanism.
📝 Abstract
The lottery ticket hypothesis posits that dense networks contain sparse subnetworks, ``winning tickets,'' that, when rewound to their initial weights and retrained in isolation, match the performance of the full model. We ask a more mechanistic question: what internal object does a winning ticket preserve? We work in a combinatorial, clause-structured toy setting that admits an interpretable feature-space representation with well-defined combinatorial distances between features. We show that winning tickets in weight space correspond to precursor locations in feature space that are already near, at initialization, to the final feature-channel codes. Dense SGD resolves these locations through structured selection: proximal locations either converge to final codes or are rejected, with rejection concentrated at more crowded neurons, implicating competition under superposition. A winning ticket is thus a family of compatible code locations that jointly balance proximity to final codes with low inter-feature interference. Sparse retraining often re-expresses the same clause/template family on a different row, so the preserved object is family-level rather than microscopic row identity. We validate this account with lightweight probes based on feature-space distance and motion; in our setting, these probes frequently outperform established weight-based ticket discovery methods in both accuracy and exact code recovery. Although these findings are grounded in a toy setting, they suggest that the lottery ticket structure is governed by hidden feature-space geometry rather than weight-space subnetwork identity.
Problem

Research questions and friction points this paper is trying to address.

lottery ticket hypothesis
feature space
combinatorial interpretability
sparse subnetworks
neural network initialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

lottery ticket hypothesis
feature-space geometry
combinatorial interpretability
structured selection
code family preservation