The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the disentanglement mechanisms between high-order logical interactions and memorization interference in shallow multilayer perceptrons (MLPs), addressing the inherent difficulty neural networks face in generalizing generative rules beyond rote memorization. Utilizing a parity-based synthetic benchmark with single-hidden-layer ReLU networks, we systematically evaluate the effects of SGD, Adam, and Muon optimizers alongside bias intervention techniques on mixed-order tasks, further introducing a noise-weight freezing strategy. Our findings reveal an intrinsic conflict between logical reasoning and mechanical memorization, as well as a critical relationship between objective symmetry and learned representations. Empirically, Muon achieves 99.21% accuracy on fourth-order tasks, while freezing noise weights elevates AdamW performance from 44.73% to 95.07%. These results demonstrate that optimization constraints can substantially enhance rule generalization capabilities.
📝 Abstract
A network can fit its training examples while failing to recover the rule that generated their labels. We examine this separation in single-hidden-layer multilayer perceptrons (MLPs), using synthetic tasks that control interaction order and the presence of nuisance inputs. We establish elementary benchmark properties: pure parity contains no predictive lower-order marginals, admits an exact Bayes posterior, and can be represented on clean latent inputs by a width-$k$ ReLU network. Experiments then identify distinct optimization outcomes. In a matched order-2--4 sweep, SGD, Adam, and Muon all reach 100\% peak test accuracy at order two; at order three they reach 96.25\%, 50.87\%, and 76.82\%, respectively, while Muon reaches 99.21\% at order four. In a separate mixed-order task, freezing only the first-layer weights connected to independent nuisance inputs raises AdamW's epoch-10 accuracy from 44.73\% to 95.07\%. Removing the same inputs only at test time raises it to 48.38\%. Thus, nuisance-weight learning changes the training outcome beyond its immediate effect on prediction. Bias interventions expose a connection between target symmetry and shallow ReLU representations. In a compact signal-only regime, both SGD and Muon learn orders five through eight, with higher SGD peak accuracy at orders nine through eleven. Together, the results show how optimization and nuisance learning constrain the higher-order rules realized by a shallow network.
Problem

Research questions and friction points this paper is trying to address.

higher-order interactions
shallow MLPs
logic vs memory
nuisance inputs
optimization dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Higher-Order Interactions
Shallow MLPs
Optimization Dynamics
Nuisance Inputs
Parity Task
🔎 Similar Papers
No similar papers found.