Kolmogorov--Arnold Networks for Small Language Models

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the replacement of conventional MLP-based feedforward networks in small language models with Kolmogorov–Arnold Networks (KANs) to enhance interpretability and compression capabilities. By employing a KAN architecture grounded in learnable univariate edge functions—such as GR-KAN—the study explicitly models feedforward pathways and integrates functional principal component analysis (fPCA) for compression alongside edge-level pruning. Systematic evaluations are conducted on benchmarks including BabyLM and Wikitext-103. Experimental results demonstrate that sparse KAN variants enable high-ratio pruning and permit auditable scalar transformations; however, they fail to consistently outperform MLP baselines in language modeling and grammaticality judgment tasks, showing no reliable gains in either performance or inference latency.
📝 Abstract
Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes negligible loss increase, although structured MLP neuron pruning tolerates comparable sparsity. The audit replicates on BabyLM, but grid-size sweeps show that near-total fPCA compression and high closed-form-fit coverage are properties of the low-capacity grid-2 basis, not universal KAN behavior. For replacement, we evaluate MLP, SwiGLU, grouped Chebyshev, and rational GR-KAN networks on BabyLM. The KAN-family and gated variants improve validation loss over the GELU MLP, but this ordering does not transfer to standardized benchmarks: across ten seeds and 59,875 BLiMP pairs, accuracies span 62.4--63.1\%, EWoK remains at chance, and a (+0.7)-point GR-KAN effect on BLiMP reverses on the supplement. Larger tests are also cautionary: parameter-matched MLPEdge underperforms the MLP on Wikitext-103, and 286M-parameter GR-KAN remains below a SwiGLU ClimbMix baseline after stabilization. Thus, small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.
Problem

Research questions and friction points this paper is trying to address.

Kolmogorov–Arnold Networks
small language models
feed-forward networks
model interpretability
MLP replacement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kolmogorov–Arnold Networks
learnable edge functions
model interpretability
small language models
feed-forward network alternatives
🔎 Similar Papers
No similar papers found.