Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether compression methods inspired by solid-state physics can be effectively applied to small-scale language models. Employing rigorous research paradigms—including preregistration, three-sigma gating, and ensemble frameworks—the authors empirically evaluate five physical mapping hypotheses utilizing techniques such as Wannier functions, tight-binding models, DMRG truncation, and Wilsonian renormalization group. The work reports four negative results, falsifying most hypotheses. Notably, it reveals a prediction reversal phenomenon at small scales, confirms the absence of performance gains from rank-surface methods, and demonstrates that attention mechanisms exhibit critical or glassy characteristics rather than nearest-neighbor insulator behavior. All data have been made publicly available, providing crucial counterintuitive evidence and methodological references for interdisciplinary model compression.
📝 Abstract
We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
Problem

Research questions and friction points this paper is trying to address.

LLM compression
solid-state physics
negative results
pre-registration
small language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pre-registration
LLM Compression
Solid-state Physics
Autonomous Research Agent
Negative Results
💼 Related Jobs
No related jobs found.
J
Jun-qiang Lu
Department of Physics, University of Puerto Rico, Mayagüez, PR 00680, USA