SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether latent sequence compression of large language model prompt prefixes can preserve reasoning capabilities. To this end, it proposes SemanticFold, a framework that achieves prefix compression by learning to fold boundary hidden states. Employing fixed-target protocols, linear probes, and paired bootstrap statistics, the authors systematically evaluate its impact on modeling, decoding, and reasoning across multi-scale models. The findings reveal that compression effects vary non-monotonically across metrics without a universal threshold, demonstrating the absence of any single scalar certificate for fidelity. Furthermore, the observed reduction in negative log-likelihood for Qwen models is primarily attributed to residual adaptation rather than sequence shortening, with endpoint responses exhibiting divergent directional shifts.
📝 Abstract
We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B. We use a fixed-target protocol: a frozen prefix is executed natively or compressed, and both arms teacher-force identical continuation tokens. This design rules out target-selection explanations for likelihood changes. We examine five endpoint families: fixed-target negative log-likelihood, finite-label reasoning accuracy, linear probe accessibility, open-ended generation, and systems-level memory and latency. We find that compression moves these endpoints non-monotonically and that they do not share a single compression threshold. On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws. On SmolLM2 at R=1.2, the mean change is 0.013 higher than native. On both Pythia checkpoints, NLL is effectively unchanged. An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone. MLP-only, which applies the transform without shortening, achieves 0.082 lower NLL than Full SemanticFold. Linear probe accuracy and macro AUC change by less than 0.03 in absolute value across conditions, with confidence intervals crossing zero. We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.
Problem

Research questions and friction points this paper is trying to address.

latent sequence compression
large language models
language modeling
decodability
reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Sequence Compression
SemanticFold
Fixed-Target Protocol
NLL Decomposition
Residual Adaptation
🔎 Similar Papers
No similar papers found.