Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing pretraining objectives are limited by the expressive capacity of symbolic primitives and data scale, hindering a clear understanding of skill emergence in large language models. This work proposes Logic-PPT (Logic-based Pre-Pretraining), a novel approach that incorporates formal deductive reasoning into the pretraining phase to endow language models with structured inductive biases, thereby accelerating the acquisition of natural language competencies. Training on 100B tokens and analyzing representation geometry, we find that Logic-PPT induces low-rank, spectrally concentrated representations amenable to lossless compression under high sparsity. Empirically, models trained with only 64B tokens via Logic-PPT match the performance of standard baselines trained on 100B tokens (achieving 80% accuracy) and retain comparable effectiveness to dense models at approximately 33% sparsity.
📝 Abstract
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.
Problem

Research questions and friction points this paper is trying to address.

pre-pretraining
formal derivations
natural language acquisition
structural bias
skill emergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

logic pre-pretraining
formal derivations
structural reorganization
model compressibility
skill acquisition
🔎 Similar Papers
2024-10-09Conference on Empirical Methods in Natural Language ProcessingCitations: 3