🤖 AI Summary
This study addresses the difficulty large language models face in preserving implicit assumptions and quantification scopes during cross-subdomain translation. To investigate this, we propose an external-criterion-free measurement tool for quantification scope sensitivity, employing a decoupled blind-testing mechanism that independently encodes truth value, content, and scope. Combined with positive controls and preregistered protocols, this framework systematically benchmarks model capacity to recover the generalization levels of source texts. Our findings reveal that when translating toward greater generality, models significantly expand quantification domains (60.6%) while rarely articulating necessary assumptions. These results demonstrate that current models lack the capacity for constrained mapping across theories, highlighting fundamental limitations in their ability to maintain logical rigor during cross-domain knowledge transfer.
📝 Abstract
Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements between the dialects of neighbouring subfields, where fidelity turns on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relation between the theories. We introduce an instrument that codes truth, content and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source, and establish its sensitivity with a planted-positive control. Across seven models from four families, translating towards the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none; translating towards the concrete framing narrows it in 28.3% and widens it in 0.3%. The hypothesis that would prevent it is stated in 21.6% of model rewrites and 4.2% of human statements. Capability does not govern the asymmetry: it appears in every model tested, and the most capable widens least. It replicates on the half of the benchmark held out by a pre-registered rule, and on statements written by mathematicians. Instructing a model to state every hypothesis it requires raises that rate but not its sensitivity to direction. We argue that these systems have acquired an object-level correspondence between subfield vocabularies without the constraint under which a translation between theories carries hypotheses to hypotheses.