🤖 AI Summary
This work addresses the tendency of large language models to prematurely commit to a single semantic interpretation when processing ambiguous inputs, thereby collapsing multiple plausible meanings into one output. To mitigate this, the authors propose a text-to-state mapping framework φ that preserves semantic ambiguity through a three-stage process: conflict detection, interpretation extraction, and state construction—operating within a non-collapsed state space. This framework establishes the first algorithmic bridge from text to a state space for non-resolution reasoning (NRR), enabling delayed collapse of ambiguity and supporting cross-lingual extension. It integrates rule-based explicit conflict markers (e.g., contrastive conjunctions) with LLM-driven enumeration of implicit ambiguities (cognitive, lexical, and structural). Evaluated on a test set of 68 ambiguous sentences, the method yields generated states with an average entropy of 1.087 bits—significantly outperforming collapsed baselines (entropy = 0)—and demonstrates cross-lingual validity using Japanese markers.
📝 Abstract
Large language models exhibit a systematic tendency toward early semantic commitment: given ambiguous input, they collapse multiple valid interpretations into a single response before sufficient context is available. We present a formal framework for text-to-state mapping ($\phi: \mathcal{T} \to \mathcal{S}$) that transforms natural language into a non-collapsing state space where multiple interpretations coexist. The mapping decomposes into three stages: conflict detection, interpretation extraction, and state construction. We instantiate $\phi$ with a hybrid extraction pipeline combining rule-based segmentation for explicit conflict markers (adversative conjunctions, hedging expressions) with LLM-based enumeration of implicit ambiguity (epistemic, lexical, structural). On a test set of 68 ambiguous sentences, the resulting states preserve interpretive multiplicity: mean state entropy $H = 1.087$ bits across ambiguity categories, compared to $H = 0$ for collapse-based baselines. We additionally instantiate the rule-based conflict detector for Japanese markers to illustrate cross-lingual portability. This framework extends Non-Resolution Reasoning (NRR) by providing the missing algorithmic bridge between text and the NRR state space, enabling architectural collapse deferment in LLM inference. Design principles for state-to-state transformations are detailed in the Appendix, with empirical validation on 580 test cases showing 0% collapse for principle-satisfying operators versus up to 17.8% for violating operators.