🤖 AI Summary
This study addresses the issues of token omission and hallucination inherent in generative approaches to multi-source interleaved text restoration by proposing the Evidence-Preserving Ownership Routing (EPOR) framework. This work pioneers a paradigm that decouples attribution inference from text reconstruction, integrating causal large language models, constrained decoding, and a deterministic index-based reconstruction algorithm to strictly preserve the original word order and frequency without loss. Evaluated on the UNMIXBENCH benchmark, EPOR reduces the error rate of a 4B-parameter model by a significant 22.3%, achieving performance comparable to state-of-the-art zero-shot large language models.
📝 Abstract
Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.