Semantic Entanglement in Vector-Based Retrieval: A Formal Framework and Context-Conditioned Disentanglement Pipeline for Agentic RAG Systems

📅 2026-04-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of semantic entanglement in vector retrieval, where multi-topic documents induce overlapping semantics in embedding space, degrading retrieval accuracy. The work formally defines semantic entanglement for the first time and introduces a quantifiable Entanglement Index (EI). To mitigate this issue, the authors propose a context-conditioned, four-stage Semantic Disentanglement Pipeline (SDP) that dynamically optimizes pre-embedding text organization through document restructuring and an agent-based feedback mechanism. Experimental evaluation on over 2,000 medical documents demonstrates that the approach substantially alleviates semantic entanglement, reducing the average EI from 0.71 to 0.14 and improving Top-K retrieval accuracy from 32% to 82%.

Technology Category

Natural Language Processing: Information ExtractionSearch and Optimization: Distributed SearchData Mining & Knowledge Management: Intelligent Query Processing

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Retrieval-Augmented Generation (RAG) systems depend on the geometric properties of vector representations to retrieve contextually appropriate evidence. When source documents interleave multiple topics within contiguous text, standard vectorization produces embedding spaces in which semantically distinct content occupies overlapping neighborhoods. We term this condition semantic entanglement. We formalize entanglement as a model-relative measure of cross-topic overlap in embedding space and define an Entanglement Index (EI) as a quantitative proxy. We argue that higher EI constrains attainable Top-K retrieval precision under cosine similarity retrieval. To address this, we introduce the Semantic Disentanglement Pipeline (SDP), a four-stage preprocessing framework that restructures documents prior to embedding. We further propose context-conditioned preprocessing, in which document structure is shaped by patterns of operational use, and a continuous feedback mechanism that adapts document structure based on agent performance. We evaluate SDP on a real-world enterprise healthcare knowledge base comprising over 2,000 documents across approximately 25 sub-domains. Top-K retrieval precision improves from approximately 32% under fixed-token chunking to approximately 82% under SDP, while mean EI decreases from 0.71 to 0.14. We do not claim that entanglement fully explains RAG failure, but that it captures a distinct preprocessing failure mode that downstream optimization cannot reliably correct once encoded into the vector space.
Problem

Research questions and friction points this paper is trying to address.

semantic entanglement
vector-based retrieval
Retrieval-Augmented Generation
embedding space
retrieval precision
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic entanglement
Entanglement Index
Semantic Disentanglement Pipeline
context-conditioned preprocessing
Retrieval-Augmented Generation
🔎 Similar Papers
No similar papers found.