Information Retrieval in long documents: Word clustering approach for improving Semantics

📅 2023-02-20
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF

career value

240K/year
🤖 AI Summary
To address the semantic gap in long-document semantic retrieval—where conventional keyword-based methods neglect lexical semantics—this paper proposes a dual-representation vector space model integrating lexical and semantic features. Our method introduces three key innovations: (1) a semantic word clustering algorithm grounded in pre-trained word embeddings to automatically identify semantically cohesive term groups; (2) a cluster-level weighting scheme that fuses intra-cluster term frequency with an enhanced TF-IDF variant; and (3) a joint lexical-semantic matching function that preserves exact lexical matching capability while substantially improving semantic coverage. Experimental evaluation on SQuAD and TREC-CAR demonstrates statistically significant improvements over pure keyword baselines: semantic retrieval accuracy increases markedly, without degrading traditional lexical retrieval performance. The approach thus achieves a balanced trade-off between precision and semantic robustness in long-document retrieval.
📝 Abstract
In this paper, we propose an alternative to deep neural networks for semantic information retrieval for the case of long documents. This new approach exploiting clustering techniques to take into account the meaning of words in Information Retrieval systems targeting long as well as short documents. This approach uses a specially designed clustering algorithm to group words with similar meanings into clusters. The dual representation (lexical and semantic) of documents and queries is based on the vector space model proposed by Gerard Salton in the vector space constituted by the formed clusters. The originalities of our proposal are at several levels: first, we propose an efficient algorithm for the construction of clusters of semantically close words using word embedding as input, then we define a formula for weighting these clusters, and then we propose a function allowing to combine efficiently the meanings of words with a lexical model widely used in Information Retrieval. The evaluation of our proposal in three contexts with two different datasets SQuAD and TREC-CAR has shown that is significantly improves the classical approaches only based on the keywords without degrading the lexical aspect.
Problem

Research questions and friction points this paper is trying to address.

Improving semantic information retrieval in long documents
Using word clustering to enhance word meaning representation
Combining lexical and semantic models for better retrieval accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Word clustering for semantic retrieval
Dual lexical-semantic document representation
Weighted clusters with embedding inputs
🔎 Similar Papers
No similar papers found.
P
Paul Mbate Mekontchou
Department of Computer Science, University of Yaoundé I - Cameroon
A
Armel Fotsoh
Sweez - R&D Department, Paris - France
B
Bernabé Batchakui
Department of Computer Science, University of Yaoundé I - Cameroon
E
Eddy Ella
Izysoft - IT Department, Yaoundé - Sydney