SmartSearch: How Ranking Beats Structure for Conversational Memory Retrieval

📅 2026-03-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a fully deterministic and efficient retrieval method for conversational memory systems, circumventing the need for complex LLM-based structured processing and learned retrieval mechanisms that incur substantial computational overhead. By eliminating structured preprocessing, the approach leverages named entity recognition–weighted substring matching, rule-driven multi-hop entity expansion, and a lightweight fusion of CrossEncoder and ColBERT reranking to achieve high recall and precision entirely on CPU. A novel score-adaptive truncation strategy is introduced to significantly reduce token consumption. The method attains 93.5% and 88.4% performance on the LoCoMo and LongMemEval-S benchmarks, respectively—surpassing all known memory systems—while reducing token usage by 8.5× compared to full-context baselines.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalNatural Language Processing: Conversational AI/Dialog SystemsMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Recent conversational memory systems invest heavily in LLM-based structuring at ingestion time and learned retrieval policies at query time. We show that neither is necessary. SmartSearch retrieves from raw, unstructured conversation history using a fully deterministic pipeline: NER-weighted substring matching for recall, rule-based entity discovery for multi-hop expansion, and a CrossEncoder+ColBERT rank fusion stage -- the only learned component -- running on CPU in ~650ms. Oracle analysis on two benchmarks identifies a compilation bottleneck: retrieval recall reaches 98.6%, but without intelligent ranking only 22.5% of gold evidence survives truncation to the token budget. With score-adaptive truncation and no per-dataset tuning, SmartSearch achieves 93.5% on LoCoMo and 88.4% on LongMemEval-S, exceeding all known memory systems under the same evaluation protocol on both benchmarks while using 8.5x fewer tokens than full-context baselines.
Problem

Research questions and friction points this paper is trying to address.

conversational memory retrieval
unstructured conversation history
token budget
retrieval recall
evidence truncation
Innovation

Methods, ideas, or system contributions that make the work stand out.

conversational memory retrieval
deterministic retrieval pipeline
score-adaptive truncation
CrossEncoder+ColBERT fusion
unstructured conversation history
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.