🤖 AI Summary
This work proposes a fully deterministic and efficient retrieval method for conversational memory systems, circumventing the need for complex LLM-based structured processing and learned retrieval mechanisms that incur substantial computational overhead. By eliminating structured preprocessing, the approach leverages named entity recognition–weighted substring matching, rule-driven multi-hop entity expansion, and a lightweight fusion of CrossEncoder and ColBERT reranking to achieve high recall and precision entirely on CPU. A novel score-adaptive truncation strategy is introduced to significantly reduce token consumption. The method attains 93.5% and 88.4% performance on the LoCoMo and LongMemEval-S benchmarks, respectively—surpassing all known memory systems—while reducing token usage by 8.5× compared to full-context baselines.
📝 Abstract
Recent conversational memory systems invest heavily in LLM-based structuring at ingestion time and learned retrieval policies at query time. We show that neither is necessary. SmartSearch retrieves from raw, unstructured conversation history using a fully deterministic pipeline: NER-weighted substring matching for recall, rule-based entity discovery for multi-hop expansion, and a CrossEncoder+ColBERT rank fusion stage -- the only learned component -- running on CPU in ~650ms. Oracle analysis on two benchmarks identifies a compilation bottleneck: retrieval recall reaches 98.6%, but without intelligent ranking only 22.5% of gold evidence survives truncation to the token budget. With score-adaptive truncation and no per-dataset tuning, SmartSearch achieves 93.5% on LoCoMo and 88.4% on LongMemEval-S, exceeding all known memory systems under the same evaluation protocol on both benchmarks while using 8.5x fewer tokens than full-context baselines.