Score
Designs, implements, and evaluates systems that index, retrieve, and rank text documents in response to user queries, covering tasks such as full‑text indexing, tokenization, query parsing, relevance modeling, and result presentation. Builds and analyzes search pipelines and their engineering concerns — inverted indexes, ranking algorithms, query optimization, scalability and distribution, latency, and retrieval effectiveness metrics.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This work addresses the challenge of efficiently constructing and querying inverted indexes over large-scale text corpora by designing and implementing a high-performance, memory-safe, and scalable inverted index library in Rust. Leveraging Rust’s zero-cost abstractions, concurrency safety guarantees, and expressive trait-based generics, the system flexibly integrates multiple classical inverted indexing techniques and supports efficient full-text retrieval algorithms. Experimental evaluation across several standard datasets and query workloads demonstrates that the proposed library achieves up to twice the query performance of state-of-the-art alternatives, substantially enhancing both retrieval efficiency and practical applicability.
Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.
Existing retrieval evaluation benchmarks predominantly rely on simple, single-point queries, failing to reflect model capabilities under realistic, complex retrieval scenarios involving multiple constraints and intents. Method: We introduce ComplexRetrieval-Bench—the first systematic, diverse, and realistic benchmark for complex retrieval tasks—covering multi-condition filtering, multi-hop reasoning, and natural-language constraints. Contribution/Results: Our benchmark reveals severe performance degradation of state-of-the-art retrieval models under complex queries (average nDCG@10 = 0.346, R@100 = 0.587). Notably, LLM-based query rewriting—widely assumed beneficial—degrades performance even for strong retrievers, challenging prevailing assumptions. Extensive experiments across modern retrieval architectures (e.g., dense, sparse, hybrid) and LLM-augmented strategies provide reproducible evaluation protocols and critical insights for next-generation general-purpose retrieval models.
This study addresses the lack of systematic evaluation regarding the impact of document selection strategies in query-focused text analysis, a gap that has led to ad hoc methodological choices. It establishes document selection as a critical methodological decision rather than merely a computational compromise and systematically evaluates seven selection strategies—ranging from random sampling to semantic and hybrid retrieval—across four prominent topic modeling approaches: LDA, BERTopic, TopicGPT, and HiCode. Experiments conducted on two datasets involving 26 open-ended queries demonstrate that semantic and hybrid retrieval strategies consistently achieve a robust balance between output quality and computational efficiency, warranting their recommendation as default choices for query-driven text analysis.
Deep learning models achieve state-of-the-art performance in NLP and information retrieval, yet their opacity severely hinders trustworthy deployment. This paper presents the first systematic, cross-model (word embeddings, RNNs/LSTMs, Transformers, BERT) and cross-task (text classification, question answering, document ranking) survey of interpretability methods in NLP/IR. We propose a structured taxonomy covering major paradigms—including feature attribution (e.g., LIME, SHAP), attention analysis, surrogate modeling, saliency mapping, and counterfactual explanation. Our framework constitutes the most comprehensive synthesis of textual interpretability techniques to date. We rigorously identify critical limitations—particularly the lack of standardized evaluation protocols and insufficient task-specific adaptation—and highlight key research gaps. The work establishes both theoretical foundations and practical guidelines for developing interpretable, reliable NLP systems.
This work addresses the architectural challenges faced by industrial-scale web retrieval systems under stringent constraints of latency, scalability, and resource efficiency. It proposes a unified multi-stage abstraction termed “Retrieval-as-a-Service” (RaaS), which, for the first time, integrates infrastructure-aware components—including efficient candidate generation, embedding-based semantic matching, and resource-conscious re-ranking—into a cohesive framework. The study systematically models the impact of incorporating large language models (LLMs) on both system performance and operational overhead. By analyzing real-world production deployments, the authors uncover fundamental trade-offs between system design choices and quality-of-service (QoS) objectives, thereby offering practical, scalable, and QoS-aware architectural guidelines for building high-performance web-scale retrieval systems.
Existing RAG systems rely heavily on heuristic configurations, lacking systematic evaluation and reproducibility. This work formalizes RAG design as an architecture search problem and introduces RAISE, a unified benchmark that establishes a standardized framework to enable controlled and reproducible hyperparameter optimization research. Within a standardized search space and computational budget, we integrate 13 search algorithms and conduct comprehensive experiments across seven textual and multimodal datasets. Our results demonstrate that the effectiveness of optimization strategies is highly task-dependent, with no single method consistently outperforming others across all settings. These findings caution against drawing conclusions about universal superiority based on aggregated rankings, underscoring the necessity of task-specific evaluation in RAG system design.
This work addresses the limitation of traditional retrieval systems, which treat document representation as a static preprocessing step and thus struggle to adapt to downstream tasks. The authors propose AutoIndex, a novel framework that formulates document representation construction as a learnable program synthesis problem. AutoIndex dynamically generates retrieval-oriented representations by searching over executable transformation programs—such as slicing, augmentation, and normalization—and iteratively refines them using validation feedback. By integrating proxy-guided program search with retrieval quality evaluation, the method enables explicit optimization of document representations. Evaluated on the CRUMB benchmark across all eight tasks, AutoIndex consistently outperforms the full-document BM25 baseline, achieving average improvements of 8.4% in Recall@100 and 8.3% in nDCG@10, with peak gains reaching 30.5% and 43.6%, respectively.