Score
Designs and implements systems that convert documents into vector representations and the pipelines for their extraction, preprocessing, projection, storage, and retrieval. Also develops and analyzes methods for embedding model selection and learning, concatenation/fusion/projection/factorization/reconstruction, regularization and optimization, geometric and sensitivity analysis, evaluation and benchmarking (semantic relevance, similarity search), and other embedding manipulations for downstream tasks.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This study addresses the cross-modal retrieval challenge for multimodal educational content—particularly computer science textbooks containing interleaved text and figures. We propose a multi-vector representation method that jointly encodes textual and visual semantics using a vision-language model (VLM), generating fine-grained multimodal embeddings and indexing them in a vector database to enable efficient cross-modal retrieval. Evaluated on over 3,600 pages of textbook material, we systematically compare four similarity metrics and find cosine similarity significantly outperforms alternatives. Benchmarking against 75 natural-language queries confirms substantial improvements in retrieval precision and practical utility within digital library settings. Our approach delivers a reproducible, scalable technical framework for intelligent discovery of multimodal educational resources, advancing the state of cross-modal semantic search in academic and pedagogical contexts.
This work addresses structured clinical document classification, a task requiring accurate and efficient retrieval in domain-specific, highly standardized medical texts. Method: We comparatively evaluate lexical vectorization (TF-IDF with cosine similarity) against semantic embedding approaches (BERT and Sentence-BERT), assessing both classification accuracy and inference latency. Results: On constrained, format-regular clinical documents, a lightweight, domain-adapted lexical method achieves a marginal yet consistent accuracy gain (+0.8%) over neural semantic baselines while accelerating inference by 67%. These findings challenge the prevailing assumption that transformer-based semantic models inherently dominate such structured, professional text classification tasks. Contribution: The study demonstrates the sufficiency and efficiency of lexical matching in structured clinical documentation, establishing interpretable, low-overhead lexical methods as viable alternatives—particularly valuable in resource-constrained clinical settings where computational efficiency and model transparency are critical.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
Deep learning models achieve state-of-the-art performance in NLP and information retrieval, yet their opacity severely hinders trustworthy deployment. This paper presents the first systematic, cross-model (word embeddings, RNNs/LSTMs, Transformers, BERT) and cross-task (text classification, question answering, document ranking) survey of interpretability methods in NLP/IR. We propose a structured taxonomy covering major paradigms—including feature attribution (e.g., LIME, SHAP), attention analysis, surrogate modeling, saliency mapping, and counterfactual explanation. Our framework constitutes the most comprehensive synthesis of textual interpretability techniques to date. We rigorously identify critical limitations—particularly the lack of standardized evaluation protocols and insufficient task-specific adaptation—and highlight key research gaps. The work establishes both theoretical foundations and practical guidelines for developing interpretable, reliable NLP systems.
Traditional vector retrieval relies on pairwise geometric similarity, which struggles to simultaneously achieve semantic alignment and consistency with the head-tail distribution of data. This work proposes a Graph Wiring framework combined with Spectral Indexing, modeling the embedding space as an energy network induced by the topology of feature column vectors. By integrating geometric similarity with spectral structural information and introducing τ-modulation for adaptive retrieval, the method leverages spectral graph theory, energy-based modeling, and epiplexity analysis. Implemented using the open-source arrowspace library, it significantly outperforms purely geometric retrieval across multiple benchmarks and industrial applications, effectively enhancing both semantic alignment and distributional consistency to meet the demands of modern RAG systems for flexible and efficient retrieval.
Traditional vector retrieval systems exhibit high encapsulation, limiting the ability of AI agents to flexibly intervene in embedding generation and scoring during query time, thereby hindering fine-grained control over retrieval. To address this, this work proposes flexvec, a novel retrieval kernel featuring the first programmable embedding modulation (PEM) mechanism, which exposes embedding matrices and score arrays as programmable interfaces to enable dynamic, composable semantic manipulations at query time. By natively integrating with SQL, supporting runtime arithmetic operations, and leveraging query materialization, flexvec efficiently executes complex modulations without relying on approximate indexing: it completes three types of compositional operations end-to-end over 240k text chunks in just 19 milliseconds, and scales to 82 milliseconds on million-scale datasets—all on a desktop CPU.
This work addresses the limitations of current scientific document retrieval methods, which predominantly rely on document image representations and struggle to effectively leverage critical evidence embedded in structured content such as text, tables, and mathematical formulas. To this end, the authors introduce ArXivDoc, a novel benchmark constructed from LaTeX source code that enables controlled query generation, facilitating a systematic evaluation of textual, visual, and multimodal representations for retrieval. Experimental results demonstrate that textual representations consistently outperform others across diverse query types; multimodal approaches combining text and images achieve substantial gains over image-only methods without requiring specialized training; and image-based representations exhibit significant performance degradation with increasing document length, particularly for structured content. This study thus exposes fundamental shortcomings of image-centric paradigms and establishes a new foundation for advancing scientific document retrieval.
Traditional vector retrieval supports only a single query vector, limiting its effectiveness in complex reasoning and multi-example retrieval scenarios. This work proposes a novel multi-query vector retrieval method that introduces anomaly pattern detection into the task for the first time. By analyzing the consistency of anomalies across dimensions among multiple query vectors, the approach dynamically identifies discriminative dimensions and retrieves items from the vector database that exhibit similar anomalous patterns along those dimensions. Integrating multi-query embeddings, high-dimensional anomaly detection, and similarity analysis, the method achieves significant performance gains across image, text, and tabular datasets. Notably, retrieval effectiveness consistently improves as the number of query examples increases from one to eight.
This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.