Score
Designs and implements systems that use large language models to rank, match, and search candidate items and to produce human‑readable explanations for those rankings and matches. This includes building relevance‑scoring and retrieval pipelines, generating interpretable placement or match explanations, and adding controls or interfaces that preserve end‑user or decision‑maker discretion over suggested items.
Current information retrieval (IR) systems face challenges including shallow semantic understanding, limited reasoning capabilities, and low decision interpretability. Method: This work systematically reviews large language model (LLM)-powered agents for search and recommendation, proposing the first IR-oriented LLM agent taxonomy grounded in three dimensions: role definition, capability boundaries, and evolutionary paradigms. It integrates key techniques—including multi-agent coordination, tool-augmented execution, memory enhancement, reflective reasoning, and retrieval-augmented generation (RAG). Contribution/Results: We introduce the first unified, IR-specific classification framework for LLM agents; establish an open-source literature index repository on GitHub; and provide both theoretical foundations and practical guidelines for developing next-generation IR systems that are interpretable, adaptive, and task-driven.
Information retrieval (IR) faces persistent challenges including data scarcity, limited interpretability, and insufficient accuracy in generative response generation. To address these, this work systematically surveys recent advances in large language model (LLM)-enhanced IR, covering five core stages: query rewriting, retrieval, re-ranking, reading comprehension, and search agents. Methodologically, it integrates classical sparse retrieval (e.g., BM25), neural dense retrieval (e.g., DPR, ANCE), LLM-driven rewriting and re-ranking, generative reading comprehension, and chain-of-thought–enabled search agents. The paper makes three key contributions: (1) the first comprehensive, full-stack technical taxonomy of LLM-IR; (2) a unified classification framework for LLM-augmented IR techniques; and (3) identification of a novel co-evolutionary paradigm—where sparse retrieval efficiency and neural semantic understanding mutually reinforce one another. Synthesizing over 100 studies, it clarifies critical bottlenecks and open problems, delivering a foundational, theoretically grounded yet practically actionable roadmap for LLM-enhanced IR.
This work investigates the internal mechanisms by which large language models (LLMs) perform relevance judgment in information retrieval. Addressing the open question—“How do LLMs understand and model query-document relevance?”—we propose the first mechanism-based, multi-stage interpretability framework. Our analysis reveals that early layers extract semantic features, intermediate layers activate relevance-specific reasoning pathways conditioned on instructions, and late-layer attention heads generate structured relevance judgments. Methodologically, we integrate activation patching with layer- and head-level attribution analysis to precisely localize functional roles across model components. Empirical results demonstrate that LLMs possess explicit, stage-wise relevance modeling capabilities—not merely opaque matching. This study uncovers an interpretable cognitive architecture underlying LLM-based IR and provides both theoretical foundations and design principles for developing trustworthy, controllable LLM-driven retrieval systems.
This work identifies systematic biases in using large language models (LLMs) for information retrieval (IR) evaluation: LLM-based evaluators exhibit pronounced “source preference” toward LLM-generated rankings, struggle to discern fine-grained performance differences, and are susceptible to artifacts from AI assistant outputs—yet show no inherent bias against AI-generated content. To rigorously characterize these biases, the study employs a multi-model collaborative experimental design, controlled prompt engineering, and human calibration—constituting the first empirical validation of such phenomena. Based on these findings, the authors propose an integrated LLM-IR ecosystem evaluation framework, accompanied by a reproducible bias diagnostic protocol and a structured research roadmap. This advances IR evaluation toward greater reliability, transparency, and methodological rigor in LLM-driven systems. (132 words)
Deep learning models achieve state-of-the-art performance in NLP and information retrieval, yet their opacity severely hinders trustworthy deployment. This paper presents the first systematic, cross-model (word embeddings, RNNs/LSTMs, Transformers, BERT) and cross-task (text classification, question answering, document ranking) survey of interpretability methods in NLP/IR. We propose a structured taxonomy covering major paradigms—including feature attribution (e.g., LIME, SHAP), attention analysis, surrogate modeling, saliency mapping, and counterfactual explanation. Our framework constitutes the most comprehensive synthesis of textual interpretability techniques to date. We rigorously identify critical limitations—particularly the lack of standardized evaluation protocols and insufficient task-specific adaptation—and highlight key research gaps. The work establishes both theoretical foundations and practical guidelines for developing interpretable, reliable NLP systems.
Legal case retrieval suffers from heavy reliance on expert judgment for relevance assessment, poor interpretability, and low efficiency. Method: This paper proposes a few-shot, multi-stage large language model (LLM) reasoning framework that emulates the incremental reasoning process of human legal experts. It integrates legal fact extraction, expert-aligned fine-tuning, and annotation-based knowledge distillation to transfer capabilities from large models to compact ones. Contribution/Results: We introduce the first domain-specific, multi-stage legal reasoning paradigm, balancing interpretability with high-fidelity annotation alignment. Experiments show κ > 0.82 agreement between the framework’s outputs and human expert annotations. With only minimal expert labeling, a lightweight model achieves 92% of the performance of its large-model counterpart, substantially reducing domain adaptation costs while preserving legal reasoning fidelity.
The effectiveness and applicability boundaries of large language models (LLMs) in recommendation tasks remain poorly understood. Method: We propose a unified prompt engineering framework that reformulates recommendation as natural language inference, enabling zero-shot and cross-scenario generalization. We conduct controlled, multi-dimensional experiments on MovieLens and Amazon datasets to isolate the independent effects of LLM architecture, parameter scale, context length, and four prompt components—task description, user interest modeling, candidate item construction, and prompting strategy. Contribution/Results: Our study establishes a reproducible evaluation paradigm and demonstrates that LLMs possess intrinsic zero-shot recommendation capability. However, prompt quality and fidelity of user interest modeling constitute critical bottlenecks. Structurally optimizing prompts yields substantial performance gains. This work provides both an empirically grounded benchmark and a practical, deployable technical pathway for LLM-based recommender systems.
This work addresses the challenge of uniformly supporting search, recommendation, and reasoning tasks over large-scale heterogeneous product catalogs by proposing NEO, a framework that adapts decoder-only large language models into an end-to-end system capable of directly generating real products without external tools. NEO interleaves natural language with typed product identifiers (SIDs) in a unified sequence and treats SIDs as an independent modality through a language-guided mechanism. By combining staged alignment and instruction tuning, it enables controllable generation across tasks, entity types, and output formats. Experiments on a catalog containing tens of millions of products demonstrate that NEO significantly outperforms strong baselines across multiple tasks and exhibits exceptional cross-task transfer capabilities.
This work proposes a fine-grained evaluation framework to assess the capability of large language models (LLMs) as relevance judges in information retrieval, extending beyond holistic document-level judgments to identify the specific textual spans that support those judgments. Leveraging a Wikipedia test collection derived from INEX, the study employs prompt engineering to guide LLMs in simultaneously performing document-level relevance assessment and span-level annotation, followed by comparative analysis against human annotations. By introducing fine-grained relevance evaluation into the LLMs-as-Judges paradigm, this research is the first to examine whether models are “right for the right reasons,” thereby substantially enhancing the credibility of automated evaluation. Experimental results demonstrate that, under human supervision, LLMs can accurately identify both relevant documents and the key evidence spans within them.
This study empirically evaluates large language models (LLMs) against industry-standard technical hiring assessments for algorithm and software engineering roles. Method: We administered realistic, industrial-grade programming, system design, and reasoning questions—commonly used by leading technology firms—to state-of-the-art LLMs (e.g., GPT-4, Claude 3, Gemini) and conducted multi-stage comparative analysis against official corporate reference solutions, assessing correctness, completeness, engineering soundness, and consistency. Contribution/Results: Our analysis reveals systematic structural gaps between LLM outputs and industrial expectations: no tested model met enterprise hiring thresholds. Critical deficiencies were observed in boundary-case handling, explicit modeling of resource constraints (e.g., time/space complexity, scalability), and maintainability-aware design. These findings challenge the prevailing assumption that LLMs can directly substitute for entry-level engineers. Moreover, this work introduces the first benchmark framework specifically tailored to industrial recruitment scenarios, providing empirically grounded insights for AI capability evaluation in real-world engineering hiring.
This study systematically investigates, for the first time, the capability of large language models (LLMs) to generate aspect-oriented search explanations (AOSE)—concise, interpretable rationales that enhance users’ comprehension efficiency and information localization speed. To address the low factual accuracy and poor user acceptability of conventional explanation methods, we comparatively evaluate two dominant LLM architectures—encoder-decoder models (e.g., T5) and decoder-only models (e.g., LLaMA, ChatGLM)—on the AOSE task. Experimental results demonstrate that the best-performing LLMs significantly outperform multiple baselines—including retrieval-augmented and rule-based approaches—in factual accuracy, logical coherence, and user acceptability. Our key contributions are: (1) formalizing and empirically validating AOSE as a novel search explanation paradigm; (2) identifying LLM architecture choice as a critical determinant of explanation quality; and (3) introducing the first LLM-oriented benchmark specifically designed for search explanation generation.
To address the high cost and low efficiency of manually constructing categorical features in large-scale recommender systems, this paper proposes a large language model (LLM)-driven multi-agent collaborative framework for automatically extracting high-quality, multivalued categorical features from unstructured text. The method integrates dynamic feedback–guided prompt engineering with an oracle-evaluated closed-loop optimization mechanism, enabling automated feature discovery, generation, and validation. Experiments demonstrate that the approach significantly improves recommender model performance (e.g., +2.3% average AUC gain), accelerates feature construction by 5.8×, and reduces human intervention by over 90%. Its core innovation lies in unifying multi-LLM agent collaboration, dynamic feedback loops, and evaluable automated prompt tuning within a single feature engineering closed loop—establishing a scalable, interpretable, end-to-end paradigm for categorical feature construction in recommender systems.