llm-assisted ranking and explanation

Designs and implements systems that use large language models to rank, match, and search candidate items and to produce human‑readable explanations for those rankings and matches. This includes building relevance‑scoring and retrieval pipelines, generating interpretable placement or match explanations, and adding controls or interfaces that preserve end‑user or decision‑maker discretion over suggested items.

llm-assistedrankingandexplanation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Large Language Models for Information Retrieval: A Survey

Aug 14, 2023
YZ
Yutao Zhu
🏛️ Renmin University of China

Information retrieval (IR) faces persistent challenges including data scarcity, limited interpretability, and insufficient accuracy in generative response generation. To address these, this work systematically surveys recent advances in large language model (LLM)-enhanced IR, covering five core stages: query rewriting, retrieval, re-ranking, reading comprehension, and search agents. Methodologically, it integrates classical sparse retrieval (e.g., BM25), neural dense retrieval (e.g., DPR, ANCE), LLM-driven rewriting and re-ranking, generative reading comprehension, and chain-of-thought–enabled search agents. The paper makes three key contributions: (1) the first comprehensive, full-stack technical taxonomy of LLM-IR; (2) a unified classification framework for LLM-augmented IR techniques; and (3) identification of a novel co-evolutionary paradigm—where sparse retrieval efficiency and neural semantic understanding mutually reinforce one another. Synthesizing over 100 studies, it clarifies critical bottlenecks and open problems, delivering a foundational, theoretically grounded yet practically actionable roadmap for LLM-enhanced IR.

Addressing challenges like data scarcity and response accuracy in IRLeveraging large language models to improve information retrieval systemsSurveying integration of LLMs in query, retrieval, and ranking components

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the internal mechanisms by which large language models (LLMs) perform relevance judgment in information retrieval. Addressing the open question—“How do LLMs understand and model query-document relevance?”—we propose the first mechanism-based, multi-stage interpretability framework. Our analysis reveals that early layers extract semantic features, intermediate layers activate relevance-specific reasoning pathways conditioned on instructions, and late-layer attention heads generate structured relevance judgments. Methodologically, we integrate activation patching with layer- and head-level attribution analysis to precisely localize functional roles across model components. Empirical results demonstrate that LLMs possess explicit, stage-wise relevance modeling capabilities—not merely opaque matching. This study uncovers an interpretable cognitive architecture underlying LLM-based IR and provides both theoretical foundations and design principles for developing trustworthy, controllable LLM-driven retrieval systems.

How LLMs internally assess relevance for IR tasksMechanistic process of generating relevance judgmentsRoles of different LLM modules in relevance judgment

This work identifies systematic biases in using large language models (LLMs) for information retrieval (IR) evaluation: LLM-based evaluators exhibit pronounced “source preference” toward LLM-generated rankings, struggle to discern fine-grained performance differences, and are susceptible to artifacts from AI assistant outputs—yet show no inherent bias against AI-generated content. To rigorously characterize these biases, the study employs a multi-model collaborative experimental design, controlled prompt engineering, and human calibration—constituting the first empirical validation of such phenomena. Based on these findings, the authors propose an integrated LLM-IR ecosystem evaluation framework, accompanied by a reproducible bias diagnostic protocol and a structured research roadmap. This advances IR evaluation toward greater reliability, transparency, and methodological rigor in LLM-driven systems. (132 words)

Assesses LLM judges' bias towards LLM rankersExamines biases in LLM-based IR evaluation componentsExplores limitations in LLM judges' performance discernment

Explainability of Text Processing and Retrieval Methods: A Critical Survey

Dec 14, 2022
SS
Sourav Saha
🏛️ Indian Statistical Institute

Deep learning models achieve state-of-the-art performance in NLP and information retrieval, yet their opacity severely hinders trustworthy deployment. This paper presents the first systematic, cross-model (word embeddings, RNNs/LSTMs, Transformers, BERT) and cross-task (text classification, question answering, document ranking) survey of interpretability methods in NLP/IR. We propose a structured taxonomy covering major paradigms—including feature attribution (e.g., LIME, SHAP), attention analysis, surrogate modeling, saliency mapping, and counterfactual explanation. Our framework constitutes the most comprehensive synthesis of textual interpretability techniques to date. We rigorously identify critical limitations—particularly the lack of standardized evaluation protocols and insufficient task-specific adaptation—and highlight key research gaps. The work establishes both theoretical foundations and practical guidelines for developing interpretable, reliable NLP systems.

Addressing non-linear model inscrutability in NLP and IRReviewing interpretability techniques for transformers and ranking modelsSurveying explainability methods for deep learning text processing

Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval

Mar 27, 2024
SM
Shengjie Ma
🏛️ Renmin University of China | Huawei | Beijing Normal University

Legal case retrieval suffers from heavy reliance on expert judgment for relevance assessment, poor interpretability, and low efficiency. Method: This paper proposes a few-shot, multi-stage large language model (LLM) reasoning framework that emulates the incremental reasoning process of human legal experts. It integrates legal fact extraction, expert-aligned fine-tuning, and annotation-based knowledge distillation to transfer capabilities from large models to compact ones. Contribution/Results: We introduce the first domain-specific, multi-stage legal reasoning paradigm, balancing interpretability with high-fidelity annotation alignment. Experiments show κ > 0.82 agreement between the framework’s outputs and human expert annotations. With only minimal expert labeling, a lightweight model achieves 92% of the performance of its large-model counterpart, substantially reducing domain adaptation costs while preserving legal reasoning fidelity.

Automating legal case relevance judgments using LLMsEnhancing accuracy via expert-aligned few-shot LLM approachImproving interpretability of legal case similarity data

The effectiveness and applicability boundaries of large language models (LLMs) in recommendation tasks remain poorly understood. Method: We propose a unified prompt engineering framework that reformulates recommendation as natural language inference, enabling zero-shot and cross-scenario generalization. We conduct controlled, multi-dimensional experiments on MovieLens and Amazon datasets to isolate the independent effects of LLM architecture, parameter scale, context length, and four prompt components—task description, user interest modeling, candidate item construction, and prompting strategy. Contribution/Results: Our study establishes a reproducible evaluation paradigm and demonstrates that LLMs possess intrinsic zero-shot recommendation capability. However, prompt quality and fidelity of user interest modeling constitute critical bottlenecks. Structurally optimizing prompts yields substantial performance gains. This work provides both an empirically grounded benchmark and a practical, deployable technical pathway for LLM-based recommender systems.

Large Language ModelsPerformance EvaluationRecommendation Systems

Latest Papers

What's happening recently
View more

This work addresses the challenge of uniformly supporting search, recommendation, and reasoning tasks over large-scale heterogeneous product catalogs by proposing NEO, a framework that adapts decoder-only large language models into an end-to-end system capable of directly generating real products without external tools. NEO interleaves natural language with typed product identifiers (SIDs) in a unified sequence and treats SIDs as an independent modality through a language-guided mechanism. By combining staged alignment and instruction tuning, it enables controllable generation across tasks, entity types, and output formats. Experiments on a catalog containing tens of millions of products demonstrate that NEO significantly outperforms strong baselines across multiple tasks and exhibits exceptional cross-task transfer capabilities.

entity groundinglarge language modelsreasoning

This work proposes a fine-grained evaluation framework to assess the capability of large language models (LLMs) as relevance judges in information retrieval, extending beyond holistic document-level judgments to identify the specific textual spans that support those judgments. Leveraging a Wikipedia test collection derived from INEX, the study employs prompt engineering to guide LLMs in simultaneously performing document-level relevance assessment and span-level annotation, followed by comparative analysis against human annotations. By introducing fine-grained relevance evaluation into the LLMs-as-Judges paradigm, this research is the first to examine whether models are “right for the right reasons,” thereby substantially enhancing the credibility of automated evaluation. Experimental results demonstrate that, under human supervision, LLMs can accurately identify both relevant documents and the key evidence spans within them.

Fine-grained EvaluationInformation RetrievalLLMs-as-Judges

"You Are Rejected!": An Empirical Study of Large Language Models Taking Hiring Evaluations

Oct 21, 2025
DF
Dingjie Fu
🏛️ Huazhong University of Science and Technology | Independent Researcher

This study empirically evaluates large language models (LLMs) against industry-standard technical hiring assessments for algorithm and software engineering roles. Method: We administered realistic, industrial-grade programming, system design, and reasoning questions—commonly used by leading technology firms—to state-of-the-art LLMs (e.g., GPT-4, Claude 3, Gemini) and conducted multi-stage comparative analysis against official corporate reference solutions, assessing correctness, completeness, engineering soundness, and consistency. Contribution/Results: Our analysis reveals systematic structural gaps between LLM outputs and industrial expectations: no tested model met enterprise hiring thresholds. Critical deficiencies were observed in boundary-case handling, explicit modeling of resource constraints (e.g., time/space complexity, scalability), and maintainability-aware design. These findings challenge the prevailing assumption that LLMs can directly substitute for entry-level engineers. Moreover, this work introduces the first benchmark framework specifically tailored to industrial recruitment scenarios, providing empirically grounded insights for AI capability evaluation in real-world engineering hiring.

Assessing performance consistency between models and company standardsInvestigating LLMs' ability to pass hiring evaluationsRevealing LLMs' failure in professional competency assessments

Generating Search Explanations using Large Language Models

Jul 22, 2025
AL
Arif Laksito
🏛️ University of Sheffield

This study systematically investigates, for the first time, the capability of large language models (LLMs) to generate aspect-oriented search explanations (AOSE)—concise, interpretable rationales that enhance users’ comprehension efficiency and information localization speed. To address the low factual accuracy and poor user acceptability of conventional explanation methods, we comparatively evaluate two dominant LLM architectures—encoder-decoder models (e.g., T5) and decoder-only models (e.g., LLaMA, ChatGLM)—on the AOSE task. Experimental results demonstrate that the best-performing LLMs significantly outperform multiple baselines—including retrieval-augmented and rule-based approaches—in factual accuracy, logical coherence, and user acceptability. Our key contributions are: (1) formalizing and empirically validating AOSE as a novel search explanation paradigm; (2) identifying LLM architecture choice as a critical determinant of explanation quality; and (3) introducing the first LLM-oriented benchmark specifically designed for search explanation generation.

Comparing encoder-decoder and decoder-only LLM performanceExploring LLMs for generating search result explanationsImproving explanation accuracy over baseline models

To address the high cost and low efficiency of manually constructing categorical features in large-scale recommender systems, this paper proposes a large language model (LLM)-driven multi-agent collaborative framework for automatically extracting high-quality, multivalued categorical features from unstructured text. The method integrates dynamic feedback–guided prompt engineering with an oracle-evaluated closed-loop optimization mechanism, enabling automated feature discovery, generation, and validation. Experiments demonstrate that the approach significantly improves recommender model performance (e.g., +2.3% average AUC gain), accelerates feature construction by 5.8×, and reduces human intervention by over 90%. Its core innovation lies in unifying multi-LLM agent collaboration, dynamic feedback loops, and evaluable automated prompt tuning within a single feature engineering closed loop—establishing a scalable, interpretable, end-to-end paradigm for categorical feature construction in recommender systems.

Automate feature extraction from unstructured text for recommendationsDevelop automated prompt-engineering tuning with dynamic feedbackIdentify valuable text aspects for models and AutoML pipelines

Hot Scholars

RM

Robert Moro

Senior Researcher at Kempelen Institute of Intelligent Technologies
Artificial IntelligenceMachine LearningUser ModelingPersonalization
JS

Jakub Simko

Expert researcher, Kempelen Institute of Intelligent Technologies
user modellingdata analysismachine learningcrowdsourcing
YX

Yuchen Xiao

Lead of Embodied AI R&D, Unitree | Research Scientist, J.P. Morgan | Ph.D. Northeastern University
Generative ModelsRobot LearningReinforcement LearningMulti-Agent Systems
HL

Horst Lichter

Professor at RWTH Aachen University
Software EngineeringSoftware Quality AssuranceDevelopment ProcessesArchitecture Evolution
DM

Deyu Meng

Professor, Xi'an Jiaotong University
Machine LearningApplied MathematicsComputer VisionArtificial Intelligence