Score
Designs and implements systems that use one or more LLMs to score, rank, and accept or reject candidate items (such as candidate paths, retrieved passages, or pieces of evidence) by enforcing relevance, temporal or system context and removing noise or staleness; this includes prompt designs and ensemble-consensus techniques to produce compact contexts for downstream models. Builds evaluation and aggregation procedures to combine multi-LLM judgments, calibrate confidence, and tune the trade-off between strict filtering, recall, and context size.
The reliability of LLM-as-a-Judge—particularly concerning evaluation accuracy, inter-annotator consistency, and fairness—remains a critical bottleneck. Method: This paper introduces the first systematic reliability assessment framework, featuring a dedicated benchmark (ReliBench) and synthesizing three core strategies: consistency enhancement, bias mitigation, and scenario adaptation. Our methodology integrates multi-model cross-validation, adversarial testing, interpretability analysis, and structured prompt engineering, underpinned by a standardized evaluation protocol. Contributions: (1) The first comprehensive, domain-wide survey of LLM-as-a-Judge reliability; (2) An open-source, reproducible evaluation toolkit and benchmark; (3) A paradigm shift from heuristic, experience-driven LLM evaluation toward rigorous, scientific, and standardized assessment practices.
研究提出了一种增量池化LLM评估方法,通过重用先前的文档判断来有效降低新检索模型选择的成本。
Large language models (LLMs) exhibit positional bias—where candidate ordering influences ranking/evaluation outcomes—and low repetition consistency—yielding unstable predictions for identical inputs—thereby undermining reliability. To address these issues, we propose a dynamic repetition strategy featuring the first confidence-driven early-stopping mechanism: for each input instance, it adaptively estimates the minimal required number of repetitions, then integrates majority voting with explicit positional bias modeling for fine-grained correction. Unlike static repetition schemes, our method eliminates the need for pre-specified repetition counts. We validate it across three LLM scales and two distinct task categories. Results show that our approach reduces average model calls by 81%–87% compared to static repetition, while preserving high ranking accuracy. This yields significant improvements in both computational efficiency and robustness without sacrificing performance.
This paper identifies validity risks in using large language models (LLMs) for evaluating information retrieval (IR) systems: when LLM-based assessments simultaneously guide system development and performance evaluation, they risk reinforcing biases, undermining reproducibility, and introducing methodological inconsistency—leading to spurious success claims and misleading conclusions. To address this, the authors propose a verifiable risk analysis framework comprising (1) quantitative detection methods for three core validity threats, (2) lightweight mitigation guardrails, and (3) a human-in-the-loop paradigm for constructing reusable, auditable test collections. Grounded in empirical analysis, assessment validity theory, and cross-institutional collaboration, the work delivers an open-source validation toolkit and an industry consensus guideline. These contributions establish responsible, reproducible, and accountable best practices for LLM-augmented IR evaluation.
Pointwise large language model (LLM) rankers suffer from limited adherence to standardized comparative guidelines and insufficient capability in holistically evaluating complex passages. To address this, we propose a dynamic multi-perspective evaluation criterion generation method: leveraging prompt engineering to instantiate interpretable, dimension-specific criteria—covering semantics, relevance, structure, and more—in real time, and jointly aggregating scores across these criteria. This mechanism is the first to achieve decomposability, interpretability, and synergistic enhancement in LLM-based evaluation. Evaluated on the BEIR benchmark across eight diverse datasets, our approach significantly improves ranking performance, yielding an average 3.2% relative gain in NDCG@10. Results demonstrate that dynamic, multi-perspective guidance effectively enhances the ranking capability of pointwise LLM rankers.
To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.
Large language models are susceptible to selection bias in adaptive prompting and program search, leading to an overestimation of the winning candidate’s performance under real-world deployment. This work proposes the SIREN protocol, which enables unbiased performance inference for the full tuning-to-deployment pipeline under a fixed tuning budget by freezing the candidate set, decoupling selection and evaluation data, and incorporating an entry-wise Gaussian multiplier bootstrap. SIREN is the first method to simultaneously support accurate estimation of program-level performance curves on limited-budget grids and construct confidence intervals for both within-budget and cross-budget comparisons. Empirical results demonstrate that conventional winner-reporting practices exhibit substantial optimistic bias, whereas SIREN closely approximates the true evaluation target under finite-sample conditions, offering reliable guidance for deployment decisions.
This study challenges the prevailing assumption that iterative updates to large language models (LLMs) inherently enhance consistency in relevance judgment. Specifically, it investigates how backbone network evolution affects the stability of LLM-based evaluation. Employing the UMBRELA single-prompt and EXAM criteria-prompting methodologies, this work systematically assesses multiple generations of commercial and open-source models—including Gemini, GPT, Qwen, and Llama—under fixed prompting conditions. The findings reveal that improvements in aggregated performance do not necessarily translate to greater judgment stability. Notably, newer model versions fail to significantly improve relevance judgment quality and exhibit regression phenomena, wherein correct assessments made by older versions are lost in subsequent iterations. By exposing the latent risks associated with cross-version prompt transferability, this research provides a critical cautionary insight for LLM evaluation practices.
This study reveals the performance degradation and evaluation bias of large language models (LLMs) in decision-making tasks involving large-scale candidate sets. By establishing candidate set size as a critical evaluation variable, we demonstrate that performance advantages observed in small-scale settings do not guarantee robustness at scale, identifying confidence collapse and positional bias as two primary failure modes. To address these challenges, this work proposes a hierarchical partitioning strategy and a permutation-based reasoning method. Experimental results indicate that these approaches effectively mitigate performance degradation, improving accuracy by approximately 20 percentage points when the number of candidates reaches N=160. Ultimately, this research establishes a new paradigm for the reliable application and rigorous evaluation of LLMs in large-scale decision-making scenarios.
This study investigates the impact of user history length on recommendation quality in large language model (LLM)-based recommender systems, challenging the common assumption that more context yields better performance. Using four state-of-the-art models—GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flash—we conduct within-user experiments on the REGEN dataset to evaluate the effectiveness of contextual histories ranging from 5 to 50 items, while also measuring inference latency. Our results demonstrate that increasing context length from 5 to 50 items does not significantly improve recommendation performance, with NDCG scores remaining stable between 0.17 and 0.23. Notably, as few as 5–10 historical items suffice to maintain recommendation quality while reducing inference cost by approximately 88%. This work provides the first empirical evidence of the efficiency of short-context inputs in LLM-based recommendation, offering a practical foundation for low-cost deployment.
Traditional peer review faces scalability bottlenecks, while large language model (LLM)-driven automated review lacks systematic investigation into its reliability, robustness, and security. This work addresses this gap by offering the first system-oriented analysis, focusing on two core tasks: critique generation and score prediction. It establishes a taxonomy of LLM-based reviewing approaches and comprehensively evaluates key technical strategies, including prompt engineering, supervised fine-tuning, retrieval augmentation, and alignment optimization. The study uncovers emerging security threats such as prompt injection and data poisoning, examines challenges arising from subjective disagreement and cross-domain generalization, and highlights limitations and domain biases in current benchmarks. Building on these insights, the paper proposes a roadmap toward developing reliable, transparent, and trustworthy AI-assisted scientific review systems.