llm-guided filtering

Designs and implements systems that use one or more LLMs to score, rank, and accept or reject candidate items (such as candidate paths, retrieved passages, or pieces of evidence) by enforcing relevance, temporal or system context and removing noise or staleness; this includes prompt designs and ensemble-consensus techniques to produce compact contexts for downstream models. Builds evaluation and aggregation procedures to combine multi-LLM judgments, calibrate confidence, and tune the trade-off between strict filtering, recall, and context size.

llm-guidedfiltering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Large language models (LLMs) exhibit positional bias—where candidate ordering influences ranking/evaluation outcomes—and low repetition consistency—yielding unstable predictions for identical inputs—thereby undermining reliability. To address these issues, we propose a dynamic repetition strategy featuring the first confidence-driven early-stopping mechanism: for each input instance, it adaptively estimates the minimal required number of repetitions, then integrates majority voting with explicit positional bias modeling for fine-grained correction. Unlike static repetition schemes, our method eliminates the need for pre-specified repetition counts. We validate it across three LLM scales and two distinct task categories. Results show that our approach reduces average model calls by 81%–87% compared to static repetition, while preserving high ranking accuracy. This yields significant improvements in both computational efficiency and robustness without sacrificing performance.

Adaptively determining repetitions per instance dynamicallyMitigating position bias in LLM-based ranking tasksReducing computational costs of repetition strategies

LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations

Apr 27, 2025
LD
Laura Dietz
🏛️ University of New Hampshire | RMIT University | Canva | University of Waterloo | University of Edinburgh | Radboud University | Microsoft

This paper identifies validity risks in using large language models (LLMs) for evaluating information retrieval (IR) systems: when LLM-based assessments simultaneously guide system development and performance evaluation, they risk reinforcing biases, undermining reproducibility, and introducing methodological inconsistency—leading to spurious success claims and misleading conclusions. To address this, the authors propose a verifiable risk analysis framework comprising (1) quantitative detection methods for three core validity threats, (2) lightweight mitigation guardrails, and (3) a human-in-the-loop paradigm for constructing reusable, auditable test collections. Grounded in empirical analysis, assessment validity theory, and cross-institutional collaboration, the work delivers an open-source validation toolkit and an industry consensus guideline. These contributions establish responsible, reproducible, and accountable best practices for LLM-augmented IR evaluation.

Assessing reliability of LLM-based evaluations in IR systemsIdentifying risks like bias reinforcement in LLM judgmentsProposing solutions for responsible LLM use in evaluations

Generating Diverse Criteria On-the-Fly to Improve Point-wise LLM Rankers

Apr 18, 2024
FG
Fang Guo
🏛️ Westlake University | South China University of Technology | Google

Pointwise large language model (LLM) rankers suffer from limited adherence to standardized comparative guidelines and insufficient capability in holistically evaluating complex passages. To address this, we propose a dynamic multi-perspective evaluation criterion generation method: leveraging prompt engineering to instantiate interpretable, dimension-specific criteria—covering semantics, relevance, structure, and more—in real time, and jointly aggregating scores across these criteria. This mechanism is the first to achieve decomposability, interpretability, and synergistic enhancement in LLM-based evaluation. Evaluated on the BEIR benchmark across eight diverse datasets, our approach significantly improves ranking performance, yielding an average 3.2% relative gain in NDCG@10. Results demonstrate that dynamic, multi-perspective guidance effectively enhances the ranking capability of pointwise LLM rankers.

Inadequate comprehensive analysis for complex passagesNeed multi-perspective criteria to enhance ranking performanceStandardized comparison guidance lacking in LLM rankers

RankLLM: A Python Package for Reranking with LLMs

May 25, 2025
SS
Sahel Sharifymoghaddam
🏛️ University of Waterloo

To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.

Addresses reliability concerns in LLM APIs and MoE modelsDevelops modular Python package for LLM-based document rerankingEnables quick reproduction of results for research and applications

Latest Papers

What's happening recently
View more

Large language models are susceptible to selection bias in adaptive prompting and program search, leading to an overestimation of the winning candidate’s performance under real-world deployment. This work proposes the SIREN protocol, which enables unbiased performance inference for the full tuning-to-deployment pipeline under a fixed tuning budget by freezing the candidate set, decoupling selection and evaluation data, and incorporating an entry-wise Gaussian multiplier bootstrap. SIREN is the first method to simultaneously support accurate estimation of program-level performance curves on limited-budget grids and construct confidence intervals for both within-budget and cross-budget comparisons. Empirical results demonstrate that conventional winner-reporting practices exhibit substantial optimistic bias, whereas SIREN closely approximates the true evaluation target under finite-sample conditions, offering reliable guidance for deployment decisions.

adaptive benchmarkingLLM evaluationperformance estimation

This study challenges the prevailing assumption that iterative updates to large language models (LLMs) inherently enhance consistency in relevance judgment. Specifically, it investigates how backbone network evolution affects the stability of LLM-based evaluation. Employing the UMBRELA single-prompt and EXAM criteria-prompting methodologies, this work systematically assesses multiple generations of commercial and open-source models—including Gemini, GPT, Qwen, and Llama—under fixed prompting conditions. The findings reveal that improvements in aggregated performance do not necessarily translate to greater judgment stability. Notably, newer model versions fail to significantly improve relevance judgment quality and exhibit regression phenomena, wherein correct assessments made by older versions are lost in subsequent iterations. By exposing the latent risks associated with cross-version prompt transferability, this research provides a critical cautionary insight for LLM evaluation practices.

Backbone EvolutionJudgment StabilityLarge Language Models

This study reveals the performance degradation and evaluation bias of large language models (LLMs) in decision-making tasks involving large-scale candidate sets. By establishing candidate set size as a critical evaluation variable, we demonstrate that performance advantages observed in small-scale settings do not guarantee robustness at scale, identifying confidence collapse and positional bias as two primary failure modes. To address these challenges, this work proposes a hierarchical partitioning strategy and a permutation-based reasoning method. Experimental results indicate that these approaches effectively mitigate performance degradation, improving accuracy by approximately 20 percentage points when the number of candidates reaches N=160. Ultimately, this research establishes a new paradigm for the reliable application and rigorous evaluation of LLMs in large-scale decision-making scenarios.

Candidate SelectionDecision MakingEvaluation

This study investigates the impact of user history length on recommendation quality in large language model (LLM)-based recommender systems, challenging the common assumption that more context yields better performance. Using four state-of-the-art models—GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flash—we conduct within-user experiments on the REGEN dataset to evaluate the effectiveness of contextual histories ranging from 5 to 50 items, while also measuring inference latency. Our results demonstrate that increasing context length from 5 to 50 items does not significantly improve recommendation performance, with NDCG scores remaining stable between 0.17 and 0.23. Notably, as few as 5–10 historical items suffice to maintain recommendation quality while reducing inference cost by approximately 88%. This work provides the first empirical evidence of the efficiency of short-context inputs in LLM-based recommendation, offering a practical foundation for low-cost deployment.

Context LengthInference CostLarge Language Models

Traditional peer review faces scalability bottlenecks, while large language model (LLM)-driven automated review lacks systematic investigation into its reliability, robustness, and security. This work addresses this gap by offering the first system-oriented analysis, focusing on two core tasks: critique generation and score prediction. It establishes a taxonomy of LLM-based reviewing approaches and comprehensively evaluates key technical strategies, including prompt engineering, supervised fine-tuning, retrieval augmentation, and alignment optimization. The study uncovers emerging security threats such as prompt injection and data poisoning, examines challenges arising from subjective disagreement and cross-domain generalization, and highlights limitations and domain biases in current benchmarks. Building on these insights, the paper proposes a roadmap toward developing reliable, transparent, and trustworthy AI-assisted scientific review systems.

LLM-based peer reviewreliabilityrobustness

Hot Scholars

YZ

Yue Zhao

Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for Science
QR

Qihan Ren

Shanghai Jiao Tong University
Explainable AIMachine LearningComputer VisionNatural Language Processing
ZT

Zhen Tan

Ph.D. at Arizona State University
Data MiningMachine LearningAI for ScienceUser-centric Explanation
RF

Rogerio Feris

Research Manager, MIT-IBM Watson AI Lab
Computer VisionMachine LearningArtificial Intelligence
YK

Yu Kong

Michigan State U, Assistant Professor; ACTION Lab, Director
computer visionmachine learningdata mining