Whose Name Comes Up? Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

📅 2026-02-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in auditing large language models (LLMs) for academic expert recommendation: the frequent neglect of user-side interventions, which obscures whether performance bottlenecks stem from the model itself or its deployment. To this end, we introduce LLMScholarBench, a novel benchmark that systematically incorporates user interventions into the evaluation framework, jointly assessing both model infrastructure and intervention strategies across multiple tasks in terms of technical quality and social representativeness. Focusing on physics, we evaluate 22 LLMs under interventions including temperature tuning, representativeness-constrained prompting, and retrieval-augmented generation (RAG). Our findings reveal that user interventions do not uniformly improve performance but instead redistribute errors between factuality and diversity: higher temperatures reduce factuality, constrained prompts enhance diversity at the cost of factuality, and RAG improves technical metrics while diminishing diversity.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsHumans and AI: Intelligent User Interfaces

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) are increasingly used for academic expert recommendation. Existing audits typically evaluate model outputs in isolation, largely ignoring end-user inference-time interventions. As a result, it remains unclear whether failures such as refusals, hallucinations, and uneven coverage stem from model choice or deployment decisions. We introduce LLMScholarBench, a benchmark for auditing LLM-based scholar recommendation that jointly evaluates model infrastructure and end-user interventions across multiple tasks. LLMScholarBench measures both technical quality and social representation using nine metrics. We instantiate the benchmark in physics expert recommendation and audit 22 LLMs under temperature variation, representation-constrained prompting, and retrieval-augmented generation (RAG) via web search. Our results show that end-user interventions do not yield uniform improvements but instead redistribute error across dimensions. Higher temperature degrades validity, consistency, and factuality. Representation-constrained prompting improves diversity at the expense of factuality, while RAG primarily improves technical quality while reducing diversity and parity. Overall, end-user interventions reshape trade-offs rather than providing a general fix. We release code and data that can be adapted to other disciplines by replacing domain-specific ground truth and metrics.
Problem

Research questions and friction points this paper is trying to address.

LLM-based scholar recommendation
auditing
end-user interventions
representation bias
recommendation failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM auditing
scholar recommendation
intervention-based evaluation
representation fairness
retrieval-augmented generation
🔎 Similar Papers
No similar papers found.
L
Lisette Espin-Noboa
Complexity Science Hub, Vienna, Austria
G
Gonzalo Gabriel Mendez
Universitat Politècnica de València, Valencia, Spain