Score
Designs and implements methods, tools, and metrics to evaluate and verify recommendations and other outputs produced by large language models and recommender systems, including auditing model outputs and recommender decisions. Builds provenance and retrieval-trace mechanisms, links claims to evidence, runs lightweight reproducibility checks, captures human confirmations and annotations, detects fabricated or missing recommendations, and measures accuracy, usefulness, inclusion/invisibility and related bias rates.
This work addresses a critical gap in auditing large language models (LLMs) for academic expert recommendation: the frequent neglect of user-side interventions, which obscures whether performance bottlenecks stem from the model itself or its deployment. To this end, we introduce LLMScholarBench, a novel benchmark that systematically incorporates user interventions into the evaluation framework, jointly assessing both model infrastructure and intervention strategies across multiple tasks in terms of technical quality and social representativeness. Focusing on physics, we evaluate 22 LLMs under interventions including temperature tuning, representativeness-constrained prompting, and retrieval-augmented generation (RAG). Our findings reveal that user interventions do not uniformly improve performance but instead redistribute errors between factuality and diversity: higher temperatures reduce factuality, constrained prompts enhance diversity at the cost of factuality, and RAG improves technical metrics while diminishing diversity.
This work addresses the lack of transparency and verifiability in existing large language model (LLM)-driven user simulators, which are prone to biases stemming from user background assumptions. The authors propose the first verifiable user simulator framework, conceptualizing the simulator as an auditable engineering component composed of seven key elements: structured Persona, task-aware Contract, human-aligned Execution, traceable Trace, profile-consistency Verification, structured Feedback, and continuous Refinement. By integrating an end-to-end auditing mechanism that combines LLMs, task contracts, bias detection, and feedback-driven optimization, the framework significantly enhances the credibility, fairness, and interpretability of simulated user behaviors. Empirical evaluations in recommendation list assessment and search query generation demonstrate the simulator’s high fidelity and its effectiveness in mitigating demographic biases.
This study replicates and validates the XRec interpretable recommendation framework, focusing on its adaptability and generalization to open-source large language models (LLMs), particularly Llama 3. Method: We propose an enhanced Mixture of Experts (MoE) embedding architecture that jointly incorporates collaborative signals and language modeling capabilities. Systematic ablation studies are conducted via instruction tuning, collaborative learning, and input/output embedding optimization. Contribution/Results: (1) First successful replication of XRec’s core performance on Llama 3, outperforming baseline methods on several metrics; (2) identification of expert module embedding structure as critical for explanation stability and personalization; (3) public release of evaluation code and experimental configurations to foster reproducible research in interpretable recommendation. Results show that integrating collaborative information significantly improves explanation consistency, yet overall performance does not uniformly surpass all strong baselines.
This study addresses the poorly understood sources of output variability in scholar recommendation by large language models (LLMs), particularly the lack of systematic evaluation of how prompt design—such as identity, language, and geographic cues—affects performance. The authors introduce the first benchmark to audit 43 LLMs across multidimensional prompts (identity, language, geography) and interdisciplinary contexts, quantitatively assessing factual accuracy, coverage, diversity, and fairness in recommending scholars across six academic disciplines. Findings reveal that model choice primarily governs baseline technical quality, while identity and geographic cues in prompts significantly influence diversity and fairness: for instance, prompts referencing South Africa reduce factual accuracy, whereas those referencing Japan enhance it but induce homogenization. This work systematically disentangles model- and prompt-level factors, demonstrating that identity framing is a critical, nontrivial determinant of recommendation quality.
The effectiveness and applicability boundaries of large language models (LLMs) in recommendation tasks remain poorly understood. Method: We propose a unified prompt engineering framework that reformulates recommendation as natural language inference, enabling zero-shot and cross-scenario generalization. We conduct controlled, multi-dimensional experiments on MovieLens and Amazon datasets to isolate the independent effects of LLM architecture, parameter scale, context length, and four prompt components—task description, user interest modeling, candidate item construction, and prompting strategy. Contribution/Results: Our study establishes a reproducible evaluation paradigm and demonstrates that LLMs possess intrinsic zero-shot recommendation capability. However, prompt quality and fidelity of user interest modeling constitute critical bottlenecks. Structurally optimizing prompts yields substantial performance gains. This work provides both an empirically grounded benchmark and a practical, deployable technical pathway for LLM-based recommender systems.
This study addresses the absence of dedicated benchmarks for automatically generating policy and operational recommendations from institutional reports. It introduces, for the first time, the task of report-driven policy and operational recommendation generation, accompanied by the first high-quality dataset and a structured evaluation framework tailored to this objective. Distinct from conventional recommender systems, this work focuses on leveraging large language models (LLMs) to identify critical issues within textual reports and generate reflective, actionable suggestions. Experimental results demonstrate that state-of-the-art LLMs perform effectively on this task, and the proposed framework establishes a reliable benchmark and methodological foundation for future research in this emerging domain.
This work addresses the lack of reproducible and calibratable prompt engineering pipelines for evidence synthesis tasks in current large language models (LLMs). It proposes an innovative workflow that decouples scientific task specifications from prompting frameworks for the first time, leveraging annotated data and explicit metrics to drive prompt optimization. The approach operationalizes the entire pipeline into artifacts using the DSPy and GEPA toolchains, employing a small student model to execute tasks while a larger reflection model guides iterative refinement. Supporting structured task definitions, metric-driven search, and cross-framework portability, the method demonstrates strong compilability and artifact consistency in title and abstract screening tasks. Empirical validation further reveals the critical impact of optimization budgets on small-model performance, significantly enhancing the reliability and transparency of LLM-based applications.
This work systematically investigates the novel opportunities and challenges in trustworthiness arising from the deep integration of large language models (LLMs) into recommender systems. Based on a comprehensive review of over 200 studies, it establishes the first unified taxonomy encompassing six key dimensions—fairness, robustness, privacy, explainability, accountability, and safety—and explicitly identifies 13 critical opportunities alongside 18 emerging challenges, including bias amplification and hallucination risks. By synthesizing mainstream datasets and evaluation metrics, the study proposes the first systematic research framework tailored to trustworthy LLM-based recommendation, offering a holistic landscape of the field, summarizing current advancements, and outlining pivotal directions for future research to foster rigorous and principled development in this interdisciplinary domain.
Traditional peer review faces scalability bottlenecks, while large language model (LLM)-driven automated review lacks systematic investigation into its reliability, robustness, and security. This work addresses this gap by offering the first system-oriented analysis, focusing on two core tasks: critique generation and score prediction. It establishes a taxonomy of LLM-based reviewing approaches and comprehensively evaluates key technical strategies, including prompt engineering, supervised fine-tuning, retrieval augmentation, and alignment optimization. The study uncovers emerging security threats such as prompt injection and data poisoning, examines challenges arising from subjective disagreement and cross-domain generalization, and highlights limitations and domain biases in current benchmarks. Building on these insights, the paper proposes a roadmap toward developing reliable, transparent, and trustworthy AI-assisted scientific review systems.