Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services

📅 2026-02-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates whether open-source large language models exhibit systematic biases related to race/ethnicity or gender in the context of library reference services, with the aim of ensuring equitable service delivery. We introduce a novel fairness auditing framework tailored to library settings by integrating diagnostic classification with linguistic provenance analysis, and apply it to prominent open-source models including Llama-3.1 8B, Gemma-2 9B, and Mistral 8B. Our findings indicate no significant disparities in model responses across racial or ethnic groups, with only one model showing a minor gender-associated bias. These results support the cautious deployment of such models under professional ethical guidelines and offer both methodological innovation and empirical evidence for the responsible application of artificial intelligence in public cultural services.

Technology Category

Natural Language Processing: Ethics — Bias, Fairness, Transparency & PrivacyMachine Learning: Ethics, Bias, and FairnessPhilosophy and Ethics of AI: Bias, Fairness & Equity

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Large language models for searchSocial Networks and Social Media: Fairness and bias in social network and social media analysis
📝 Abstract
As libraries explore large language models (LLMs) as a scalable layer for reference services, a core fairness question follows: can LLM-based services support all patrons fairly, regardless of demographic identity? While LLMs offer great potential for broadening access to information assistance, they may also reproduce societal biases embedded in their training data, potentially undermining libraries' commitments to impartial service. In this chapter, we apply a systematic evaluation approach that combines diagnostic classification to detect systematic differences with linguistic analysis to interpret their sources. Across three widely used open models (Llama-3.1 8B, Gemma-2 9B, and Ministral 8B), we find no compelling evidence of systematic differentiation by race/ethnicity, and only minor evidence of sex-linked differentiation in one model. We discuss implications for responsible AI adoption in libraries and the importance of ongoing monitoring in aligning LLM-based services with core professional values.
Problem

Research questions and friction points this paper is trying to address.

fairness
large language models
library reference services
bias
demographic identity
Innovation

Methods, ideas, or system contributions that make the work stand out.

fairness audit
large language models
diagnostic classification
linguistic analysis
responsible AI
🔎 Similar Papers
No similar papers found.