LIBRA: Measuring Bias of Large Language Model from a Local Context

📅 2025-02-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM bias evaluation suffers from two key limitations: overreliance on U.S.-centric cultural contexts and inability to distinguish genuine bias from hallucinatory responses arising from knowledge boundary gaps. This paper introduces LIBRA, a localized bias assessment framework that systematically measures LLM bias toward New Zealand-specific culture and lexicon—the first such effort in this context. We propose EiCAT, an integrated scoring paradigm combining idealized CAT scores, Boundary-Breach Scores (BBS), and distributional divergence metrics. We construct the first large-scale New Zealand–local bias benchmark, comprising 360,000 test cases, and develop a crowdsourcing-free, corpus-based automated evaluation methodology grounded in local linguistic resources. Experimental results reveal that mainstream models—including BERT, GPT-2, and Llama-3—consistently struggle with New Zealand–specific vocabulary. Notably, while Llama-3 exhibits higher measured bias, it demonstrates superior cultural adaptability compared to its counterparts.

Technology Category

Natural Language Processing: Ethics — Bias, Fairness, Transparency & PrivacyMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Bias, Fairness & Privacy

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Large Language Models (LLMs) have significantly advanced natural language processing applications, yet their widespread use raises concerns regarding inherent biases that may reduce utility or harm for particular social groups. Despite the advancement in addressing LLM bias, existing research has two major limitations. First, existing LLM bias evaluation focuses on the U.S. cultural context, making it challenging to reveal stereotypical biases of LLMs toward other cultures, leading to unfair development and use of LLMs. Second, current bias evaluation often assumes models are familiar with the target social groups. When LLMs encounter words beyond their knowledge boundaries that are unfamiliar in their training data, they produce irrelevant results in the local context due to hallucinations and overconfidence, which are not necessarily indicative of inherent bias. This research addresses these limitations with a Local Integrated Bias Recognition and Assessment Framework (LIBRA) for measuring bias using datasets sourced from local corpora without crowdsourcing. Implementing this framework, we develop a dataset comprising over 360,000 test cases in the New Zealand context. Furthermore, we propose the Enhanced Idealized CAT Score (EiCAT), integrating the iCAT score with a beyond knowledge boundary score (bbs) and a distribution divergence-based bias measurement to tackle the challenge of LLMs encountering words beyond knowledge boundaries. Our results show that the BERT family, GPT-2, and Llama-3 models seldom understand local words in different contexts. While Llama-3 exhibits larger bias, it responds better to different cultural contexts. The code and dataset are available at: https://github.com/ipangbo/LIBRA.
Problem

Research questions and friction points this paper is trying to address.

Measure biases in Large Language Models
Address cultural context limitations in bias evaluation
Develop framework for local bias assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local context bias measurement
Beyond knowledge boundary score
Enhanced Idealized CAT Score