🤖 AI Summary
This study addresses the unreliability of pairwise comparisons and the scarcity of multilingual datasets in evaluating stereotypes within large language models. To this end, it introduces a novel dual minimal pair framework that leverages data augmentation to generate substitute attributes, thereby bridging cross-lingual data gaps. Furthermore, by incorporating mutual information theory to model group–attribute associations, the work establishes a new, robust evaluation metric system. This approach effectively rectifies logically inconsistent preferences and significantly enhances the robustness of cross-lingual bias measurement. Consequently, it enables reliable comparisons of stereotype intensity across diverse languages and models, offering a more rigorous methodological foundation for assessing representational biases in multilingual AI systems.
📝 Abstract
A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models.
Our code is available at https://github.com/stepanat/missing-minimal-pair/.