🤖 AI Summary
This study investigates how non-native English characteristics affect the response quality of large language models (LLMs). To our knowledge, this work is the first to systematically quantify the impact of non-native rhetorical and lexical features on model outputs. We construct FABLE, a large-scale controlled dataset, and conduct comparative evaluations alongside linguistic feature analyses across 34 open-source LLMs. Our findings reveal that models mirror the higher-order linguistic properties present in user prompts, confirming that non-native speakers tend to receive lower-quality responses. This demonstrates a significant linguistic fairness issue inherent in current LLMs.
📝 Abstract
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.