Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how non-native English characteristics affect the response quality of large language models (LLMs). To our knowledge, this work is the first to systematically quantify the impact of non-native rhetorical and lexical features on model outputs. We construct FABLE, a large-scale controlled dataset, and conduct comparative evaluations alongside linguistic feature analyses across 34 open-source LLMs. Our findings reveal that models mirror the higher-order linguistic properties present in user prompts, confirming that non-native speakers tend to receive lower-quality responses. This demonstrates a significant linguistic fairness issue inherent in current LLMs.
📝 Abstract
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Non-Native English
Performance Disparity
Prompt Fluency
Response Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Non-Native English
Controlled Dataset
Performance Disparity
Prompt Fluency
💼 Related Jobs
No related jobs found.
Y
Yusheng Zhou
Electrical Engineering and Computer Science Department, University of Michigan, Ann Arbor, MI 48105, USA
E
Eleanor Lin
Electrical Engineering and Computer Science Department, University of Michigan, Ann Arbor, MI 48105, USA
David Jurgens
David Jurgens
Associate Professor, School of Information and Dept. of Computer Science, University of Michigan
Natural Language ProcessingComputational Social ScienceComputational Sociolinguistics