🤖 AI Summary
This study addresses persistent challenges in cross-lingual translation of mathematical word problems by large language models, particularly insufficient cultural consistency, diversity compression, and misjudgment of regional context. Through the first large-scale corpus analysis, the authors systematically audit how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro handle culturally embedded entities—such as names, foods, and locations—when translating 60 English problems into seven languages spanning high- and low-resource settings. Combining human annotation with quantitative evaluation, they perform fine-grained coding of 6,489 cultural transformation instances, revealing widespread entropy collapse in diversity, surface-level token preferences, systematic regional misattribution, and frequent cross-cultural contamination errors (e.g., “Easter eggs used in Eid celebrations”). The work introduces the first fine-grained annotation framework for cultural translation, finding model agreement on transformation type in only 62.5% of cases and exact substitution alignment in just 33.5%.
📝 Abstract
Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into Bengali, Hindi, Punjabi (India), Urdu, Sindhi (Pakistan), Italian, and Sicilian (Italy), a language set spanning the full resource spectrum, from high-resource Italian and Hindi to under-studied Sindhi, Sicilian, and Punjabi. We annotate 6,489 entity transformations, coding whether models preserve, localize, generalize, omit, or change entities such as names, foods, and places. Models agree on transformation type in 62.5% of cases and on specific substitutions in only 33.5%, meaning model choice directly shapes which cultural world students encounter. All 21 language-model combinations show entropy collapse, with adaptation compressing rather than expanding cultural diversity. Models prioritize surface markers such as names, foods, and currencies while preserving deeper structural features such as grade-level systems that embed culturally specific assumptions. Despite prompts specifying target countries, models misattribute regional context by using Bangladeshi taka for Indian Bengali students and produce cross-cultural contamination, such as adapting egg hunts as Eid activities. Some failures are visible in individual translations. Others, including diversity collapse, systematic preference for surface markers, and consistent regional misattribution, emerge only through corpus-level analysis. The surface plausibility that makes adapted problems look correct is precisely what makes deeper failures easy to overlook.