cross-lingual prompting

Designs, builds, and evaluates prompts, prompt templates, and prompting procedures that elicit desired reasoning and factual responses from models when inputs or outputs involve multiple languages. This includes exploring cross‑lingual prompt variants and multilingual prompt engineering, probing models with alternate languages to surface latent parametric knowledge, inducing models to reason in English, and optimizing or measuring compute–accuracy tradeoffs, factual consistency, and performance disparities across languages.

cross-lingualprompting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Feb 05, 2024
PS
Pranab Sahoo
🏛️ Indian Institute of Technology Patna | Stanford University | Amazon AI

Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.

Analysis of strengths and limitations of prompting approachesOverview of advancements in prompt engineering techniquesSystematic organization of prompt engineering methods

Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages

Dec 02, 2025
LZ
Lechen Zhang
🏛️ University of Illinois Urbana-Champaign | University of Michigan | LG AI Research

This study investigates how system prompts enhance the accuracy and behavioral robustness of large language models (LLMs) in multilingual settings. To this end, we propose the first four-dimensional system prompt evaluation and optimization framework tailored for multilingual scenarios, integrating key prompt components—including chain-of-thought reasoning, affective cues, and contextual grounding—and analyze over ten million inference units across five languages, three mainstream LLMs, and three cross-lingual benchmarks. Experimental results show that high-performing prompts induce more structured and consistent reasoning paths while significantly suppressing code-switching. Our optimization method yields average improvements of 5–10% across all metrics, substantially enhancing cross-lingual reasoning consistency and deployment reliability. The core contributions are: (1) uncovering interpretable associations between prompt components and multilingual performance, and (2) establishing a scalable, principled evaluation paradigm for multilingual system prompts.

Evaluating prompt components for cross-lingual robustnessOptimizing system prompts for multilingual LLM behaviorReducing language-switching in multilingual reasoning patterns

Towards Better Understanding of Program-of-Thought Reasoning in Cross-Lingual and Multilingual Environments

Feb 25, 2025
PP
Patomporn Payoungkhamdee
🏛️ VISTEC | KAIST | Cohere | SCB 10X | AI Singapore | Chulalongkorn University

In multilingual settings, large language models (LLMs) struggle with multi-step reasoning—especially for non-English languages—due to tight coupling between reasoning and execution, rendering chain-of-thought (CoT) prompting ineffective. Method: This work systematically evaluates the Program-of-Thought (PoT) paradigm, which decouples multilingual reasoning generation from executable code execution. We investigate (i) how instruction fine-tuning affects cross-lingual alignment between questions and reasoning steps, and (ii) how reasoning quality—quantified by functional correctness of generated code—determines final answer accuracy. Contribution/Results: We introduce, for the first time, PoT reasoning quality as a heuristic metric for test-time performance prediction and adaptive optimization. Experiments show that PoT-finetuned models significantly outperform CoT baselines on multilingual reasoning tasks; moreover, reasoning quality exhibits strong positive correlation with answer accuracy, revealing the intrinsic mechanism behind PoT’s superior generalization in multilingual contexts.

Enhance multilingual reasoning in LLMsImprove answer accuracy through reasoning qualitySeparate reasoning from code execution

CRaFT: An Explanation-Based Framework for Evaluating Cultural Reasoning in Multilingual Language Models

Oct 15, 2025
SH
Shehenaz Hossain
🏛️ ADAPT Centre | Computer Science Department | Munster Technological University

Current evaluations of multilingual large language models’ (LLMs’) cultural reasoning capabilities rely predominantly on answer accuracy, neglecting interpretability and cross-linguistic comparability. To address this, we propose CRaFT—the first explanation-based framework for cross-cultural reasoning assessment. CRaFT introduces a four-dimensional explanatory quality metric: cultural fluency, deviation, consistency, and linguistic adaptability. Leveraging the World Values Survey, we construct a culturally grounded, multilingual question–explanation dataset covering Arabic, Bengali, and Spanish (2,100+ instances). Empirical analysis reveals salient language-specific patterns: Arabic responses exhibit lower cultural fluency; Bengali reasoning achieves higher overall quality; GPT-4 demonstrates strong linguistic adaptability but weak consistency; conversely, FANAR shows high stability yet limited flexibility. CRaFT establishes a novel, interpretable, decomposable, and cross-linguistically comparable paradigm for evaluating culturally intelligent multilingual LLMs.

Assessing model explanations across cultural contextsEvaluating cultural reasoning in multilingual language modelsMeasuring cross-lingual variation in cultural understanding

Reasoning with Large Language Models, a Survey

Jul 16, 2024
AP
Aske Plaat
🏛️ Leiden University

Large language models (LLMs) exhibit limited multi-step reasoning capabilities—e.g., in elementary mathematics, logical deduction, combinatorial games, and robotic planning—when deployed without fine-tuning. Method: We systematically survey prompt-driven reasoning mechanisms, introducing the first structured taxonomy for LLM reasoning; empirically demonstrate that prompts can elicit metacognitive behaviors such as self-reflection and self-correction; formally define “reasoning by LLMs” as a distinct challenge beyond pattern matching; and integrate chain-of-thought prompting, self-consistency decoding, reasoning-path evaluation, and reinforcement-learning-inspired controllable reasoning frameworks. Contribution/Results: Our work clarifies the fundamental boundaries of LLM reasoning, identifies key open challenges, establishes a unified research paradigm, and proposes a verifiable, controllable, and systematic research agenda for advancing reasoning in foundation models.

Developing methods to control and optimize reasoning processesEnhancing multi-step reasoning in large language modelsEvaluating reasoning performance on diverse benchmarks

Latest Papers

What's happening recently
View more

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

This work addresses the significant performance gaps of current large reasoning models in multilingual settings, where English-centric reasoning patterns are often erroneously imposed on other languages. The authors propose a de-anglocentric approach that first defines measurable multilingual reasoning features and analyzes their association with answer accuracy via logistic regression. They then employ sparse autoencoders to uncover language-specific latent reasoning concepts, which inform a test-time path selection strategy. Experiments across two mathematical benchmarks, four models, and ten languages reveal that while most reasoning features correlate positively with accuracy, their strength—and sometimes even direction—varies substantially across languages. These findings provide an empirical foundation for developing language-adapted evaluation frameworks and reward mechanisms.

language-specific patternsLarge Reasoning Modelsmultilingual reasoning

This study addresses a critical bottleneck in cross-lingual parametric knowledge transfer for large reasoning language models: script mismatch, rather than linguistic family or language divergence, is identified as the primary barrier. Through analysis of the ECLeKTic and MultiLoKo datasets, the work reveals that disparities in writing systems significantly impede knowledge generalization across languages. To mitigate this, the authors propose enhancing the model’s capacity to handle transliteration ambiguity during inference, complemented by regression-based analysis, entity back-translation prompting, synthetic data generation, and targeted supervised fine-tuning (SFT). Experimental results demonstrate that this integrated approach effectively narrows the knowledge transfer gap in cross-script scenarios, establishing the feasibility of improving cross-lingual parametric knowledge transfer through post-training interventions.

cross-lingual knowledge transferlarge reasoning modelsparametric knowledge

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

Dec 27, 2025
AO
Anaelia Ovalle
🏛️ Meta Superintelligence Labs

This study reveals a severe reasoning-conclusion misalignment problem in multilingual large language models (LLMs) for non-Latin-script languages—misalignment rates exceeding those for Latin-script languages by over 2×—indicating that conventional evaluation metrics substantially overestimate their true reasoning capabilities. To address this, we propose the first human-validated, cross-lingual reasoning alignment evaluation framework. It comprises: (i) a human-annotated taxonomy of reasoning errors (primarily evidence omission and logical breaks); (ii) the GlobalMMLU benchmark; (iii) 65K manually verified reasoning chains; (iv) cross-lingual consistency scoring; and (v) fine-grained error annotation. Empirical evaluation across six languages and six state-of-the-art LLMs demonstrates a significant decoupling between reasoning correctness and task accuracy. Crucially, reasoning-conclusion alignment rates for non-Latin-script languages range only from 31% to 47%, markedly below the 68–79% observed for Latin-script languages.

Analyzes error types like unsupported claims and illogical reasoning stepsEvaluates reasoning-conclusion alignment across languages in LLMsIdentifies higher misalignment in non-Latin scripts compared to Latin

Hot Scholars

IV

Ivan Vykopal

Research Assistant, Kempelen Institute of Intelligent Technologies
NLPDeep LearningMachine LearningComputer Vision
FF

Felix Friedrich

postdoc @ Meta FAIR, Montreal
Multimodal AIGenerative AIAI AlignmentAI Safety
JL

Jefrey Lijffijt

Professor at Ghent University - Data Science, Visualisation, ML & AI
Knowledge discoverymachine learningdata visualizationvisual analytics
MB

Maarten Buyl

Postdoctoral researcher at Ghent University
AI safetyfairnesslarge language modelsrecommender systems
FK

Fajri Koto

Assistant Professor (tenure-track), MBZUAI
Computational LinguisticsNatural Language ProcessingMultilingual NLPHuman-centered NLP