Score
Designs, builds, and evaluates prompts, prompt templates, and prompting procedures that elicit desired reasoning and factual responses from models when inputs or outputs involve multiple languages. This includes exploring cross‑lingual prompt variants and multilingual prompt engineering, probing models with alternate languages to surface latent parametric knowledge, inducing models to reason in English, and optimizing or measuring compute–accuracy tradeoffs, factual consistency, and performance disparities across languages.
Multilingual large language models (LLMs) exhibit poor generalization on low-resource languages and heavily rely on parameter-intensive fine-tuning. Method: We systematically review 36 papers (2021–2023), covering 250 languages, 30 NLP tasks, and 39 prompting techniques, and propose the first multidimensional classification and analytical framework integrating language families and resource levels (high/low). We introduce model-agnostic prompting strategies—including natural-language prompt design, zero-/few-shot cross-lingual transfer, knowledge elicitation, and templating—empirically validated on mT5, XGLM, and LLaMA-2-Multilingual. Results: The synthesized state-of-the-art prompting strategies yield an average performance gain of 12.7% on low-resource language tasks without any parameter updates. Our core contribution is the establishment of the first interpretable, transferable theoretical framework and practical guide for multilingual prompt engineering.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
This study investigates how system prompts enhance the accuracy and behavioral robustness of large language models (LLMs) in multilingual settings. To this end, we propose the first four-dimensional system prompt evaluation and optimization framework tailored for multilingual scenarios, integrating key prompt components—including chain-of-thought reasoning, affective cues, and contextual grounding—and analyze over ten million inference units across five languages, three mainstream LLMs, and three cross-lingual benchmarks. Experimental results show that high-performing prompts induce more structured and consistent reasoning paths while significantly suppressing code-switching. Our optimization method yields average improvements of 5–10% across all metrics, substantially enhancing cross-lingual reasoning consistency and deployment reliability. The core contributions are: (1) uncovering interpretable associations between prompt components and multilingual performance, and (2) establishing a scalable, principled evaluation paradigm for multilingual system prompts.
In multilingual settings, large language models (LLMs) struggle with multi-step reasoning—especially for non-English languages—due to tight coupling between reasoning and execution, rendering chain-of-thought (CoT) prompting ineffective. Method: This work systematically evaluates the Program-of-Thought (PoT) paradigm, which decouples multilingual reasoning generation from executable code execution. We investigate (i) how instruction fine-tuning affects cross-lingual alignment between questions and reasoning steps, and (ii) how reasoning quality—quantified by functional correctness of generated code—determines final answer accuracy. Contribution/Results: We introduce, for the first time, PoT reasoning quality as a heuristic metric for test-time performance prediction and adaptive optimization. Experiments show that PoT-finetuned models significantly outperform CoT baselines on multilingual reasoning tasks; moreover, reasoning quality exhibits strong positive correlation with answer accuracy, revealing the intrinsic mechanism behind PoT’s superior generalization in multilingual contexts.
Current evaluations of multilingual large language models’ (LLMs’) cultural reasoning capabilities rely predominantly on answer accuracy, neglecting interpretability and cross-linguistic comparability. To address this, we propose CRaFT—the first explanation-based framework for cross-cultural reasoning assessment. CRaFT introduces a four-dimensional explanatory quality metric: cultural fluency, deviation, consistency, and linguistic adaptability. Leveraging the World Values Survey, we construct a culturally grounded, multilingual question–explanation dataset covering Arabic, Bengali, and Spanish (2,100+ instances). Empirical analysis reveals salient language-specific patterns: Arabic responses exhibit lower cultural fluency; Bengali reasoning achieves higher overall quality; GPT-4 demonstrates strong linguistic adaptability but weak consistency; conversely, FANAR shows high stability yet limited flexibility. CRaFT establishes a novel, interpretable, decomposable, and cross-linguistically comparable paradigm for evaluating culturally intelligent multilingual LLMs.
Large language models (LLMs) exhibit limited multi-step reasoning capabilities—e.g., in elementary mathematics, logical deduction, combinatorial games, and robotic planning—when deployed without fine-tuning. Method: We systematically survey prompt-driven reasoning mechanisms, introducing the first structured taxonomy for LLM reasoning; empirically demonstrate that prompts can elicit metacognitive behaviors such as self-reflection and self-correction; formally define “reasoning by LLMs” as a distinct challenge beyond pattern matching; and integrate chain-of-thought prompting, self-consistency decoding, reasoning-path evaluation, and reinforcement-learning-inspired controllable reasoning frameworks. Contribution/Results: Our work clarifies the fundamental boundaries of LLM reasoning, identifies key open challenges, establishes a unified research paradigm, and proposes a verifiable, controllable, and systematic research agenda for advancing reasoning in foundation models.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This work addresses the significant performance gaps of current large reasoning models in multilingual settings, where English-centric reasoning patterns are often erroneously imposed on other languages. The authors propose a de-anglocentric approach that first defines measurable multilingual reasoning features and analyzes their association with answer accuracy via logistic regression. They then employ sparse autoencoders to uncover language-specific latent reasoning concepts, which inform a test-time path selection strategy. Experiments across two mathematical benchmarks, four models, and ten languages reveal that while most reasoning features correlate positively with accuracy, their strength—and sometimes even direction—varies substantially across languages. These findings provide an empirical foundation for developing language-adapted evaluation frameworks and reward mechanisms.
This study addresses a critical bottleneck in cross-lingual parametric knowledge transfer for large reasoning language models: script mismatch, rather than linguistic family or language divergence, is identified as the primary barrier. Through analysis of the ECLeKTic and MultiLoKo datasets, the work reveals that disparities in writing systems significantly impede knowledge generalization across languages. To mitigate this, the authors propose enhancing the model’s capacity to handle transliteration ambiguity during inference, complemented by regression-based analysis, entity back-translation prompting, synthetic data generation, and targeted supervised fine-tuning (SFT). Experimental results demonstrate that this integrated approach effectively narrows the knowledge transfer gap in cross-script scenarios, establishing the feasibility of improving cross-lingual parametric knowledge transfer through post-training interventions.
This study reveals a severe reasoning-conclusion misalignment problem in multilingual large language models (LLMs) for non-Latin-script languages—misalignment rates exceeding those for Latin-script languages by over 2×—indicating that conventional evaluation metrics substantially overestimate their true reasoning capabilities. To address this, we propose the first human-validated, cross-lingual reasoning alignment evaluation framework. It comprises: (i) a human-annotated taxonomy of reasoning errors (primarily evidence omission and logical breaks); (ii) the GlobalMMLU benchmark; (iii) 65K manually verified reasoning chains; (iv) cross-lingual consistency scoring; and (v) fine-grained error annotation. Empirical evaluation across six languages and six state-of-the-art LLMs demonstrates a significant decoupling between reasoning correctness and task accuracy. Crucially, reasoning-conclusion alignment rates for non-Latin-script languages range only from 31% to 47%, markedly below the 68–79% observed for Latin-script languages.