analyze prompt sensitivity

Designs and runs controlled experiments and diagnostic suites that systematically vary prompts and context to measure how prompt wording, framing, and in‑context cues change model outputs; this includes creating prompt perturbations, ablations, automated and iterative evaluation workflows, and cross‑model testbeds. Builds and applies quantitative and qualitative metrics to quantify syntactic and semantic sensitivity, response stability, accuracy differences, and judgment shifts attributable solely to prompt or contextual changes (e.g., comparing zero‑shot, chain‑of‑thought, and source‑grounded prompts).

analyzepromptsensitivity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether existing probing methods genuinely capture large language models’ awareness of evaluation contexts or merely reflect superficial structural cues from prompt formats. For the first time, it systematically disentangles evaluation context from prompt format by constructing a controlled 2×2 dataset and applying diagnostic text rewrites, thereby assessing linear probe effectiveness under partially constrained prompt structures. The results demonstrate that probe signals primarily stem from structural features of benchmark formulations rather than semantic understanding: probe performance substantially degrades when free-form prompts are used. This finding reveals the confounding influence of structural artifacts in current research on model awareness, undermining the reliability of prior conclusions and highlighting the critical need to distinguish genuine semantic comprehension from spurious dependencies on prompt formatting.

evaluation awarenessformat sensitivitylinear probes

This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.

prompt artifactspsychological constructspsychometrics

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Jun 18, 2024
FE
Federico Errica
🏛️ NEC Italia | NEC Laboratories Europe

Large language models (LLMs) exhibit unstable outputs in software applications when prompts undergo minor rephrasings, hindering reliable deployment. Method: This paper introduces two label-free, quantifiable metrics—sensitivity (cross-prompt prediction variance) and consistency (prediction stability across semantically equivalent prompts)—to formally decouple and evaluate LLM robustness to prompt perturbations. Leveraging text classification tasks, we conduct systematic, multi-round prompt rewriting and statistical analysis of prediction distributions. Contribution/Results: Empirical evaluation reveals that mainstream LLMs consistently exhibit high sensitivity and low consistency, exposing a critical robustness gap. Our framework provides a reproducible, ground-truth-label-free diagnostic paradigm for prompt engineering, enabling joint optimization of accuracy and robustness. This work establishes the first formal, measurement-driven approach to assessing and improving LLM resilience against prompt variations.

Large Language ModelsPrediction InstabilitySemantic Sensitivity

Paraphrase Types Elicit Prompt Engineering Capabilities

Jun 28, 2024
JP
Jan Philip Wahle
🏛️ University of Göttingen | University of Toronto

This study investigates how linguistic dimensions of prompt formulation—morphology, syntax, and lexis—affect large language model (LLM) task performance. Method: We conduct controlled experiments across 120 diverse tasks using five LLMs, rigorously matching confounding factors such as prompt length and lexical diversity. This enables the first fine-grained, linguistically grounded attribution analysis of prompt rewriting. Contribution/Results: Semantic-preserving rewrites at morphological and lexical levels yield the largest performance gains, revealing LLMs’ high sensitivity to surface-form variation. We propose a multidimensional prompt rewriting generation framework that integrates cross-model behavioral comparison and median gain statistics. Evaluated on Mixtral 8x7B and LLaMA-3-8B, it achieves median task performance improvements of +6.7% and +5.5%, respectively—demonstrating that linguistically informed rewriting systematically enhances prompt robustness and effectiveness.

Language ModelsPerformance OptimizationPrompting Strategies

Latest Papers

What's happening recently
View more

This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.

bias reinforcementfactual consistencylarge language models

This study addresses the challenge of disentangling whether opinion shifts in large language models (LLMs) stem from relevant evidence or user pressure. To this end, it proposes a black-box evaluation framework that quantifies response variations across multiple scales, including probability, judgment, and action. By innovatively constructing conditional belief response profiles, the approach isolates these two influences, establishes the boundaries of diagnostic validity, and reveals the context-dependence of such evaluations. The framework is systematically validated through controlled variable analysis, synthetic data verification, and targeted ablation experiments. Experiments on models such as Qwen and Llama demonstrate that while surface-level response discrepancies are primarily driven by decoding strategies, certain internal belief patterns remain stable. These findings offer a novel paradigm for evaluating the robustness of LLMs.

attributionbelief revisionevidence-driven

This study addresses the lack of systematic empirical analysis on how format selection, instruction count, and context length in prompt design jointly influence instruction following and hallucination in large language models. Using a unified synthetic corpus—the “Book of Veyra”—the authors conduct over 30,000 API evaluations across five mainstream models, simultaneously manipulating all three variables and measuring performance through multidimensional metrics including instruction-following rate, recall accuracy, sycophancy, and fabrication rate. Key findings reveal that perfect response rates drop to zero when instruction counts exceed 80; recall performance sharply declines once context lengths reach 64–128k tokens, though hallucination remains minimal and refusal rates surge to 90%; and prompt format efficacy varies significantly across models and is constrained by token overhead.

context lengthhallucinationinstruction adherence

This study addresses a critical oversight in existing probing methods for assessing whether language models are aware of being evaluated: the decisive influence of prompt selection on measurement outcomes, which undermines cross-model comparability. Treating prompts as a core component of measurement design, the authors fix task content while systematically varying prompts and employ controlled experiments, probe direction analysis, variance decomposition, and surface-form ablation tests to quantify the contributions of prompts, models, and their interaction to observed scores. Findings reveal that models account for only a small fraction of score variance; prompt choice can reverse apparent scaling trends; and surface-level prompt features alone suffice to reproduce most published results. The work demonstrates that single-prompt probing yields unreliable comparisons and specifies the minimum number of prompts required for robust evaluation.

evaluation sensingmeasurement variancemodel comparison

Hot Scholars

WC

Wanxiang Che

Professor of Harbin Institute of Technology
Natural Language Processing
YC

Yujun Cai

NTU → Meta → Lecturer(Assistant Professor) @UQ
Multi-Modal PerceptionVision-Language Models
BP

Barbara Plank

Professor, LMU Munich, Visiting Prof ITU Copenhagen
Natural Language ProcessingComputational LinguisticsMachine LearningTransfer Learning
YS

Yangqiu Song

HKUST
Artificial IntelligenceData MiningNatural Language ProcessingKnowledge Graphs
BH

Bryan Hooi

National University of Singapore
Machine LearningNatural Language ProcessingGraphsTrustworthy AI