Score
Design and run controlled, often randomized output experiments and analysis pipelines that probe the observable behavior of trained generative and conditional models (primarily text and multimodal language models) using prompt-based elicitation, sampling regimes, and task-style tests to reveal distributional outputs, prompt sensitivity, instability, and systematic biases. Build evaluation metrics, statistical analyses, simulation or integration checks (including NLI-style and task-based diagnostics) to quantify variability, compare models, and identify failure modes and biases in model outputs.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
Existing approaches struggle to systematically compare the outputs of large language models under varying generation conditions. This work proposes a “visual fingerprint” framework that models model outputs as distributions over multidimensional linguistic choices—encompassing content, expression, and structure—and enables cross-condition comparison of generative behavior through an integrated natural language processing pipeline and distribution visualization techniques. For the first time, this method facilitates intuitive, distribution-level insights into stable behavioral patterns that persist across diverse settings yet remain undetectable via single-sample inspection or conventional aggregate metrics. The efficacy of the framework is demonstrated across four distinct application scenarios.
Quantifying output variability of large language models (LLMs) under input perturbations or model variants remains challenging due to the intractability of explicit probability distribution modeling and sensitivity to stochasticity. Method: This paper proposes a black-box robustness auditing framework that formulates output divergence detection as a statistical hypothesis test in semantic embedding space (e.g., BERTScore, STS). It constructs empirical null distributions via Monte Carlo sampling—bypassing explicit modeling of output distributions and mitigating randomness-induced bias. Contribution/Results: We introduce the first distributed perturbation analysis paradigm, enabling model-agnostic, multi-perturbation joint testing, interpretable p-values, scalar effect sizes, and integrated multiple-testing correction (e.g., Bonferroni). Experiments demonstrate accurate quantification of response shifts, reliable estimation of true/false positive rates, and cross-model consistency assessment—achieving significantly improved reliability and reproducibility in LLM robustness evaluation without distributional assumptions.
Large language model (LLM) prompts exhibit poor robustness across model deployments and iterative updates, and lack verifiable, formal specifications. Method: This paper introduces the first automated test generation framework grounded in prompt semantic parsing. It formalizes natural-language prompts as input-output specifications and integrates constraint-guided test case generation with multi-model consistency evaluation to enable specification-driven, interpretable validation. Contribution/Results: It pioneers treating prompts as testable software artifacts, supporting self-supervised, reasoning-driven test generation and cross-model behavioral regression detection. Evaluated on eight benchmark prompts, the method increases the detection rate of invalid outputs across four major LLM families (GPT-4, Claude, Llama, etc.) by an average of 37% over baselines. An open-source tool implementing the framework is production-ready for industrial prompt debugging.
Non-expert users struggle to efficiently optimize LLM prompts due to limited domain knowledge and insufficient feedback mechanisms. To address this, we propose a beginner-oriented visual prompt engineering system featuring a novel tri-strategy collaborative optimization framework—integrating keyword perturbation, semantic paraphrasing, and optimal few-shot example recommendation. We design a multi-view synchronized interface, an interactive prompt editing environment, and a real-time evaluation mechanism grounded in both semantic similarity and task-specific accuracy. Experiments demonstrate that our system reduces user prompt iteration time by 37%, increases prompt diversity by 2.1×, and improves average accuracy by 11.4% across multiple NLP tasks—significantly outperforming existing prompt interfaces. This work lowers the cognitive barrier to prompt engineering and establishes a new paradigm for LLM interaction that is interpretable, iterative, and empirically evaluable for non-experts.
This study investigates whether large language models (LLMs) can unsupervisedly detect latent regularities in noisy, non-linguistic acoustic inputs—motivated by the challenge of recognizing communication intent from extraterrestrial intelligence (SETI) with unknown signaling conventions. Method: We propose the “Generative Reactivity” framework, which abandons conventional symbolic decoding assumptions and instead treats structural coherence in model outputs as evidence of underlying order in the input. We introduce the Semantic Induction Potential (SIP), a composite metric integrating entropy, syntactic coherence, compression gain, and repetition penalty to quantify response strength. Contribution/Results: Using zero-shot, cross-modal evaluation on GPT-2 small (117M), we observe statistically significant SIP increases for humpback whale song and nightingale vocalizations versus white noise (p < 0.01), while human speech elicits only moderate reactivity. These findings demonstrate that LLMs can autonomously perceive statistical structure in non-human acoustic signals without prior encoding knowledge—establishing a hypothesis-free paradigm for SETI signal detection.
This study addresses the threat posed by the stochasticity of large language model (LLM) outputs to research reproducibility, demonstrating that randomness arises not only from sampling strategies but also from non-sampling factors such as silent model updates, numerical rounding, and expert routing in mixture-of-experts architectures. The work is the first to explicitly model LLM outputs as draws from a probability distribution and systematically evaluates the impact of various sources of randomness on downstream outcomes through regression analysis in a sentiment classification task. Empirical comparisons across multiple software and hardware environments—using both API-accessible and locally deployed open-source models—reveal substantial output variability even at zero temperature. Building on these findings, the paper proposes standardized reporting guidelines for research papers and reproduction packages to encourage the adoption of stricter LLM usage protocols within the community.
Current interactions with language models often rely on single-sample outputs, which fail to reveal structural properties of the output distribution—such as multimodality or sensitivity to prompt perturbations—and can lead users to overgeneralize in open-ended tasks. To address this limitation, this work proposes GROVE, the first interactive tool that visualizes language model generations as textual graphs, representing shared structures, branching points, and clusters through overlapping paths while preserving access to original outputs. Through three crowdsourced user studies, we demonstrate that GROVE significantly enhances users’ structural understanding of generation diversity. Moreover, combining graph-based summaries with direct access to raw outputs enables both macro-level insights and fine-grained analysis, effectively bridging the gap in distribution-aware workflows.
This study addresses the instability of large language models (LLMs) across different prompting styles, a phenomenon whose underlying mechanisms remain poorly understood. By systematically comparing instruction-based and exemplar-based prompts through attention head analysis, representational probing, and behavioral experiments, the work identifies for the first time a shared “lexical task attention head” that operates consistently across prompting paradigms. The activation strength of this attention head strongly correlates with model performance and reliably triggers task-relevant answer generation. Furthermore, the research reveals that competition among internal task representations is a key driver of prompt sensitivity. These findings offer an interpretable framework for understanding LLM internal dynamics and suggest novel directions for enhancing prompt robustness.
We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their free-form, multi-token generations across aligned prompt contexts and produces human-readable, statistically validated natural-language hypotheses describing how the models differ, along with recurring themes that summarize patterns across validated hypotheses. We evaluate the approach in synthetic setting by injecting known behavioral changes and showing that the pipeline reliably recovers them. We then apply it to three real-world interventions, reasoning distillation, knowledge editing and unlearning, demonstrating that the method surfaces both intended and unexpected behavioral shifts, distinguishes large from subtle interventions, and does not hallucinate differences when effects are absent or misaligned with the prompt bank. Overall, the pipeline provides a statistically grounded and interpretable tool for post-hoc auditing of intervention-induced changes in model behavior.