Score
Designs and runs empirical benchmark suites and evaluation protocols for prompting methods—covering text and visual prompts—by building datasets, tasks, metrics, and reporting tools to measure prompt effectiveness, topic-level accuracy variation, clustered error causes, robustness (including image→video attack scenarios), and certification-level performance.
Existing LM evaluation frameworks (e.g., HELM) rely on fixed prompts, suffering from poor generalizability and consequently underestimating model performance and yielding inconsistent cross-model rankings. To address this, we propose DSPy+HELM—a novel integrated framework that systematically incorporates structured prompting strategies (including chain-of-thought, self-consistency, least-to-most, and program-of-thought) into standardized evaluation. We conduct reproducible, large-scale assessments of state-of-the-art LMs across seven diverse benchmarks—four general-purpose and three medical—using declarative prompt optimization to explicitly elicit and enhance model reasoning capabilities. Our approach significantly improves evaluation robustness: average accuracy increases by 4%, result variance decreases by 2%, and ranking reversals occur in 3/7 leaderboards, yielding more accurate estimates of true model capability ceilings. All prompt optimization pipelines and integration tools are open-sourced to strengthen decision utility and experimental reproducibility.
This study investigates whether existing probing methods genuinely capture large language models’ awareness of evaluation contexts or merely reflect superficial structural cues from prompt formats. For the first time, it systematically disentangles evaluation context from prompt format by constructing a controlled 2×2 dataset and applying diagnostic text rewrites, thereby assessing linear probe effectiveness under partially constrained prompt structures. The results demonstrate that probe signals primarily stem from structural features of benchmark formulations rather than semantic understanding: probe performance substantially degrades when free-form prompts are used. This finding reveals the confounding influence of structural artifacts in current research on model awareness, undermining the reliability of prior conclusions and highlighting the critical need to distinguish genuine semantic comprehension from spurious dependencies on prompt formatting.
This work addresses the lack of systematic evaluation and underutilization of metadata in text-to-image (T2I) model assessment and recommendation. Methodologically, we propose the first open-source unified evaluation framework, built upon DeepFashion-MultiModal, integrating multimodal metrics—including CLIP similarity, LPIPS, FID, and retrieval-based scores—and introducing metadata-driven prompt enhancement. We present the first systematic analysis revealing how structured metadata improves visual realism, semantic fidelity, and model robustness. Building on these insights, we design a multi-objective, metric-balanced model-prompt co-recommendation strategy. Experiments demonstrate that our framework significantly enhances state-of-the-art T2I models across perceptual realism, semantic consistency, and cross-architecture stability, enabling fine-grained, task-adaptive model selection and prompt optimization.
Systematic validation of prompt–image alignment evaluation in text-to-image (T2I) generation remains lacking; existing automatic metrics have not been rigorously assessed for quality, reliability, or cross-metric comparability against human judgments. Method: We propose the first comprehensive framework for alignment evaluation—introducing a skill-graded benchmark and a large-scale, multi-template human evaluation dataset with over 100K annotations; designing a skill-driven prompt taxonomy and a multi-round consistency scoring protocol; and developing QA-Metric, a question-answering–based automatic metric aligned with human judgment. Contribution/Results: Experiments demonstrate that QA-Metric significantly outperforms state-of-the-art methods on both our benchmark and TIFA160. Moreover, our analysis uncovers, for the first time, intrinsic connections among prompt ambiguity, model bias, and metric bias—revealing critical limitations in current alignment assessment paradigms.
Multimodal large language models (MLLMs) benchmarks suffer from pervasive non-visual shortcut learning—models achieve high scores by exploiting textual biases, linguistic priors, or superficial statistical patterns, severely compromising the validity of visual understanding evaluation. Method: We propose a “test-set stress testing” and “iterative bias pruning” framework that leverages LLMs to actively detect and quantify textual biases in benchmarks. Using k-fold cross-validation, we fine-tune an LLM and integrate it with random forests and handcrafted features to score and prune biased samples. Contribution/Results: Our method systematically identifies and eliminates non-visually solvable instances across four mainstream benchmarks, yielding the debiased benchmark VSI-Bench-Debiased. It exhibits significantly reduced non-visual solvability, widened performance gaps on visually blind tasks, and robustly advances a vision-centric, reliable paradigm for multimodal evaluation.
Current evaluations of instruction-based embedding models commonly rely on a single prompt, overlooking the models’ sensitivity to variations in instruction wording. This work systematically examines 15 prompt variants per task across six prominent embedding models and eleven datasets—amounting to 990 experimental configurations—and reveals, for the first time at scale, that single-prompt evaluation can severely mislead performance assessment. The default prompt may systematically over- or under-estimate model capabilities, and any model can be made to top leaderboards through favorable prompt selection. To address this issue, the study advocates for adopting multi-prompt evaluation protocols or reporting prompt sensitivity metrics to enhance the robustness and fairness of model assessments.
Large language models (LLMs) exhibit high sensitivity to prompt formulations in vulnerability detection, yet this issue lacks systematic evaluation. This work proposes PromptAudit, a framework that treats prompt sensitivity as a first-order property of vulnerability detection. By controlling dataset, decoding, and parsing variables while varying only the prompting strategy, the study evaluates five prompting methods—including chain-of-thought, few-shot, and self-consistency—across five open-source LLMs on 6,074 CVE samples spanning 16 programming languages. Results show that standard chain-of-thought prompting achieves the best overall performance; few-shot prompting significantly improves performance for prompt-sensitive models; and adaptive chain-of-thought and self-consistency techniques reduce recall and induce excessive abstention, respectively.
Current evaluations of prompt injection and jailbreak detectors for large language models often suffer from dataset-specific threshold tuning and opaque operating points. This work proposes a unified evaluation framework that enforces a single global operating point—maximizing F1 score under a false positive rate (FPR) constraint of ≤1%—across 16 public benchmarks. The framework employs a dual-channel cross-validation strategy combining StratifiedKFold with a StratifiedGroupKFold variant based on parent prompt IDs and MinHash+LSH clustering to mitigate data leakage. To systematically assess detector robustness and generalization, it integrates multidimensional diagnostic mechanisms, including adversarial validation, permutation feature importance, and paraphrase invariance probes. This approach enables consistent, reproducible cross-dataset performance comparisons and establishes quantifiable criteria for evaluating detector generalization.
This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.
Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.