Score
Designs and analyzes scenario-based attitude measurement instruments and item batteries that use vignettes or situational prompts to elicit granular, item-level endorsement or acceptance patterns. Work includes constructing and scoring vignette items, detecting scenario-specific acceptance gradients and within-domain normative hierarchies, testing for null effects at the item or domain level, and linking patterned responses to external predictors such as authority-related variables.
This study addresses a critical oversight in existing probing methods for assessing whether language models are aware of being evaluated: the decisive influence of prompt selection on measurement outcomes, which undermines cross-model comparability. Treating prompts as a core component of measurement design, the authors fix task content while systematically varying prompts and employ controlled experiments, probe direction analysis, variance decomposition, and surface-form ablation tests to quantify the contributions of prompts, models, and their interaction to observed scores. Findings reveal that models account for only a small fraction of score variance; prompt choice can reverse apparent scaling trends; and surface-level prompt features alone suffice to reproduce most published results. The work demonstrates that single-prompt probing yields unreliable comparisons and specifies the minimum number of prompts required for robust evaluation.
This study evaluates whether large language models (LLMs) can automatically generate questionnaires capable of effectively measuring social attitudes and rivaling established expert-designed scales. Through a within-subjects experimental design, the performance of GPT-4–generated questionnaires—elicited via structured prompts—was systematically compared against validated human-crafted scales across three domains: climate change, immigration, and diversity and inclusion. This work presents the first multi-domain, within-participant comparison between LLM-generated instruments and standard psychometric scales. Results indicate that LLM-generated questionnaires reliably capture major attitudinal divides and are suitable for exploratory, large-scale attitude assessment. However, they exhibit lower resolution in uncovering belief structures and reduced precision in differentiating subpopulations compared to expert-developed scales, suggesting promising yet supplementary utility in social science research.
针对大语言模型生成人格测验条目质量不稳定问题,提出AI-GENIE框架结合自适应提示工程与网络心理测量方法,有效提升结构效度并减少语义冗余。
Current psychological measurement item validation for large language models (LLMs) lacks efficient construct validity assessment methods. Method: This paper proposes a virtual validation framework grounded in mediation modeling: LLMs generate trait–response mediators—such as cognitive biases and social desirability tendencies—that reflect individual differences and drive simulated respondents’ diverse response behaviors, thereby evaluating items’ robustness in measuring target constructs (Big Five, Schwartz Values, VIA Strengths). Contribution/Results: This work is the first systematic investigation of LLMs’ potential for psychometric validity validation without requiring large-scale human-annotated data. Experiments demonstrate that LLMs reliably generate theoretically grounded mediators and accurately reproduce expected response patterns across all three major theoretical frameworks. The framework successfully supports item selection and validity evaluation, achieving substantial reductions in validation cost while maintaining methodological rigor.
This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.
This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.
This study investigates whether the stability of large language model (LLM) responses under minor prompt perturbations varies by question type—specifically, objective factual versus subjective belief questions. By applying multidimensional prompt perturbations (pertaining to wording, framing, and formatting) to four prominent instruction-tuned models across six benchmark datasets, and quantifying response consistency using binomial generalized estimating equations (GEE), the work provides the first empirical evidence that prompt robustness critically depends on the interaction between question type and perturbation method. The findings reveal significantly lower response consistency for subjective questions compared to objective ones, challenging the common assumption that LLM outputs directly reflect their internal beliefs or values, and highlighting the pronounced sensitivity of current models to prompt variations in value-laden tasks.
Current practices of directly applying human psychometric instruments to large language models (LLMs) to construct “psychological profiles” suffer from fundamental biases that may mislead research on their usability, safety, and agentic behavior. This study employs a psychometric framework to administer multiple personality and risk-preference scales to 56 instruction-tuned LLMs alongside large human samples, integrating self-report questionnaires, behavioral tasks, and variance decomposition into a multimodal assessment system. Findings reveal that 81–90% of inter-model differences stem from directional response biases rather than genuine traits; this bias diminishes with increasing model capability but persists nonetheless. Scale reliability is almost entirely predicted by a newly proposed metric—“response orthogonality.” These results demonstrate that LLM “psychological profiles” can be artificially manipulated through item selection, exposing critical limitations in prevailing evaluation paradigms.
为理解社交媒体使用如何影响福祉,提出CAST框架,通过多维度测量和建模个体在不同时间尺度上的行为、生理及体验。
该研究通过构建框架来定义GEO可见性评分,解决生成答案中来源出现、引用或品牌提及的评估问题,方法包括设定情况注释、提示表述、执行条件、权重和评分规则。