Score
Designs, builds, and refines measurement instruments and indices—such as Likert and other rating scales, survey instruments, item banks, and composite indices—by generating items, choosing response formats and anchors, piloting, and empirically operationalizing constructs. Performs psychometric analyses and validation (reliability assessment, factor analysis, item response modeling, appropriate handling of ordinal responses, latent-trait scoring, and evaluation of scale properties) to ensure the instrument measures the intended psychological or behavioral constructs.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
Current psychometric practice relies on separate descriptive statistics (mean and standard deviation) to assess item quality, lacking a standardized diagnostic tool that integrates both to quantify raw deviation from scale midpoints and its uncertainty—especially problematic in small-sample settings. Method: We propose a Standardized Projected Deviation Index (SPDI), derived from Cohen’s *d*, which unifies the magnitude and variability of an item’s raw deviation from the scale midpoint into a single, bounded, scale-invariant, and bias-controlled quality metric. Results: Through theoretical derivation and small-sample simulation studies, we demonstrate that SPDI is interpretable, invariant across items, and robust under limited data. It provides empirically grounded, actionable thresholds for identifying formative indicator redundancy and evaluating reflective indicator consistency—thereby enabling objective, quantitative item-level diagnostics in both exploratory and confirmatory measurement contexts.
Current psychological measurement item validation for large language models (LLMs) lacks efficient construct validity assessment methods. Method: This paper proposes a virtual validation framework grounded in mediation modeling: LLMs generate trait–response mediators—such as cognitive biases and social desirability tendencies—that reflect individual differences and drive simulated respondents’ diverse response behaviors, thereby evaluating items’ robustness in measuring target constructs (Big Five, Schwartz Values, VIA Strengths). Contribution/Results: This work is the first systematic investigation of LLMs’ potential for psychometric validity validation without requiring large-scale human-annotated data. Experiments demonstrate that LLMs reliably generate theoretically grounded mediators and accurately reproduce expected response patterns across all three major theoretical frameworks. The framework successfully supports item selection and validity evaluation, achieving substantial reductions in validation cost while maintaining methodological rigor.
Current practices of directly applying human psychometric instruments to large language models (LLMs) to construct “psychological profiles” suffer from fundamental biases that may mislead research on their usability, safety, and agentic behavior. This study employs a psychometric framework to administer multiple personality and risk-preference scales to 56 instruction-tuned LLMs alongside large human samples, integrating self-report questionnaires, behavioral tasks, and variance decomposition into a multimodal assessment system. Findings reveal that 81–90% of inter-model differences stem from directional response biases rather than genuine traits; this bias diminishes with increasing model capability but persists nonetheless. Scale reliability is almost entirely predicted by a newly proposed metric—“response orthogonality.” These results demonstrate that LLM “psychological profiles” can be artificially manipulated through item selection, exposing critical limitations in prevailing evaluation paradigms.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the distortion in reliability reporting common in psychological research due to selective computation. For the first time, it applies a unified marginal reliability estimator across 889 psychometric datasets from the Item Response Warehouse, enabling standardized reliability assessment of cognitive tests, clinical screening instruments, and personality and attitude scales. Findings reveal that, depending on the reliability definition used, between 30% and 52% of datasets fall below the conventional .80 reliability threshold, with observed variability primarily attributable to genuine differences rather than estimation noise. The work underscores the substantial impact of reliability definition on substantive conclusions and advocates for more transparent and consistent reporting standards in psychometric practice.
This work addresses the limitations of existing LLM-as-a-Judge evaluation methods, which predominantly focus on output quality and lack a systematic framework for assessing the reliability of large language models (LLMs) as measurement instruments. To this end, the study introduces item response theory (IRT)—specifically the graded response model (GRM)—into this domain, proposing a two-stage diagnostic framework that evaluates LLM judges along two interpretable dimensions: internal consistency and human alignment. By integrating prompt perturbations with human rating data, the approach generates interpretable diagnostic signals that effectively identify unreliable LLM judgments. Empirical results demonstrate the method’s capacity to validate the reliability of LLM-as-a-Judge systems, offering both theoretical grounding and practical guidance for their trustworthy deployment in evaluation tasks.
This study addresses the susceptibility of psychometric assessments to misclassification under low reliability and the computational burden of reliability estimation in small samples. It constructs an interpretable reliability distortion cost function to quantify its impact on extreme quantile identification. Building upon a common factor model and latent variable percentile statistics, this work derives a novel variant of Cronbach’s Alpha that achieves high-precision approximation of McDonald’s Omega without requiring parameter estimation, while further optimizing confidence interval construction. The contributions include precisely quantifying the classification consequences of insufficient reliability and proposing a practical alternative for reliability estimation that simultaneously ensures high computational efficiency and low sampling variability.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.