Score
Operationalizing latent dimensions into measurement items, iteratively selecting and validating questionnaire items, and using expert feedback and empirical testing to produce a short, reliable psychometric scale.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
This study addresses the long-standing divide between latent variable models and network models in psychometrics, which has hindered theoretical integration and methodological innovation. Through an exploratory literature review, cross-disciplinary comparison of statistical models, and visualization techniques, it systematically examines the intrinsic connections among Item Response Theory (IRT), Structural Equation Modeling (SEM), Generalized Linear Models (GLM), and network analysis. The work proposes a unified modeling paradigm that elucidates both the commonalities and complementarities across these approaches, establishing an integrative framework bridging latent variable and network perspectives. This framework not only offers a novel lens for addressing longstanding debates about the nature of psychological constructs but also facilitates the development of reproducible, modular psychometric tools, thereby advancing interdisciplinary collaboration and methodological synthesis.
This work addresses the limitations of existing LLM-as-a-Judge evaluation methods, which predominantly focus on output quality and lack a systematic framework for assessing the reliability of large language models (LLMs) as measurement instruments. To this end, the study introduces item response theory (IRT)—specifically the graded response model (GRM)—into this domain, proposing a two-stage diagnostic framework that evaluates LLM judges along two interpretable dimensions: internal consistency and human alignment. By integrating prompt perturbations with human rating data, the approach generates interpretable diagnostic signals that effectively identify unreliable LLM judgments. Empirical results demonstrate the method’s capacity to validate the reliability of LLM-as-a-Judge systems, offering both theoretical grounding and practical guidance for their trustworthy deployment in evaluation tasks.
HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.
This study addresses the common practice in confirmatory factor analysis of accepting standardized factor loadings as low as 0.50, which leads to elevated measurement error, compromised construct validity, and unstable factor solutions. Building on the logic of average variance extracted (AVE) and communality, the authors propose and justify a uniform item-level threshold of λ ≥ 0.70, aligning it with construct-level validity requirements. Through theoretical derivation, Monte Carlo simulations, and structural equation modeling, the research systematically evaluates the impact of weak loadings on measurement quality, factor score determinacy, and model fit. Findings demonstrate that retaining indicators with λ < 0.70 significantly undermines model accuracy and robustness, whereas enforcing the λ ≥ 0.70 criterion enhances the explanatory power and overall quality of latent variable models.
This study addresses the challenge of determining the number of latent dimensions in multidimensional graded response models by proposing an adaptive Bayesian framework for dimensionality selection. The approach introduces a cumulative ordered spike-and-slab (COSS) prior on the column variances of the item loading matrix, which automatically shrinks redundant dimensions while preserving meaningful structure. Coupled with Albert–Chib latent variable augmentation, it enables an efficient Gibbs sampler. A key innovation lies in integrating ordered shrinkage with Bayesian nonparametric principles, allowing adaptive inference of the latent dimensionality without pre-specifying a candidate set and naturally quantifying uncertainty in dimension selection. Simulation studies and analyses of real psychological assessment data demonstrate that the method accurately recovers true dimensional structures, yields more precise parameter estimates, and maintains computational efficiency.
This study addresses the challenge of integrating qualitative, interview-based assessments into quantitative psychometric evaluation while preserving scale structural validity. The authors propose ConvScale, a novel method that seamlessly embeds structured psychological scales into naturalistic conversational interviews. Leveraging an AI-driven dialogue system, ConvScale guides participants through contextually rich interactions and employs natural language processing to map utterances to specific scale items, which are then aggregated into construct-level quantitative scores. This approach enables quantitative analysis without sacrificing contextual depth, thereby expanding the applicability of interviews in psychometrics. In an experiment with 18 participants, item- and construct-level scores generated by ConvScale showed strong agreement with self-reports and demonstrated moderate internal consistency. Although structural validity requires further refinement, the findings substantiate ConvScale’s feasibility as an innovative tool for quantitative psychological assessment.
针对大语言模型生成人格测验条目质量不稳定问题,提出AI-GENIE框架结合自适应提示工程与网络心理测量方法,有效提升结构效度并减少语义冗余。