scale development

Operationalizing latent dimensions into measurement items, iteratively selecting and validating questionnaire items, and using expert feedback and empirical testing to produce a short, reliable psychometric scale.

scaledevelopment

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A formative measurement validation methodology for survey questionnaires

Oct 16, 2025
MD
Mark Dominique Dalipe Munoz
🏛️ Iloilo Science and Technology University

Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.

Addresses model misspecification issues in formative survey indicatorsIntegrates diagnostic checks to ensure psychometric and statistical integrityProvides validation methodology for formative constructs in questionnaires

This study addresses the long-standing divide between latent variable models and network models in psychometrics, which has hindered theoretical integration and methodological innovation. Through an exploratory literature review, cross-disciplinary comparison of statistical models, and visualization techniques, it systematically examines the intrinsic connections among Item Response Theory (IRT), Structural Equation Modeling (SEM), Generalized Linear Models (GLM), and network analysis. The work proposes a unified modeling paradigm that elucidates both the commonalities and complementarities across these approaches, establishing an integrative framework bridging latent variable and network perspectives. This framework not only offers a novel lens for addressing longstanding debates about the nature of psychological constructs but also facilitates the development of reproducible, modular psychometric tools, thereby advancing interdisciplinary collaboration and methodological synthesis.

latent variablesmethodological reconciliationnetwork psychometrics

This work addresses the limitations of existing LLM-as-a-Judge evaluation methods, which predominantly focus on output quality and lack a systematic framework for assessing the reliability of large language models (LLMs) as measurement instruments. To this end, the study introduces item response theory (IRT)—specifically the graded response model (GRM)—into this domain, proposing a two-stage diagnostic framework that evaluates LLM judges along two interpretable dimensions: internal consistency and human alignment. By integrating prompt perturbations with human rating data, the approach generates interpretable diagnostic signals that effectively identify unreliable LLM judgments. Empirical results demonstrate the method’s capacity to validate the reliability of LLM-as-a-Judge systems, offering both theoretical grounding and practical guidance for their trustworthy deployment in evaluation tasks.

automated evaluationItem Response TheoryLLM-as-a-Judge

HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.

Improving rigor and efficiency in HCI measurement designLeveraging LLMs and prior literature for construct developmentStandardizing measurement item design process for researchers

Two-step estimation of latent trait models

Mar 28, 2023
JK
J. Kuha
🏛️ London School of Economics and Political Science | Leiden University

To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.

Evaluating performance compared to one-step and three-step methodsExamining properties through simulation studies and applicationsTwo-step estimation for latent trait models

Latest Papers

What's happening recently
View more

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

This study addresses the common practice in confirmatory factor analysis of accepting standardized factor loadings as low as 0.50, which leads to elevated measurement error, compromised construct validity, and unstable factor solutions. Building on the logic of average variance extracted (AVE) and communality, the authors propose and justify a uniform item-level threshold of λ ≥ 0.70, aligning it with construct-level validity requirements. Through theoretical derivation, Monte Carlo simulations, and structural equation modeling, the research systematically evaluates the impact of weak loadings on measurement quality, factor score determinacy, and model fit. Findings demonstrate that retaining indicators with λ < 0.70 significantly undermines model accuracy and robustness, whereas enforcing the λ ≥ 0.70 criterion enhances the explanatory power and overall quality of latent variable models.

Average Variance Extractedconfirmatory factor analysisconstruct validity

This study addresses the challenge of determining the number of latent dimensions in multidimensional graded response models by proposing an adaptive Bayesian framework for dimensionality selection. The approach introduces a cumulative ordered spike-and-slab (COSS) prior on the column variances of the item loading matrix, which automatically shrinks redundant dimensions while preserving meaningful structure. Coupled with Albert–Chib latent variable augmentation, it enables an efficient Gibbs sampler. A key innovation lies in integrating ordered shrinkage with Bayesian nonparametric principles, allowing adaptive inference of the latent dimensionality without pre-specifying a candidate set and naturally quantifying uncertainty in dimension selection. Simulation studies and analyses of real psychological assessment data demonstrate that the method accurately recovers true dimensional structures, yields more precise parameter estimates, and maintains computational efficiency.

dimension selectionlatent dimensionsmodel uncertainty

This study addresses the challenge of integrating qualitative, interview-based assessments into quantitative psychometric evaluation while preserving scale structural validity. The authors propose ConvScale, a novel method that seamlessly embeds structured psychological scales into naturalistic conversational interviews. Leveraging an AI-driven dialogue system, ConvScale guides participants through contextually rich interactions and employs natural language processing to map utterances to specific scale items, which are then aggregated into construct-level quantitative scores. This approach enables quantitative analysis without sacrificing contextual depth, thereby expanding the applicability of interviews in psychometrics. In an experiment with 18 participants, item- and construct-level scores generated by ConvScale showed strong agreement with self-reports and demonstrated moderate internal consistency. Although structural validity requires further refinement, the findings substantiate ConvScale’s feasibility as an innovative tool for quantitative psychological assessment.

conversational interviewsitem-level scoringpsychometric scales

Hot Scholars

EM

Eduardo Miranda

Professor, Dept. of Civil and Environmental Engineering, Stanford University
Structural EngineeringEarthquake EngineeringPerformance-Based DesignLoss Estimation