Score
Designing, validating, and analyzing measurement instruments (scales, assessments) to quantify latent constructs—ensuring reliability, validity, and comparability across experimental conditions and populations for constructs like embodiment, privacy perceptions, or acceptance.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
Traditional structural equation modeling (SEM) relies predominantly on reflective latent variable specifications, limiting its capacity to flexibly represent composite constructs—linear combinations of observed indicators. Existing compositional modeling approaches either compromise core SEM functionalities (e.g., overall model fit assessment, missing data handling, multi-group comparison) or inflate model complexity via auxiliary latent variables. Method: We propose the first SEM framework that unifies composites and latent variables within a single covariance structure model, leveraging maximum likelihood and generalized least squares estimation to directly specify the implied covariance matrix incorporating composites. Contribution/Results: Our approach eliminates the need for redundant latent variables while fully preserving SEM’s diagnostic and inferential capabilities—including fit evaluation, standard error estimation, and hypothesis testing. It significantly enhances expressive power and analytical flexibility for hybrid constructs (reflective + formative) and extends SEM’s applicability to more complex theoretical models.
HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.
Large language models (LLMs) are increasingly deployed in psychological research—as tools, targets of assessment, and cognitive models—yet recent evidence reveals severe measurement unreliability: factor structures of personality traits collapse, moral judgments reverse with minor punctuation changes, and theory-of-mind performance fluctuates dramatically under syntactic rephrasing. These “measurement ghosts” reflect statistical artifacts rather than substantive phenomena, threatening construct validity. Method: We propose the first validity-driven, six-stage workflow integrating psychometric principles and causal inference frameworks, dynamically calibrating validation rigor to research objectives and systematically governing the entire LLM psychology research lifecycle. Our approach includes construct validity verification, computational confound control, modeling of non-independent observations, and transparent experimental design. Contribution/Results: Applied to assessing “LLM selfhood,” our framework successfully disentangles genuine computational phenomena from measurement artifacts, establishing a reproducible empirical paradigm for AI psychology.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the frequent neglect of local cultural perspectives in existing automated evaluations of AI-generated images, particularly regarding “cultural appropriateness.” It introduces a novel evaluation framework that deeply integrates diverse community participation from the outset, collaborating with blind and visually impaired individuals in the UK and residents of Kerala and Tamil Nadu in India to systematically translate lived cultural experiences and community concerns into actionable assessment dimensions. Leveraging multimodal large language models as judges (LLM-as-a-judge), the approach operationalizes community consensus into structured scoring rules, enabling automated evaluation of cultural appropriateness. The work not only establishes a conceptual framework grounded in community values and demonstrates its feasibility but also exposes critical limitations in current AI models’ understanding of cultural context.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.
This study addresses the lack of effective instruments for assessing computational empowerment and self-beliefs among adolescents engaged in constructing and deconstructing AI/ML systems. It develops and validates a six-factor self-belief measurement model tailored to youth, encompassing creative expression, problem-solving self-belief, auditing self-efficacy, interest in auditing, beliefs about design justice, and perceived value of learning AI/ML. Confirmatory factor analysis of survey data from 124 adolescents supports the validity of the proposed six-factor structure and reveals significant positive associations between design justice beliefs and problem-solving self-belief, auditing self-efficacy, and creative expression. By integrating dimensions of construction, deconstruction, and design justice, this work offers a novel instrument for evaluating adolescent AI literacy.
This study addresses the long-standing divide between latent variable models and network models in psychometrics, which has hindered theoretical integration and methodological innovation. Through an exploratory literature review, cross-disciplinary comparison of statistical models, and visualization techniques, it systematically examines the intrinsic connections among Item Response Theory (IRT), Structural Equation Modeling (SEM), Generalized Linear Models (GLM), and network analysis. The work proposes a unified modeling paradigm that elucidates both the commonalities and complementarities across these approaches, establishing an integrative framework bridging latent variable and network perspectives. This framework not only offers a novel lens for addressing longstanding debates about the nature of psychological constructs but also facilitates the development of reproducible, modular psychometric tools, thereby advancing interdisciplinary collaboration and methodological synthesis.
Existing AI psychometrics predominantly repurpose human personality inventories (e.g., Big Five, HEXACO) or ad hoc role definitions, resulting in behavioral distortion and poor domain adaptability. To address this, we propose the first Situation Judgment Test (SJT) framework specifically designed for AI systems, integrating industrial-organizational psychology and personality theory to construct fine-grained, socioemotionally capable virtual personas. Our method innovatively incorporates demographic prior modeling and autobiographical narrative generation, coupled with Pydantic-based structured generation, enabling interpretable and reproducible AI personality modeling and behavioral analysis. We instantiate this framework in a law enforcement assistant scenario, curating a large-scale benchmark: 8,500 virtual personas, 4,000 situational judgment items, and 300,000 AI responses—spanning eight archetype categories and eleven competency dimensions. All data and code are publicly released.