psychometric assessment design

Designs, builds, and analyzes psychometric instruments—such as tests, questionnaires, scales, and item banks—and their scoring systems using classical test theory methods. Tasks include item writing and selection, item difficulty and discrimination analysis, estimation of reliability/internal consistency (e.g., Cronbach’s alpha), collection of validity evidence, test form construction and equating, and score interpretation.

psychometricassessmentdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the susceptibility of psychometric assessments to misclassification under low reliability and the computational burden of reliability estimation in small samples. It constructs an interpretable reliability distortion cost function to quantify its impact on extreme quantile identification. Building upon a common factor model and latent variable percentile statistics, this work derives a novel variant of Cronbach’s Alpha that achieves high-precision approximation of McDonald’s Omega without requiring parameter estimation, while further optimizing confidence interval construction. The contributions include precisely quantifying the classification consequences of insufficient reliability and proposing a practical alternative for reliability estimation that simultaneously ensures high computational efficiency and low sampling variability.

Common-factor modelsCronbach's alphaMcDonald's omega

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

Standardized Descriptive Index for Measuring Deviation and Uncertainty in Psychometric Indicators

Dec 24, 2025
MD
Mark Dominique Dalipe Muñoz
🏛️ Iloilo Science and Technology University

Current psychometric practice relies on separate descriptive statistics (mean and standard deviation) to assess item quality, lacking a standardized diagnostic tool that integrates both to quantify raw deviation from scale midpoints and its uncertainty—especially problematic in small-sample settings. Method: We propose a Standardized Projected Deviation Index (SPDI), derived from Cohen’s *d*, which unifies the magnitude and variability of an item’s raw deviation from the scale midpoint into a single, bounded, scale-invariant, and bias-controlled quality metric. Results: Through theoretical derivation and small-sample simulation studies, we demonstrate that SPDI is interpretable, invariant across items, and robust under limited data. It provides empirically grounded, actionable thresholds for identifying formative indicator redundancy and evaluating reflective indicator consistency—thereby enabling objective, quantitative item-level diagnostics in both exploratory and confirmatory measurement contexts.

Develop a standardized index for item deviation in psychometricsEstablish thresholds for redundancy and consistency in indicatorsMeasure item quality by combining mean and standard deviation

Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality

Oct 13, 2025
JJ
Jana Jung
🏛️ University of Mannheim | GESIS - Leibniz Institute for the Social Sciences | Complexity Science Hub Vienna

This study investigates the applicability and ecological validity of human psychometric instruments—such as gender/racial bias and moral judgment scales—when adapted for evaluating large language models (LLMs). Employing a comprehensive assessment framework comprising multi-round item design, prompt variation testing, convergent validity analysis, and behavioral alignment with downstream tasks, we systematically evaluated the reliability and validity of 12 widely used psychological tests across LLMs. Results indicate moderate internal consistency (Cronbach’s α ≈ 0.65–0.78) but critically low ecological validity: model psychometric scores exhibit weak or even negative correlations with actual discriminatory outputs and fairness-related decision-making in realistic scenarios. The core contribution is the first proposal and empirical validation of an ecological validity evaluation paradigm specifically tailored for LLMs, demonstrating fundamental limitations in directly transplanting human-centered scales. These findings provide critical empirical grounding for theoretical reconceptualization and methodological innovation in AI-oriented psychological assessment.

Assessing validity of human tests on AI modelsEvaluating psychometric test reliability for LLMsTesting alignment between scores and real behavior

QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation

Mar 07, 2025
BN
Bang Nguyen
🏛️ University of Notre Dame | University of Wisconsin-Madison | University of Illinois at Urbana-Champaign

Existing question generation (QG) evaluation methods lack alignment with psychometric metrics, failing to reflect true item quality across dimensions such as topic coverage, difficulty, discrimination, and distractor efficiency. Method: This paper introduces Classical Test Theory (CTT) into QG evaluation for the first time. We construct item pairs exhibiting significant quality differences and propose QG-SMS—a large language model (LLM)-based student modeling and simulation framework for interpretable, multi-dimensional automatic assessment. QG-SMS integrates LLM-driven student behavioral modeling, simulated response generation, CTT-based metric computation, and human validation. Contribution/Results: Experiments demonstrate that QG-SMS substantially improves the discriminative accuracy and robustness of QG systems in evaluating educational item quality. Its assessments strongly correlate with actual student performance and outperform conventional automated metrics.

Evaluating test item quality in educational assessmentsIdentifying shortcomings in existing QG evaluation methodsProposing QG-SMS for enhanced test item analysis

Latest Papers

What's happening recently
View more

This work addresses the complex challenges of continuously monitoring item pool quality and health in large-scale AI-driven assessments. It proposes AQuAP, a dashboard system integrated with an item factory framework that leverages operational data analytics to support item generation and pool management. The system introduces novel metrics such as Effective Bank Size (EBS), which combines exposure rates and usage frequency to holistically evaluate the security, diversity, and efficiency of the item pool. By integrating psychometric indicators, exposure control algorithms, and advanced visualization techniques, AQuAP enables real-time monitoring of item pool vitality. The system has been successfully deployed in the Duolingo English Test, significantly enhancing the intelligence and responsiveness of item pool management.

AI-driven testingeducational assessmentitem bank health

This study addresses the cold-start problem in item parameter estimation when newly developed test items lack empirical response data. The authors propose a prediction approach leveraging textual embeddings and regularized regression, accompanied by an evaluation framework integrating resampling-based cross-validation, reliability ceilings, and design ceilings. Innovatively employing a dual “ceiling” analysis, the work demonstrates that differences in parameter predictability stem primarily from measurement reliability rather than the strength of textual information, underscoring the necessity of repeated validation. In the EEDI mathematics item bank, predicted difficulty parameters achieved an R² of 0.53, representing 57% of the reliability ceiling, whereas pseudo-guessing parameters in the three-parameter logistic model proved largely unpredictable due to near-zero reliability ceilings. BEA benchmark experiments further reveal that relying solely on RMSE can obscure extremely low explained variance, highlighting the critical role of dimensionless metrics in model evaluation.

cold start problemitem calibrationpsychometric parameters

This study addresses ongoing validity concerns regarding the direct transfer of human assessments to AI evaluation by examining human-AI construct equivalence from a psychometric perspective. Integrating exploratory factor analysis, consistency testing, and resampling methods, we systematically compared latent structures between humans and large language models in educational assessments. The results provide the first empirical evidence of significantly divergent factor structures in chemistry and quantitative reasoning tasks, demonstrating construct nonequivalence. These findings challenge the prevailing paradigm of interpreting AI capabilities through human normative standards. Instead, this work posits that assessment transfer must be predicated on latent structural similarity, thereby offering a critical theoretical foundation for developing scientifically rigorous AI capability evaluation frameworks.

Assessment ValidityConstruct EquivalenceLarge Language Models

This study addresses the insufficient modeling of embedding-based semantic structures of assessment items and the limited interpretability of context scores in psychometrics. To overcome these limitations, it proposes constructing context scores using external corpus similarity and introduces a partially specified two-step factor analysis procedure, with technical validation integrating vector embeddings, the Bayesian Information Criterion (BIC), and Q-matrix diagnostic models. Applied to TIMSS mathematics items, this work reveals a stable seven-cluster semantic structure that distinguishes multiple association types. Furthermore, it demonstrates that conditional semantic representations outperform single-factor models, offering new directions for text-assisted response calibration.

assessment itemscontextual scoresdiagnostic interpretation

This study presents the first systematic evaluation of whether large language models (LLMs) can capture item discriminability—the core psychometric property that enables reading comprehension items to differentiate between students of varying ability levels. Under a zero-shot setting, 42 LLMs were assessed using two complementary approaches: direct prediction of item discriminability and classical test theory (CTT) calibration based on model-synthesized responses. Results indicate that although LLM outputs contain non-random signals, their discriminability patterns show limited alignment with human-calibrated benchmarks: the best Spearman correlation for direct prediction reaches only 0.152, improving modestly to 0.241 with the CTT approach—still far from reliably replicating the discrimination structure inherent in human responses. This work thus reveals significant limitations of current LLMs in psychometric modeling.

educational assessmentitem discriminationlarge language models

Hot Scholars

YY

Yuichi Yoshida

National Institute of Informatics
Theoretical Computer Science
YP

Yash Pote

National University of Singapore
AIConstraint SolversSAT Solvers.
LF

Livio Finos

Full Professor of Statistics, University of Padova
multivariate analysispermutation testsbiostatisticspsychometrics
CH

Carl Hvarfner

Research Scientist, Meta
Machine LearningBayesian Optimization
CI

Christian Internò

PhD Student, Bielefeld University, Honda Research Institute Europe
Representation LearningMachine LearningDistributed LearningAI Safety