human evaluation

Designing and running human-subject or expert studies and psychometric protocols to measure system outputs on qualities such as coherence, novelty, aesthetics, hallucination and faithfulness, and to empirically validate attack effectiveness and safety shortcomings.

humanevaluation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional human-centric benchmarks in evaluating artificial intelligence once systems surpass human performance. To overcome this, the paper proposes a novel evaluation paradigm based on relative measurement, wherein models autonomously generate public challenges and employ adversarial psychometrics to differentiate among competing systems. By integrating incentive mechanisms with an automated adjudication protocol, the framework enables scalable, human-judgment-free assessment. It is designed to support both verifiable tasks and open-domain scenarios, establishing—for the first time—a dynamic evaluation system that co-evolves with the capabilities of intelligent agents, thereby effectively measuring superhuman-level intelligence.

benchmark saturationbeyond human scaleintelligence measurement

Human-AI Complementarity: A Goal for Amplified Oversight

Oct 30, 2025
RJ
Rishub Jain
🏛️ Google DeepMind

This study addresses the challenging human supervision task of verifying the factual accuracy of AI-generated outputs. Methodologically, we propose a human-AI collaborative fact-checking framework featuring an AI confidence estimation and explanation generation module; it dynamically integrates AI scores with human judgments while modulating human trust by presenting only verifiable search evidence—not definitive conclusions. Our key contribution is the first systematic empirical validation that “lightweight AI assistance”—defined as providing only auditable evidence—significantly mitigates human overreliance on AI. Results demonstrate that our fusion mechanism improves human verification accuracy by 12.7% over both pure-human and pure-AI baselines. These findings establish a scalable technical pathway and cognitively grounded design principles for building trustworthy AI supervision paradigms.

Enhancing fact-verification accuracy through human-AI complementary ratingsImproving human oversight quality using AI assistance systemsOptimizing AI assistance types to prevent human over-reliance

This study addresses the absence of an objective, decomposable, and scalable framework for evaluating the behavioral alignment of large language model agents with human behavior. The authors propose a novel evaluation paradigm grounded in well-established, replicable behavioral hypotheses from social science, systematically translating human-subject experiments into a standardized benchmark for AI agents. They introduce HumanStudy-Bench, an open platform, along with two new metrics—Probability Alignment Score (PAS) and Effect Consistency Score (ECS)—to quantify the consistency between agents and human populations in both inference patterns and effect sizes. Evaluating four agent designs across ten models in twelve robust experiments reveals a bimodal performance distribution, with agent architecture exerting a stronger—and non-monotonic—influence on human alignment than model scale.

agent-human alignmentAI agentsbehavioral hypotheses

Risk Psychology & Cyber-Attack Tactics

Oct 23, 2025
RK
Rubens Kim
🏛️ University of Southern California | Marshall School of Business | Raytheon Technologies | Northeastern University

This study investigates whether individual cognitive traits predict cyberattack behavior. Using red-team operation data from cybersecurity experts in simulated enterprise networks—integrated with psychometric assessments and tokenized attack action logs—we construct a multilevel mixed-effects Poisson regression model (with technical usage frequency nested within participants) to examine how cognitive attributes influence attack technique selection. Results demonstrate significant heterogeneity in the predictive power of cognitive differences across distinct attack techniques; notably, cognitive traits explain variance beyond that accounted for by experience and training. Specific dimensions—including cognitive flexibility and risk preference—robustly predict the adoption of high-stealth or high-complexity attack techniques. This work provides the first empirical evidence of micro-level cognitive mechanisms exerting a dominant influence on cyberattack decision-making. It establishes a foundational theoretical and methodological basis for developing cognition-informed active defense strategies and cognitive threat profiling techniques.

Cognitive processes predict cyber-attack technique selectionIndividual cognitive differences shape cybersecurity attack strategiesPsychological factors outweigh training in attack behavior

Position: AI Evaluation Should Learn from How We Test Humans

Jun 18, 2023
YZ
Yan Zhuang
🏛️ University of Science and Technology of China

Current AI evaluation faces challenges including high annotation costs, training-evaluation data contamination, and low-quality test items—leading to insufficient reliability and validity. This paper pioneers the systematic integration of classical psychometrics—particularly Item Response Theory (IRT)—into AI capability assessment, proposing an IRT-based adaptive testing paradigm. By jointly modeling item parameters and model abilities, it enables dynamic item selection, personalized measurement, and interpretable evaluation. Unlike static benchmark suites, our approach supports real-time ability estimation and online item parameter calibration. Experiments demonstrate that, compared to conventional benchmarks, our paradigm reduces annotation costs by over 40%, mitigates data contamination between training and evaluation, and significantly improves both reliability and validity. This work establishes a theoretical foundation and technical pathway for building robust, efficient, and scalable next-generation AI evaluation frameworks.

Adaptive testing can improve AI evaluation robustness and efficiencyPsychometrics offers solutions for modern AI assessment challengesStatic AI evaluation methods have high costs and reliability issues

Latest Papers

What's happening recently
View more

In AI safety evaluations for mental health applications, expert feedback is often treated as ground truth, yet its reliability remains questionable. This study engaged three board-certified psychiatrists to independently assess large language model–generated mental health responses using standardized scales. Inter-rater reliability was quantified via intraclass correlation coefficients (ICC) and Krippendorff’s alpha, complemented by qualitative interviews and calibrated rating instruments to explore sources of disagreement. Results revealed extremely low inter-rater reliability (ICC: 0.087–0.295), with even negative agreement (α = –0.203) on critical safety items such as suicidal or self-harm content. These discrepancies were not random but stemmed from systematic differences in clinical philosophies—such as prioritizing safety, promoting client engagement, or emphasizing cultural sensitivity—thereby challenging conventional evaluation paradigms that rely on aggregated labels and uncovering profound diversity in professional judgment.

AI safetyexpert disagreementhuman feedback

This study investigates how human-in-the-loop (HITL) feedback influences users’ perceptions of system accuracy and trust, highlighting the critical moderating role of task subjectivity. Through three controlled user experiments that systematically differentiate between objective and subjective task contexts, the research analyzes behavioral measures to assess the effects of feedback interaction. Findings reveal that in objective tasks, providing feedback significantly diminishes users’ trust in and perceived accuracy of the system, whereas this negative effect vanishes in subjective tasks. These results underscore task type as a pivotal factor shaping human–AI trust dynamics and offer important theoretical grounding and practical guidance for the design of HITL systems.

Human-in-the-Loopobjective feedbackperceived accuracy

Human factors in cybersecurity remain conceptualized as static, isolated vulnerabilities, lacking systematic modeling and empirical validation. Method: This paper introduces the first dynamic, multidimensional human-factor cybersecurity framework, integrating a cognitive-affective-behavioral model with attribution theory to systematically identify 50 human factors and 295 interaction mechanisms, and to formally define 12 types of human-factor interactions. Leveraging systematic mapping analysis and empirical psychometrics, we develop a psychometric toolkit comprising 99 actionable metrics. Contribution/Results: The framework enables a paradigm shift from static trait-based perspectives to dynamic systems thinking. It yields three applied outputs: (1) a human-risk diagnostic protocol, (2) evidence-based security training guidelines, and (3) human-centered interface design principles. This work establishes a unified theoretical foundation and engineering-ready infrastructure for human-centric cybersecurity research and practice.

Identifies and maps interactions among cognitive, emotional, and behavioral factors influencing cyberthreat susceptibility.Integrates human factors into cybersecurity as a dynamic, interconnected system.Provides tools for empirical assessment and targeted interventions in human-centered cybersecurity.

It remains unclear whether individual scores derived from large language models (LLMs) for inferring user states in operational settings exhibit psychometric stability and interpretability. This study introduces the first reproducible psychometric evaluation framework, integrating test–retest reliability and aggregate reliability analyses to systematically assess 213 user state indicators generated by multimodal LLMs—including GPT-4o audio, Gemini 2.0 Flash, and Gemini 2.5 Flash. Findings reveal that only 31 indicators meet established reliability criteria. Most individual scores demonstrate insufficient stability for real-time adaptation; however, they remain valuable in post-hoc analyses for uncovering patterns of user interaction and their associations with satisfaction, trust, and engagement.

AI trustworthinesslarge language modelsmetric stability

This study addresses a critical gap in existing vulnerability standards—such as CVE—which focus exclusively on technical flaws while neglecting systematic classification and mitigation of human behavioral and psychological vulnerabilities, including those exploited through social engineering and deception. To bridge this gap, the paper proposes the first structured, machine-readable standardization framework for human-centric vulnerabilities. Grounded in behavioral science theories—including dual-process theory, prospect theory, and models of social influence—the framework systematically identifies and categorizes cognitive, affective, and trust-based mechanisms leveraged by attackers. It introduces a Human Vulnerability Identifier (HVID), an HVSS severity scoring system, and an HVP patching mechanism, thereby establishing a scalable and actionable infrastructure. This foundation enables automated defense strategies and security assessments, effectively filling the long-standing void in human-factor security standardization.

behavioral sciencecybersecurity frameworkhuman vulnerabilities

Hot Scholars

KS

Koustuv Saha

University of Illinois Urbana-Champaign
Computational Social ScienceSocial ComputingHuman-Centered Machine LearningWellbeing
PM

Pattie Maes

Professor of Media Arts and Sciences, MIT
human computer interactionartificial intelligencedigital health
HG

Hatice Gunes

Full Professor of Affective Intelligence & Robotics, University of Cambridge
Artificial IntelligenceAffective AIHealth AIAI Fairness
DW

Dakuo Wang

Northeastern University
Human-AI CollaborationHuman-Centered AIHuman-Computer InteractionAI for Healthcare
MC

Mark Colley

University College London
Automated DrivingAugmented RealityDriver-Vehicle InteractionAccessibility