assess method validity

Designs and applies rubrics, tests, and diagnostic procedures that evaluate whether a proposed method, measurement, or solution is appropriate and valid for its intended purpose, distinguishing correctness of outcomes from validity of the procedure or instrument. Builds graduated validity scales and analytic protocols that identify specific procedural or conceptual errors and rate the adequacy of solution methods.

assessmethodvalidity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Procedural Framework for Assessing the Desirability of Process Deviations

Jun 13, 2025
MG
Michael Grohs
🏛️ University of Mannheim | SAP Signavio

Existing process conformance checking techniques identify deviations between process executions and models but cannot assess their desirability—i.e., whether they are problematic, acceptable, or beneficial—leading to subjective, inefficient, and non-reproducible manual evaluation. To address this gap, we propose the first structured, reproducible framework for assessing deviation desirability. Grounded in a systematic literature review and semi-structured expert interviews, the framework defines three mutually exclusive desirability categories—problematic, acceptable, and beneficial—each accompanied by actionable recommendations that integrate theoretical conceptualization with frontline practical insights. We empirically validate the framework through task-oriented experiments, demonstrating significant improvements in analysts’ assessment efficiency and inter-rater consistency. Crucially, it maintains comprehensiveness while supporting concise, actionable decision-making. This work provides a methodological foundation for evidence-based process deviation governance.

Assessing desirability of process deviations systematicallyProviding step-by-step framework for deviation categorizationStreamlining manual, subjective desirability evaluations

Validity in Design Science

Mar 12, 2025
KR
Kai R. Larsen
🏛️ University of Colorado | University of Virginia | Berlin School of Economics and Law | Georgia State University | Memorial University of Newfoundland | Florida International University | University of Sydney

Design science has long lacked a unified definition and systematic validation mechanism for the validity of knowledge claims, undermining scholarly credibility and interdisciplinary communication. Method: This paper introduces the Design Science Validity Framework (DSVF)—the first comprehensive validity framework tailored to design science—defining validity rigorously and classifying knowledge claims into three primary types (principled, causal, and contextual) with corresponding subtypes. It establishes an operational framework architecture and an embedded process model that integrates validity assurance throughout the entire research lifecycle. Contribution/Results: Through framework engineering, claim taxonomy modeling, methodology meta-analysis, and cross-domain case validation, DSVF has successfully supported validity arguments in over ten design science studies and achieved self-validation of its own knowledge claims. The framework significantly enhances the rigor, reproducibility, and interdisciplinary communicability of design science outcomes.

Developing a framework to validate design science artifactsEnsuring validity of knowledge claims in design scienceSystematic validation of knowledge claims in design projects

TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes

Mar 26, 2025
RS
Raj Sanjay Shah
🏛️ Georgia Institute of Technology | AWS AI Labs | OneMedical

Behavioral therapy notes lack standardized quality criteria, impeding legal compliance and clinical utility. To address this, we propose the first multidimensional quality assessment framework specifically designed for behavioral therapy notes, centered on three core dimensions: completeness, conciseness, and faithfulness. Methodologically, we replace conventional Likert-scale evaluations with a novel structured rubric, supported by a manually annotated dataset and a standardized evaluation protocol. Our approach integrates expert co-design, dual-source data augmentation (human-written and LLM-generated notes), and fine-grained human annotation guided by the rubric, complemented by inter-annotator agreement analysis. Results demonstrate that rubric-based assessment yields superior reliability and interpretability. Empirical analysis reveals widespread deficiencies in completeness and conciseness among clinician-authored notes; while LLM-generated notes exhibit faithfulness limitations—particularly hallucinations—they achieve higher preference and scores from clinical practitioners in blinded evaluations.

Assess LLMs' ability to mimic human note evaluationCompare therapist-written and LLM-generated notes using rubricDevelop rubric for evaluating behavioral therapy notes quality

This study addresses the insufficient validation of artifact effectiveness in Design Science Research (DSR). It systematically evaluates how three dominant DSR frameworks—Baskerville et al., Hevner et al., and Gregor & Jones—support five established validity types: instrumental, technical, design, purpose, and generalizability. While all frameworks explicitly emphasize purpose validity, they exhibit structural gaps in instrumental and design validity, undermining research credibility. To remedy this, we propose, for the first time, a revised DSR framework that comprehensively integrates all five validity types. Each type is formally defined and illustrated with concrete examples. Through qualitative comparative analysis and methodological reconstruction, the framework renders validity assessment systematic and operationally feasible. The revised framework enhances the rigor, transparency, and cross-contextual applicability of DSR artifacts and outcomes.

Assessing artifact validity in DSRComparing three influential DSR frameworksRevising DSR framework for systematic evaluation

Latest Papers

What's happening recently
View more

This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.

accountabilitydiagnostic justificationincident response

Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.

coding reliabilityconstruct validitylarge language models

This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.

analytical workflowcomputer-based assessmentsconsistency check

Hot Scholars

AW

Angelina Wang

Cornell Tech
machine learning fairnessevaluation and measurement
GS

Guojie Song

Professor (Research), Tenured of Peking University
Psychological AIAI Safe & Value AlignmentAgent Cognition & Behavioral ModelingLLM&GML
HY

Haoran Ye

AI PhD @ Peking University
AgentAI Safety and AlignmentAI PsychologyLearn to Optimize