Score
Designs and applies rubrics, tests, and diagnostic procedures that evaluate whether a proposed method, measurement, or solution is appropriate and valid for its intended purpose, distinguishing correctness of outcomes from validity of the procedure or instrument. Builds graduated validity scales and analytic protocols that identify specific procedural or conceptual errors and rate the adequacy of solution methods.
Existing process conformance checking techniques identify deviations between process executions and models but cannot assess their desirability—i.e., whether they are problematic, acceptable, or beneficial—leading to subjective, inefficient, and non-reproducible manual evaluation. To address this gap, we propose the first structured, reproducible framework for assessing deviation desirability. Grounded in a systematic literature review and semi-structured expert interviews, the framework defines three mutually exclusive desirability categories—problematic, acceptable, and beneficial—each accompanied by actionable recommendations that integrate theoretical conceptualization with frontline practical insights. We empirically validate the framework through task-oriented experiments, demonstrating significant improvements in analysts’ assessment efficiency and inter-rater consistency. Crucially, it maintains comprehensiveness while supporting concise, actionable decision-making. This work provides a methodological foundation for evidence-based process deviation governance.
Design science has long lacked a unified definition and systematic validation mechanism for the validity of knowledge claims, undermining scholarly credibility and interdisciplinary communication. Method: This paper introduces the Design Science Validity Framework (DSVF)—the first comprehensive validity framework tailored to design science—defining validity rigorously and classifying knowledge claims into three primary types (principled, causal, and contextual) with corresponding subtypes. It establishes an operational framework architecture and an embedded process model that integrates validity assurance throughout the entire research lifecycle. Contribution/Results: Through framework engineering, claim taxonomy modeling, methodology meta-analysis, and cross-domain case validation, DSVF has successfully supported validity arguments in over ten design science studies and achieved self-validation of its own knowledge claims. The framework significantly enhances the rigor, reproducibility, and interdisciplinary communicability of design science outcomes.
Behavioral therapy notes lack standardized quality criteria, impeding legal compliance and clinical utility. To address this, we propose the first multidimensional quality assessment framework specifically designed for behavioral therapy notes, centered on three core dimensions: completeness, conciseness, and faithfulness. Methodologically, we replace conventional Likert-scale evaluations with a novel structured rubric, supported by a manually annotated dataset and a standardized evaluation protocol. Our approach integrates expert co-design, dual-source data augmentation (human-written and LLM-generated notes), and fine-grained human annotation guided by the rubric, complemented by inter-annotator agreement analysis. Results demonstrate that rubric-based assessment yields superior reliability and interpretability. Empirical analysis reveals widespread deficiencies in completeness and conciseness among clinician-authored notes; while LLM-generated notes exhibit faithfulness limitations—particularly hallucinations—they achieve higher preference and scores from clinical practitioners in blinded evaluations.
This study addresses the insufficient validation of artifact effectiveness in Design Science Research (DSR). It systematically evaluates how three dominant DSR frameworks—Baskerville et al., Hevner et al., and Gregor & Jones—support five established validity types: instrumental, technical, design, purpose, and generalizability. While all frameworks explicitly emphasize purpose validity, they exhibit structural gaps in instrumental and design validity, undermining research credibility. To remedy this, we propose, for the first time, a revised DSR framework that comprehensively integrates all five validity types. Each type is formally defined and illustrated with concrete examples. Through qualitative comparative analysis and methodological reconstruction, the framework renders validity assessment systematic and operationally feasible. The revised framework enhances the rigor, transparency, and cross-contextual applicability of DSR artifacts and outcomes.
This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.
Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.