outcome measurement

Designing and applying concrete measures and assessments of user- and system-level outcomes to demonstrate impact and maturity, quantify how verification or controls affect reliance and decision-making, and evaluate perceptions of effectiveness under security/privacy constraints.

outcomemeasurement

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Traditional compliance assessments rely on point-in-time audits and self-attestation, which struggle to enable continuous, cross-organizational, and traceable verification of security controls in multi-vendor environments. This work proposes a permissioned blockchain-based Third-Party Risk Assessment (TPRA) framework that transforms static compliance into a dynamic, repeatable, and verifiable continuous governance mechanism through smart contract–automated evaluation workflows, multi-party governance protocols, and longitudinal state tracking. The study contributes an actionable TPRA architecture, along with complementary compliance maturity metrics and a qualitative model, enabling quantification and long-term validation of security control implementation maturity across organizational boundaries and time periods.

blockchaincompliance assessmentframework implementation

Accountability of Robust and Reliable AI-Enabled Systems: A Preliminary Study and Roadmap

Jun 20, 2025
FS
Filippo Scaramuzza
🏛️ Jheronimus Academy of Data Science | Tilburg University

AI systems deployed in real-world settings face significant safety and efficacy risks due to misalignment among robustness, reliability, and accountability—key dimensions of trustworthy AI. Method: This study proposes the first theoretical framework that explicitly embeds *accountability* as a core dimension in AI evaluation. Through conceptual evolution analysis, systematic literature review, and multi-source empirical case studies, it develops a tripartite, synergistic assessment model integrating *robustness*, *reliability*, and *accountability*. The framework innovatively unifies governance-by-design, dynamic testing, and responsibility traceability into a cross-layer analytical paradigm. Contribution/Results: It identifies six critical technical and institutional challenges and five novel categories of testing requirements, and delineates a co-evolutionary pathway for technical capabilities and regulatory infrastructure. The work provides an actionable theoretical foundation and implementation roadmap for AI standardization, regulatory practice, and liability attribution.

Assessing robustness and reliability of AI systemsEnsuring safety and effectiveness in AI applicationsIncorporating accountability for trustworthy AI development

Prescriptive Zero Trust- Assessing the impact of zero trust on cyber attack prevention

Aug 18, 2025
ST
Samuel T. Aiello
🏛️ Dakota State University

Enterprises struggle to quantitatively assess the effectiveness of Zero Trust Architecture (ZTA) implementations and lack a structured, stage-wise evolution roadmap. Method: This study proposes a four-level, data-driven Zero Trust Maturity Model (ZTMM), the first to define a quantifiable set of technical controls—including identity verification, micro-segmentation, end-to-end encryption, and automated policy orchestration—grounded in industry best practices and empirically validated. Maturity is stratified into Initial, Developing, Mature, and Optimized levels. Contribution/Results: The model enables organizations to precisely diagnose their current posture, design phased transformation roadmaps, and measure improvements in security posture. Empirical validation demonstrates that adoption significantly enhances organizational resilience against advanced persistent threats (APTs). The ZTMM provides a reusable, standardized assessment framework and implementation guidance for operationalizing ZTA.

Assessing Zero Trust's impact on cyber attack preventionDefining key technical controls for Zero Trust deploymentQuantifying cybersecurity maturity with Zero Trust guidelines

This study addresses the underexplored tension between institutional expectations and lived experience among CMMC assessors operating in non-consultative roles. Drawing on role conflict theory, it employs interpretative phenomenological analysis (IPA) to conduct semi-structured interviews with CMMC-certified assessors, systematically uncovering their subjective experiences and logics of duty fulfillment in this mode. Findings reveal that assessors navigate role conflicts through strategies centered on technical competence, procedural discipline, and boundary management. These insights not only extend theoretical understandings of professional credibility construction in cybersecurity compliance contexts but also offer empirical grounding for establishing interactional norms and boundary-setting practices within CMMC implementation frameworks.

assessor role expectationsCMMCcybersecurity certification

This study addresses the limitations of existing large language model (LLM) vulnerability detection benchmarks, which rely on a single metric and fail to accommodate the diverse evaluation needs of different security stakeholders. To bridge this gap, we propose SecLens-R—the first role-oriented, multidimensional evaluation framework—defining five role-specific weighting schemes across 35 dimensions grouped into seven categories. We evaluate 12 state-of-the-art LLMs on 406 tasks spanning 10 programming languages and 8 OWASP vulnerability types, using both Code-in-Prompt and Tool-Use paradigms. Results reveal significant performance disparities across roles, with score differences up to 31 points for the same model (e.g., Qwen3-Coder: 76.3 for an Engineering Lead vs. 45.2 for a CISO), underscoring the necessity of contextualized, multi-objective assessment and advancing vulnerability detection from uniform standards toward role-driven decision-making paradigms.

large language modelsmulti-stakeholder evaluationrole-specific assessment

Latest Papers

What's happening recently
View more

This study addresses the complex assurance challenges confronting AI-enabled Cyber-Physical Systems (AI-CPS) across perception, computation, control, human factors, and governance dimensions, noting that mere compliance with ISO/IEC 42001 fails to reveal architectural impacts or practical maturity. The authors propose CEDAR-42001, a two-stage method that uniquely maps compliance audit evidence onto a seven-layer AI-CPS architecture and governance hierarchy. By integrating a five-dimensional maturity profile, constraint identification, and rule-driven reasoning, the approach generates a traceable, architecture-aware assurance posture. Applied to an autonomous vehicle fleet case, it revealed that while 89.9% of audit items were compliant, only 34.3% met a high-assurance baseline. The method successfully reconstructed the 2023 Cruise incident, precisely identifying cross-layer deficiencies and recommending targeted mitigations to inform decision-making from strategic to operational levels.

AI-CPSarchitectural layersassurance posture

This study addresses the significant abstraction gap between security-by-design specifications—typically expressed in domain-specific languages (DSLs)—and code-level analyzers, which impedes the traceability of design intent to implementation vulnerabilities. It presents the first large-scale empirical investigation, examining 559 security checks across 36 analyzers and 66 security design DSLs. The authors introduce SecLan, a unified model that captures shared security concepts between these two layers, and validate its structure through expert evaluation involving 22 practitioners and qualitative interviews with 9 additional experts. The findings reveal a pronounced mismatch between security concepts at the design and implementation levels, with existing analyzer checks often relying on overly broad vulnerability descriptions, leading to ambiguous mappings. This work provides both an empirical foundation and a modeling framework to bridge the gap between security design and implementation.

abstraction gapcode analyzersdomain-specific languages

This work addresses the limitations of existing AI trustworthiness assessment approaches, which are either too abstract to support full lifecycle monitoring or rely on single metrics insufficient for governance needs. The paper proposes a lightweight, auditable framework for dynamic trustworthiness management that integrates formal modeling with governance processes. By employing context-sensitive trustworthiness dimension protocols and interpretable rule learning based on decision trees, the framework enables end-to-end monitoring and documentation of AI systems—from design and deployment through re-evaluation. Novel diagnostic tools, including hierarchical transitions, margin-of-boundary analysis, and profile drift detection, are introduced alongside clearly accountable human-in-the-loop checkpoints. Experiments on synthetic AI lifecycle trajectories demonstrate the framework’s effectiveness in detecting performance degradation, abrupt perturbations, and impacts of system updates, thereby establishing a transparent, traceable, and contestable evidentiary basis for AI governance.

AI governanceauditableconformity documentation

This study addresses the limitations of current AI system evaluations, which often suffer from inconsistent methodologies and metrics that yield incomparable results and poor alignment with real-world contexts and human needs. To bridge this gap, the authors propose a reproducible three-stage scenario generation pipeline that integrates human-centered design, operational feasibility, and methodological transparency. The approach begins by eliciting authentic AI use cases from domain experts via structured use case worksheets, then leverages large language model prompt engineering combined with iterative human review to transform these into human-oriented evaluation scenarios. A validation rubric is developed to assess scenario quality. Applied in the financial services sector, the method successfully distilled six high-level AI use case categories and produced 107 validated evaluation scenarios, substantially enhancing the consistency, comparability, and real-world relevance of AI assessments.

AI evaluationapples-to-apples comparisonevaluation scenarios

This work addresses the limitations of existing AI governance frameworks, which rely on static metrics and post-hoc audits and thus lack the capacity for dynamic, real-time assessment of deployment readiness in high-risk systems—particularly regarding fairness discrepancies, threshold sensitivity, and remediation progress. To bridge this gap, the paper proposes the Operational AI Deployment Assurance (OADA) framework, which uniquely models governance uncertainty as an operational challenge within the deployment pipeline. OADA introduces mechanisms such as deployment assurance scores, readiness categorization, threshold stability zones, and governance escalation states to enable closed-loop, dynamic governance from evaluation to deployment. By integrating the Fairness Discrepancy Index (FDI) and FairRisk-FDI with threshold sensitivity analysis and repair-aware assurance evolution, OADA successfully identifies models deemed “compliant” by conventional metrics yet operationally unstable, offering a scalable deployment assurance paradigm for high-stakes domains like medical AI.

AI governancedeployment assurancefairness disagreement

Hot Scholars

CB

Conrad Borchers

Carnegie Mellon University
Educational Data MiningLearning AnalyticsIntelligent Tutoring SystemsSelf-Regulated Learning
VA

Vincent Aleven

Professor of Human-Computer Interaction, Carnegie Mellon University
Learning science and technologiesintelligent tutoring systemseducational games
JL

Jionghao Lin

University of Hong Kong | Carnegie Mellon University | Monash University
Artificial Intelligence in EducationLearning AnalyticsHuman-Centered AIFeedback
JR

Jasper Roe

Durham University
EducationArtificial IntelligenceEdtechAcademic Integrity
MP

Mike Perkins

Head of the Centre for Research & Innovation | Associate Professor, British University Vietnam
Generative AIAcademic IntegrityPerformance ManagementPolicing