Score
Formulating precise, formal criteria and operational definitions for system properties, goals, or temporal representations so they are interpretable, testable, and comparable. This covers resolving open-texture in property definitions, specifying necessary conditions for interpretability principles, and defining autonomy/operational criteria to categorize methods.
High-level security properties (e.g., confidentiality, integrity) in the Software Development Life Cycle (SDLC) lack systematic refinement mechanisms, leading to semantic disconnects between these properties and concrete artifacts such as threats, defenses, and assets. Method: We propose the first SDLC-wide security property refinement taxonomy, implemented as a formal, refinable, verifiable, and traceable classification framework in Event-B. The framework integrates principles from security engineering and adaptive systems theory. Contribution: It bridges the semantic gap between high-level security objectives and mid-to-low-level security models, enabling co-evolution of security properties with threat and defense models. Rigorously verified in Event-B, the framework ensures logical consistency and correctness. It provides both theoretically sound foundations and practically actionable guidance for security requirements–driven system development.
The rise of multi-paradigm programming languages has rendered traditional paradigm classification methods inadequate, leading to interoperability issues and conceptual ambiguity. Method: We conduct a systematic literature review encompassing 74 studies to diagnose fundamental limitations in existing classification schemes—particularly their coarse conceptual granularity and lack of formal foundations—and propose a reconstructive paradigm framework grounded in type theory, category theory, and Unifying Theories of Programming (UTP). This framework identifies orthogonal atomic primitives to enable formal modeling and theoretical unification of hybrid-paradigm languages. Contributions: (1) An academic evolution map tracing the shift from empirical classification to formal reconstruction; (2) A research roadmap toward a foundational, unified programming paradigm theory; and (3) A rigorously verifiable theoretical basis for language design, tool development, and cross-paradigm integration.
This work addresses the scalability and practicality challenges in modeling operational semantics for programming languages. Methodologically, it introduces a mathematically lightweight yet semantically precise and extensible operational semantics framework, formalizing program computation steps to uniformly support semantic equivalence, reduction semantics, static analysis, compiler correctness proofs, and program property verification. Its key contributions are: (i) systematic modeling of multi-paradigm language features using minimal, accessible mathematical machinery—balancing theoretical rigor with engineering utility; and (ii) significantly enhanced portability and reusability of semantic models, demonstrated through successful formal verification of multiple production compilers and static analyzers. The framework provides a unified, scalable semantic foundation for programming language design, specification standardization, and trustworthy software construction.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.
This work addresses the lack of operationalizability in existing definitions of interpretability, which hinders their utility in guiding model design and reasoning. It proposes the first formalization of interpretability as a symmetry problem, deriving properties and categories of interpretable models from four fundamental symmetry classes. Building on this foundation, the paper constructs a unified Bayesian inversion framework that naturally integrates core reasoning tasks—such as alignment, intervention, and counterfactual inference—into a coherent structure. This approach establishes the first symmetry-based, operationally grounded theory of interpretability, offering a rigorous formal foundation for reasoning in artificial intelligence systems.
Existing AI openness assessments predominantly focus on the availability of data, models, and code, yet neglect whether such openness meaningfully advances democratization, autonomy, and other intended socio-technical outcomes. This paper identifies the root cause as a disconnection from real-world release contexts and a lack of systematic analysis of reuse stakeholders’ identities, purposes, and constraints. Methodologically, it innovatively adapts five core lessons from systems security—emphasizing context-awareness, failure modes, layered defenses, human factors, and emergent properties—and integrates them into openness evaluation for the first time. The resulting socio-technical framework reveals fundamental limitations of conventional metrics in resilience and fairness. It proposes a novel paradigm centered on ecosystem integrity, risk-informed contextualization, and outcome-oriented efficacy—providing both theoretical grounding and actionable guidance for designing substantively open AI systems.
This paper addresses hyperproperties—higher-order system requirements encompassing information-flow security, knowledge reasoning, and robustness, which span multiple execution traces—by proposing the first unified logical and algorithmic framework covering the entire verification lifecycle. Methodologically, it rigorously characterizes the expressive power and decidability boundaries of classical temporal logics (LTL, CTL, S1S) over hyperproperties; then introduces a novel multi-trace synchronization modeling and quantifier alternation handling mechanism grounded in higher-order temporal logic, constraint solving, and symbolic automata. Key contributions include: (i) a comprehensive taxonomy and complexity-theoretic characterization of hyperproperty logics; (ii) an open-source verification toolchain supporting HyperLTL and HyperCTL*; and (iii) end-to-end support for core verification tasks—including satisfiability checking, model checking, runtime monitoring, and controller synthesis.
This study addresses the longstanding challenge of treating machine learning interpretability as a non-functional requirement lacking quantifiable metrics and validation mechanisms. To bridge this gap, the work proposes an innovative approach that reframes interpretability as a verifiable functional requirement through the integration of data and model provenance. By synergizing principles from requirements engineering and machine learning engineering, the authors develop a systematic and operational verification framework. This framework enables, for the first time, the explicit specification and empirical validation of interpretability requirements, thereby substantially enhancing the engineering rigor and trustworthiness of machine learning system development.
This study addresses the pervasive lack of structural integrity in current AI governance documents, which often fail to meet critical requirements such as traceability, dynamic re-verification, and objective evidence. To bridge this gap, the work systematically adapts structural governance principles from aviation software certification standards (DO-178C/DO-330) and proposes a novel integrity framework tailored for static AI governance artifacts. The framework introduces three key concepts—“epoch constraints,” “proof surfaces,” and “structural gaps”—and establishes the seven-principle PromptQ system. Structural analysis of mainstream governance documents reveals that 37% fall below a basic quality threshold, thereby demonstrating the framework’s effectiveness and practicality in enhancing the rigor and verifiability of AI governance documentation.
This study addresses the lack of a unified and stable definition of Artificial General Intelligence (AGI), which has led to significant disagreement in its assessment. Employing a design science research approach, the paper proposes the DAF-AGI framework, which innovatively prioritizes “definition alignment” over “capability alignment” and positions “definitional sovereignty” as a critical dimension of algorithmic sovereignty. The framework evaluates AGI definitions through five sequential criteria for adjudicative fitness and incorporates a governance audit mechanism to examine the authorship, vested interests, and certification structures underlying each definition. Validation across six prominent AGI stances—including one eliminativist perspective—reveals that only performance-oriented definitions classify current generative systems as AGI, while all others either reject this classification or remain indeterminate, thereby underscoring both the ambiguity of existing definitions and the urgent need for governance.
This work addresses the challenge of formal verification for Reflex programs in industrial-scale control systems, where the generation of an excessive number of verification conditions often renders manual analysis impractical. To overcome this limitation, the authors propose a hybrid verification strategy that integrates a structured requirement annotation language with automated invariant inference based on program structure, coupled with an SMT solver to automatically discharge a substantial subset of verification conditions. By leveraging this synergistic approach, the method significantly reduces the number of verification tasks requiring human intervention, thereby enhancing the automation, feasibility, and overall efficiency of formal verification for large-scale process control systems.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.