Score
Designs and implements numeric and logical measures that quantify aspects of a system’s or model’s behavior, including methods to derive and extract behavioral metrics from specifications or traces and to construct quantitative and logical behavioral metrics. Builds analyses that prove properties such as soundness and abstraction-respecting similarity, and applies statistical and numeric techniques to compare, analyze, and interpret the resulting behavioral measurements.
This paper addresses the challenge of characterizing behavioral distances in quantitative systems—such as probabilistic and fuzzy transition systems—using modal logic, where existing logical frameworks lack precise expressive power. Methodologically, we first prove that probabilistic trace distance is not definable by unary modal logic; then introduce a graded monadic algebraic construction to establish general criteria for real-valued modal logics to characterize behavioral distances; finally, leveraging category theory and coalgebraic techniques, we develop a novel characteristic logic for trace distance in fuzzy metric transition systems. Our main contribution is a unified, scalable, and multi-granular framework for behavioral distances that overcomes the expressive limitations of classical modal logic. This framework constitutes the first metric semantic unification for quantitative transition systems—spanning both probabilistic and fuzzy settings—with rigorous theoretical foundations and broad modeling applicability.
This work addresses the limitations of existing formalisms for hyperproperties in capturing quantitative aspects inherent in real-world systems, such as numerical relationships in information flow control. To overcome this, the paper introduces Quantitative Hyper-Logic (QHL), a novel framework that reformulates hyperproperty specifications using measure theory, replacing classical Boolean quantifiers with measures to support nested quantitative structures. Leveraging Hoeffding’s inequality and extreme value theory, the authors develop an efficient statistical verification algorithm and provide rigorous analyses of sample complexity and statistical guarantees. Experimental evaluation on quantitative information-flow benchmarks demonstrates that QHL substantially outperforms conventional qualitative approaches, offering superior expressiveness and verification capabilities that better align with the demands of practical systems.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
Traditional binary correctness verification fails to capture quantitative system behaviors. Method: We propose the first automated toolkit for quantitative automata supporting six classical semantics—Inf, Sup, LimInf, LimSup, LimInfAvg, and LimSupAvg—and systematically address core decision problems: emptiness, inclusion, equivalence, and safety/liveness verification. Our approach introduces weighted transition modeling and a generalized value-function framework, integrating symbolic decision procedures, optimization solvers, and automata transformation techniques to enable extremal-value computation, safety-liveness decomposition, and real-time monitoring. Contribution/Results: Experiments demonstrate efficiency on inclusion checking, constant-function recognition, and online monitoring tasks. We release the first open-source benchmark suite for quantitative automata analysis, establishing a scalable, modular, and unified infrastructure for quantitative system verification.
Existing runtime monitors support only Boolean specification verification, making it infeasible to progressively approximate quantitative properties—such as average response time—over infinite traces. Method: This paper establishes the first unified formal framework for quantitative approximate monitoring, introducing quantitative monitors whose estimates monotonically improve as observation prefixes grow, and rigorously modeling the trade-off between estimation accuracy and resource consumption (specifically, register count). Contribution/Results: We prove that register count strictly determines the theoretical upper bound on achievable accuracy; moreover, each additional register strictly increases the attainable precision—demonstrating an irreducible, non-compensatory relationship between resources and accuracy. Our framework conservatively extends classical Boolean monitoring theory while ensuring soundness. The proposed approach provides provably optimal, resource-bounded approximate monitoring for critical performance metrics, enabling verifiable, deployment-aware runtime assurance.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
本文提出了一种形式化的语义块模型和执行评判基准来独立评估规范质量,通过结构化表示和机器可验证条件解决规范确定性问题。
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.
This study addresses the challenge of quantifying and comparing large language model (LLM) behaviors across vendors in a standardized, cost-effective manner. We propose a simple, inexpensive, and reproducible framework for investigating model behavior by applying a fixed set of public stimuli across a cross-vendor panel of models. The framework innovatively integrates three complementary evaluation methods—exact matching, LLM-judge codebooks, and instrumented environments—to enable scalable behavioral tracking at minimal cost. Experiments reveal lexical convergence among models, evolving robustness to suffix-based prompts, divergences in stance adherence, and patterns of documentation non-compliance exhibited by coding agents. Collectively, this work establishes a systematic evaluation paradigm for tracking the behavioral evolution of large language models.