Score
Designs, builds, and analyzes monitoring algorithms that evaluate discounted-sum aggregates over time-indexed observation streams, producing approximate verdicts with formal guarantees such as (ε,δ)-soundness, ε-approximate soundness, or correctness in expectation. This work includes handling stochastic inputs, implementing finite-memory monitors, and quantifying and controlling precision–resource trade-offs such as memory and observation requirements.
This work addresses the fundamental trade-off in runtime monitoring of quantitative signals, where instantaneous noise is high while long-term averaging obscures local structure. It introduces, for the first time, a formal framework for discounted monitoring, defining ε-approximately reliable monitoring in deterministic settings and (ε,δ)-reliability in stochastic ones, thereby establishing statistical optimality and fundamental limits on precision. The authors develop a memory-efficient monitoring algorithm by integrating statistical hypothesis testing with a novel arithmetic specification language supporting multiple discounting schemes, an affine register machine model, and both synchronous and asynchronous semantics. Theoretical analysis yields explicit, tight lower bounds on memory and observation requirements. Empirical evaluation demonstrates the approach’s effectiveness in practical scenarios such as algorithmic fairness.
Existing runtime monitors support only Boolean specification verification, making it infeasible to progressively approximate quantitative properties—such as average response time—over infinite traces. Method: This paper establishes the first unified formal framework for quantitative approximate monitoring, introducing quantitative monitors whose estimates monotonically improve as observation prefixes grow, and rigorously modeling the trade-off between estimation accuracy and resource consumption (specifically, register count). Contribution/Results: We prove that register count strictly determines the theoretical upper bound on achievable accuracy; moreover, each additional register strictly increases the attainable precision—demonstrating an irreducible, non-compensatory relationship between resources and accuracy. Our framework conservatively extends classical Boolean monitoring theory while ensuring soundness. The proposed approach provides provably optimal, resource-bounded approximate monitoring for critical performance metrics, enabling verifiable, deployment-aware runtime assurance.
In high-stakes domains such as credit assessment and judicial risk prediction, real-time monitoring of cross-group fairness in automated decision-making systems remains challenging. Method: This paper proposes a runtime fairness monitoring framework for data streams, formally specifying algorithmic fairness requirements in Real-Time Lola (RTLola)—a temporal stream logic—and designing a lightweight architecture supporting dynamic verification, online statistical testing, and streaming execution. Contribution/Results: The approach overcomes expressiveness and scalability limitations of conventional static fairness analysis. Evaluated on the real-world COMPAS dataset and diverse synthetic benchmarks, it achieves millisecond-scale monitoring latency while accurately detecting group-level disparities. The framework ensures theoretical soundness—grounded in formal verification—and practical deployability, bridging the gap between rigorous fairness guarantees and production-grade streaming systems.
This paper addresses real-time runtime verification of static fairness in machine learning systems with unknown but Markovian dynamics, focusing on dynamically quantifying and certifying decision bias with respect to sensitive attributes under partial or full observability. Method: We propose a formal specification language expressive enough to encode multiple fairness notions and develop two statistical monitoring algorithms—offering uniform and non-uniform error bounds—to enable progressively precise, confidence-guaranteed quantitative fairness verification. Our approach integrates Markov chain modeling, sequential observation analysis, and lightweight quantitative verification. Contribution/Results: The prototype system achieves millisecond-scale response times on loan approval and university admission benchmarks, significantly improving the timeliness, reliability, and scalability of fairness monitoring compared to existing methods.
This paper addresses the challenge of dynamically aligning probabilistic system models with their actual runtime behavior. Methodologically, it proposes a lightweight, real-time alignment monitoring framework featuring an online-computable alignment score, novel differential alignment monitoring (to detect local misalignment trends), and weighted alignment monitoring (to support task-specific customization and model comparison). The monitor is built upon sequential prediction, integrating probabilistic forecasts, distributional similarity metrics (e.g., Wasserstein distance), and high-confidence interval estimation for runtime assessment. Experiments on the PRISM benchmark demonstrate that the monitor incurs low memory overhead, responds rapidly, and effectively detects model–reality misalignment with high accuracy and strong real-time performance. This work establishes a new paradigm for trustworthy verification of probabilistic systems.
This study addresses the problem of reliably observing causal order in shared-memory concurrent systems (COP), formalizing its observability limits and proving that strong consistency—defined as both completeness and reliability—is generally unattainable. The key insight is that the placement of monitoring instrumentation, rather than the choice of timestamp mechanism, fundamentally determines observability guarantees. To this end, the work proposes three non-blocking monitor implementations: FAInc (a centralized atomic counter), Striped (a decentralized counter), and Collect (an iterative register snapshot). Theoretically, all three provide equivalent COP guarantees. Experimental evaluation on a 64-core NUMA architecture demonstrates that Striped achieves throughput comparable to Collect while maintaining linearizability and substantially alleviating the cache contention bottleneck inherent in FAInc.
This work addresses the challenge of inaccurate and noise-amplified verification in stream-based runtime monitoring caused by sensor noise. To tackle this issue, the paper introduces RLola, a novel extension of the Lola specification language that incorporates slack variables to symbolically model sensor noise, thereby avoiding the aliasing problems inherent in interval arithmetic. The approach leverages SMT solving to enable precise offline verification and formally delineates a sublanguage amenable to constant-memory online monitoring. Experimental evaluation within the RTLola framework demonstrates that RLola achieves superior accuracy and computational efficiency, effectively uncovering specification violations that are missed by purely online monitoring techniques.
This work proposes a unified residual semantics framework based on Stone–Čech compactification to address the diverse divergence behaviors in non-terminating computations, including ordinary loops, hybrid cycles, and escapes through non-compact portions of the observation space. The approach models infinite executions as residual process streams modulo structural congruence, characterizes temporal asymptotic behavior via tail clusters, and preserves observational correlations through compactified products. It is the first systematic application of Stone–Čech compactification to modeling residual behaviors of concurrent processes, offering a unified treatment of multiple divergence types and establishing residual tail laws for operations such as prefix and choice, along with their boundary behavior under parallel composition. Key algebraic laws are experimentally validated, and unbounded escapes are quantified through resource observables without explicit reference to points in the remainder of the compactification.