Score
Designs, builds, or analyzes methods and tools that estimate uncertainty by executing programs or model outputs and comparing observed versus expected behavior to quantify functional uncertainty and behavioral consistency. These methods analyze execution traces and variability across runs to detect silent logical or runtime errors and to produce executability-based uncertainty measures.
This work proposes the first general framework to systematically quantify and apportion epistemic uncertainty arising from substituting true subprocesses with approximate or learned submodels in stochastic simulation and digital twin applications. The framework constructs confidence or credible intervals for performance metrics via bootstrapping and Bayesian model averaging, and employs a tree-based decomposition to allocate total output variability to individual submodels, yielding importance scores. It is compatible with both parametric and nonparametric models, supports frequentist and Bayesian paradigms, and accommodates dynamic initialization scenarios. Validation on synthetic data and a call center digital twin demonstrates that the method effectively reveals each submodel’s contribution to overall uncertainty, significantly enhancing the interpretability and reliability of simulation outcomes.
This work addresses the deployment reliability challenges of current code language models, which often suffer from overconfidence or underconfidence due to the absence of effective uncertainty estimation and active abstention mechanisms. The authors propose a unified, deployment-oriented framework that treats uncertainty as an actionable signal, jointly optimizing model calibration, selective prediction, and lightweight program analysis tool invocation to establish an end-to-end decision-making and repair pipeline. Evaluated on both classification and generation tasks, the approach enables risk-controlled, coverage-adjustable applications, significantly improving correctness ranking and selective prediction performance while maintaining high coverage. This leads to a substantial reduction in error rates and enhances the practical reliability of code language models in real-world scenarios.
In exploratory programming, fragmented feedback and inefficient comparison hinder iterative development; current informal practices—such as relying on memory, manual annotations, or screenshots—introduce errors and impede reproducibility. To address this, we propose Exploriants, a real-time, example-based programming extension. It introduces the novel “variant point” mechanism to automatically capture probe-style outputs during exploration and designs a domain-adaptive, parallel comparison view that transforms unstructured experimentation into a traceable, reproducible, structured iteration process. Our approach integrates example-driven programming, real-time probe-based output collection, and a configurable visualization interface for side-by-side comparison. We evaluate Exploriants across three domains—image processing, data processing, and game development—demonstrating statistically significant reductions in manual comparison errors, improved exploration efficiency, and enhanced accuracy in directionality assessment during iterative development.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
Quantifying how input uncertainty propagates to model outputs remains a fundamental challenge in computational modeling. Method: This study systematically reviews and empirically compares prominent global and local sensitivity analysis (SA) techniques—including Sobol’, FAST, Morris screening, and local derivative-based methods—implemented via standard software packages, supporting both probabilistic modeling and distribution-free settings. Contribution/Results: We propose a practical decision framework that guides method selection based on problem characteristics, analytical objectives, and resource constraints—rejecting the notion of a universally “optimal” SA method and thereby addressing a critical gap in methodological implementation guidance. A reusable, open-source toolkit is developed to enhance the reliability and interpretability of uncertainty attribution. The framework and tools have been validated across multiple engineering and policy modeling applications, demonstrating robustness and scalability in real-world contexts.
This work addresses the challenges of model uncertainty and unpredictability in partially observable or black-box systems during runtime by proposing a unified theoretical framework that integrates epistemic logic with temporal logic. Leveraging automata theory, it systematically formalizes core concepts—including specification, diagnosis, opacity, and monitorability—and synthesizes lightweight online monitors through offline analysis. The approach is extended to real-time systems, resolving key issues related to their temporal semantics and algorithmic complexity. Furthermore, the study precisely characterizes the fundamental limits of runtime verification, thereby establishing a constructive and implementable foundation for practical deployment of monitoring mechanisms.
Existing sampling-based uncertainty estimation methods struggle to effectively capture behavioral discrepancies in code generation by large language models and lack fine-grained characterization of execution semantics. This work proposes the first uncertainty estimation framework that integrates semantic distance, leveraging multi-candidate program sampling, execution trace analysis, and semantic distance metrics to accurately assess the reliability of generated code—without requiring access to internal model representations or invoking LLM-as-a-judge. The approach significantly outperforms existing techniques across multiple programming languages and benchmarks, reducing runtime by 48%–79% while maintaining robustness across diverse models and settings, thereby addressing a critical gap in quantifying behavioral divergence in code generation.
This work investigates how external uncertainties propagate through structured multi-agent workflows to induce information contamination, thereby degrading reasoning trajectories and output correctness. We introduce a taxonomy of three distinct manifestations of information contamination along with their control-flow characteristics, establishing the first classification framework tailored to structured multi-agent workflows and a trajectory-based detection and localization methodology. Through systematic injection of structured perturbations across 32 GAIA tasks and 614 experimental configurations involving three diverse models, we uncover a decoupling between workflow structural divergence and answer correctness, exposing the fundamental limitations of current validation mechanisms. These findings provide empirical grounding for the design of robust, defense-oriented multi-agent workflows.