Score
Designs and builds stress-testing experiments and suites that define workloads, failure injections, environmental and resource constraints (power, load, timing, etc.), and feature- or model-level scenarios to exercise behavior under high load, edge cases, and degraded conditions. Implements stress-testing frameworks, automated load generators and harnesses, protocols, instrumentation and metrics collection, and analysis methods to evaluate stability, performance, robustness, and failure modes at component and system levels.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
To address the “simulation-to-reality gap”—the difficulty of reproducing simulation-identified failure scenarios in real-world autonomous driving—this paper proposes a verification method based on formal scenario modeling and time-series matching. The method formally translates abstract scenario programs written in the Scenic probabilistic programming language into computable temporal matching rules, enabling precise retrieval of failure-relevant patterns from large-scale real-world sensor data. A key contribution is the design of an efficient, linearly scalable query algorithm that supports real-time pattern matching over long temporal sequences. Experimental evaluation demonstrates that the approach achieves higher recall accuracy for critical failure scenarios than state-of-the-art commercial vision-language models, while accelerating query throughput by several orders of magnitude. This significantly improves both the efficiency and trustworthiness of transferring simulation-discovered failures to real-vehicle validation.
GPU reliability degrades under compute-intensive workloads due to aging, yet existing monitoring approaches lack fine-grained, interpretable metrics for early reliability prediction. Method: This paper proposes a holistic stress quantification method that fuses real-time telemetry with low-overhead hardware performance counters—including throughput, instruction issue rate, and pipeline stall events—systematically integrating multi-source signals for parallel workloads. Unlike conventional single-metric monitoring, it constructs a GPU stress assessment model explicitly targeting key functional units (e.g., SMs, memory subsystem). Results: The method accurately characterizes dynamic stress distribution across GPU units, significantly improving early reliability prediction in aging-sensitive scenarios. The proposed stress metric exhibits strong correlation (>0.89) with measured lifetime degradation trends, providing an interpretable, deployable foundation for GPU reliability modeling and runtime health management.
This work addresses the challenge of reliability verification for web services coupling multi-physics, multi-scale models in scientific computing. We propose a novel approach integrating serialized formal specifications with usage-driven statistical testing. Our method comprises constructing executable specifications, modeling realistic usage scenarios, performing adaptive statistical testing, and estimating reliability confidence—thereby overcoming the limited coverage of conventional unit testing. To our knowledge, this is the first framework that synergistically combines formal specification and statistical testing for certification of coupled services, introducing quantifiable reliability metrics. Empirical evaluation demonstrates that the method effectively uncovers critical failure paths missed by unit testing and achieves a certified reliability level of ≥0.999 for coupled service controllers at a 95% confidence level.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
研究提出GuardrailLoop测试平台,通过固定策略、计算限制和崩溃恢复方法解决自改进代理流程中的审计问题。
This work proposes a novel approach to black-box testing of Functional Mock-up Units (FMUs) by integrating large language models (LLMs) with a human-in-the-loop mechanism. Addressing the inefficiency and poor interpretability of traditional FMU-based dynamic simulation testing—which relies on manually crafted scenarios—the method automatically generates structured Given-When-Then test objectives from FMU interface and functional specifications, and constructs complete test plans comprising input sequences and assertion oracles. Upon simulation execution, the framework produces visualizable logs and statistical evaluation metrics. The approach significantly enhances test design efficiency and result interpretability, facilitates test asset reuse, and demonstrates effectiveness on a lubricating oil cooling system by autonomously generating executable test scenarios and delivering objective-level pass-rate analysis.
This study addresses the limitations of existing time series forecasting benchmarks, which evaluate a narrow scope and fail to capture structured anomalies or system-level failures in real-world scenarios. This challenge is further compounded by the unauditable pretraining data of foundation models, which impedes rigorous generalization assessment. To overcome these issues, this work proposes a scenario-based stress testing framework that transcends conventional noise perturbation by incorporating causal and system-level failure modeling. By integrating semantic scenarios, explicit failure operators, and graded difficulty levels, the framework jointly evaluates historical inputs, future targets, and multidimensional failure metrics. Ultimately, this research establishes a novel paradigm combining attribution capability with deployment relevance, effectively revealing model failure mechanisms under specific operational conditions.
本文提出一种基于Kubernetes的自动化框架,通过并行执行大规模驾驶模拟测试来提高ADS验证效率,显著缩短了测试时间。
Current evaluations based on benchmark accuracy fail to uncover safety risks of medical large language models in real-world clinical settings. This work proposes the AI-MASLD framework, which introduces narrative stress auditing into medical LLM assessment for the first time, drawing inspiration from metabolic stress testing in hepatology. The framework employs six categories of narrative perturbation probes to conduct dual stress tests on seven models. Using multidimensional metrics—including Metabolic Index (MI), Perturbation Flip Rate (PFR), and Counterfactual Fairness Index (CFI)—the study reveals that while models perform well on clean data, their performance diverges significantly under stress. Notably, open-source models demonstrate safety profiles comparable to or better than closed-source counterparts, and medical fine-tuning paradoxically undermines logical consistency and fairness, exposing a phenomenon of pseudo-normalization.