Score
Designing, deploying, and validating experiments on real physical systems (hardware-in-the-loop, embedded platforms, robotics, RF transmitters) to ensure real-world performance and robustness. Employed to validate behavior-tree strategies in competition, port NAS pipelines to low-end MCUs, and verify over-the-air attack/impersonation effectiveness across hardware variants.
This work addresses emerging security threats—such as deepfakes, semantic manipulation, and MCP protocol exploits—that challenge AI agents in cyber-physical systems (CPS), where traditional defenses fall short. To counter these risks, the authors propose SENTINEL, a holistic security framework spanning the AI agent lifecycle. SENTINEL integrates threat modeling, feasibility assessment, defense selection, and continuous validation, while incorporating physical constraints and data provenance mechanisms. Empirical evaluation in a smart grid case study demonstrates that detection alone is insufficient for securing safety-critical CPS; instead, robust protection requires synergistic enforcement of physical laws and trustworthy data lineage to achieve defense-in-depth. This research provides a systematic design methodology for building trustworthy, AI-enabled cyber-physical systems.
This work addresses the inefficiency and insufficient coverage in security verification of complex hardware systems by proposing the first structured integration of artificial intelligence and large language models (LLMs) across the entire hardware security verification pipeline—spanning asset identification, threat modeling, test generation, simulation analysis, formal verification, and countermeasure reasoning. Using the open-source NVDLA accelerator as a case study, the authors construct a multidimensional trustworthy verification framework that synergistically combines simulation, formal methods, and benchmark evaluation. Empirical results demonstrate that the proposed approach significantly accelerates the verification process and validates an effective, AI-driven pathway for hardware security assurance through the fusion of multiple evidence sources.
This work addresses the current lack of a systematic understanding of the capabilities of AI sandboxes in ensuring safety, security, and regulatory compliance, particularly within physical AI and cyber-physical systems. It proposes the first unified, assurance-oriented framework for AI sandboxes, introducing a formal boundary definition, a comprehensive sandbox taxonomy, a threat model targeting the assurance mechanisms themselves, and a quantifiable evaluation methodology spanning six dimensions—including fidelity and controllability. Through formal modeling, threat analysis, and multi-case validation, the study clarifies what aspects of AI behavior can be effectively tested in sandboxes, which risk categories can be meaningfully controlled, and what forms of evidence such environments can generate to support safety and compliance claims, thereby establishing foundational tools for trustworthy AI verification.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
Traditional simulation and formal verification struggle to effectively uncover security vulnerabilities in system-on-chip (SoC) designs under realistic hardware-software interactions and adversarial scenarios. This work proposes the first security-oriented hardware emulation validation framework, integrating assertion checking, coverage-driven exploration, adversarial testing, information flow tracking, fault injection, and side-channel analysis, while introducing security-aware coverage metrics. The established workflow encompasses instrumentation, stimulus generation, runtime monitoring, and forensic analysis, positioning hardware emulation as a foundational pre-silicon methodology for addressing security challenges in heterogeneous SoCs and third-party intellectual property. The paper further outlines forward-looking directions, including AI-assisted emulation, digital security twins, and chiplet-level security exploration.
In industrial cyber-physical systems (ICPS) research, demonstration objectives are often ill-defined, and technical feasibility assessment is frequently decoupled from outcome validation. Method: This paper proposes a five-level demonstration framework grounded in Maslow’s hierarchy of needs, systematically mapping demonstration goals to concrete research tasks and industrial use cases. It explicitly links work packages, verification metrics, and real-world scenarios, overcoming the vagueness and low operationality inherent in conventional Technology Readiness Level (TRL) frameworks. The approach integrates requirements engineering, modeling of software-intensive systems, and hierarchical framework design to support cross-phase requirements evolution analysis. Contribution/Results: Applied in two ICPS research projects, the framework effectively identified demonstration misalignments, refined requirement specifications, and significantly enhanced the precision, consistency, and rigor of feasibility assessment and project planning.
To address the lack of scalable, security-testing environments for real-time systems, this paper proposes SCART: the first testing framework that deeply integrates fine-grained network attack modeling into cycle-accurate simulation and digital twin environments. SCART supports multi-level fault and attack injection—including single/multi-sensor tampering, subsystem anomalies, and coordinated attacks—and enables end-to-end quantification of their impact on ML-driven navigation systems. By unifying cycle-accurate simulation, real-time fault injection, attack behavior modeling, and joint verification of ML models, SCART significantly improves UAV flight control model accuracy (32% reduction in prediction error) and anomaly detection robustness (58% reduction in false positives). The framework has been empirically validated on a physical UAV swarm under representative attack scenarios. SCART establishes a reproducible, scalable paradigm for security and reliability assessment of real-time cyber-physical systems.
This work addresses the formal specification falsification problem for safety-critical cyber-physical systems (CPS). Methodologically, it proposes a Bayesian optimization framework that integrates adaptive local surrogate modeling—built upon Gaussian processes—with domain-informed prior knowledge to guide initial sampling, and employs a customized acquisition function (e.g., an enhanced Expected Improvement) tailored for counterexample search. The core contributions are: (i) a local–global collaborative modeling strategy that balances exploration and exploitation, and (ii) a prior-driven, budget-aware optimization policy. Experiments on multiple benchmark CPS demonstrate that the approach significantly reduces simulation cost: local modeling accelerates convergence, while prior knowledge boosts counterexample detection success by up to 3× under tight budgets (<50 simulations). Moreover, the designed acquisition function substantially improves efficiency over standard Bayesian optimization baselines.
This study addresses the methodological gap in IoT security research between low-fidelity simulations and costly, hard-to-reproduce physical testbeds. To bridge this divide, the authors propose BYOT-CPS, a hybrid cyber-physical testbed that integrates real IoT devices—such as smart bulbs and cameras—with virtual networks emulated in GNS3. Designed to meet core requirements of fidelity, heterogeneity, scalability, reproducibility, and isolation, the platform implements a structured experimental environment comprising enterprise, service, attack, and monitoring zones. The system successfully demonstrates mixed physical-virtual networking, penetration testing, Mirai-style DDoS attack emulation, and fine-grained traffic monitoring. BYOT-CPS effectively narrows the gap between simulation and physical experimentation, offering a robust infrastructure for IoT security research, education, and third-party evaluation.
This work proposes a novel “betting”-based methodology for efficiently and accurately evaluating robotic performance in real-world settings where physical experimentation is constrained. By introducing betting theory into sim-to-real performance assessment—a first in the field—the approach constructs an estimator theoretically superior to Monte Carlo estimation. The method integrates control variate approximation, cross-fidelity simulation, and statistical decision rules to enable practical deployment. Experimental results demonstrate its efficacy on synthetic data and simulated environments, and it successfully infers real-world robotic grasping accuracy with significantly fewer physical trials while improving evaluation precision.
This study addresses a critical security vulnerability in unmanned aerial vehicle (UAV) flight control systems, demonstrating that their fail-safe mechanisms are susceptible to non-invasive voltage glitching attacks, thereby compromising system reliability. For the first time, the authors integrate ARMORY software-based fault simulation with the ChipWhisperer hardware voltage glitching platform to launch timing-precise attacks against the fail-safe logic of a UAV autopilot implemented on an STM32 microcontroller. This hardware-software co-design approach successfully identifies narrow execution windows harboring exploitable security flaws, enabling the suppression or manipulation of emergency responses—such as disabling failsafe landing protocols. The results validate the feasibility of timing-sensitive fault injection targeting flight controllers and underscore the tangible hardware-level security risks confronting cyber-physical systems.
Existing cyber ranges struggle to faithfully evaluate the timing behavior of Byzantine Fault Tolerance (BFT) protocols in cyber-physical systems, while real-world testing poses significant security risks. This work proposes ByzTwin-Range, a novel two-layer architecture that uniquely integrates digital twins with BFT protocol testing, enabling controlled experimentation, fault injection, and what-if analysis grounded in real operational data. By overcoming the temporal fidelity limitations of conventional ranges, the system incorporates industrial standards such as OPC UA, TSN, FMI/HLA, and QUIC/mTLS to support continuous validation, adaptive hardening, and differential privacy analysis. Empirical evaluations successfully uncovered critical vulnerabilities—including timeout misconfigurations, timing misjudgments, and adversarial delay attacks—and demonstrated enhanced system resilience through a secure feedback channel.
This work addresses the inefficiency of manual VCD file analysis in hardware security research, where practitioners often laboriously inspect waveform traces to identify software-hardware interface vulnerabilities. To overcome this bottleneck, the paper introduces RTL-Arrow, a novel framework that automatically transforms VCD execution traces—generated from hardware simulation—into cloud-ready, structured data frames compatible with modern data science workflows. RTL-Arrow integrates VCD parsing, structured data frame construction, and cloud-native format encapsulation, complemented by an automated compilation pipeline that produces a high-performance toolchain. Released as an open-source library, RTL-Arrow substantially lowers the barrier to hardware-software co-verification, significantly enhancing the efficiency and scalability of cross-layer vulnerability detection and analysis.