Score
Designs and implements frameworks that systematically test detectors for failure modes (including noise and other perturbations), producing empirical robustness evaluations often via automated test generation and targeted stress tests. Analyzes observed weaknesses and develops targeted hardening measures — such as preprocessing, model changes, or data-augmentation and mitigation procedures — to eliminate or reduce specific failure modes and verify improved robustness.
To address the insufficient speed, reliability, and maintainability of testing in modern software systems, this paper designs and implements a modular automated testing framework that deeply integrates Cucumber-BDD with Java. The framework introduces a novel natural-language-driven test design and engineering implementation co-development mechanism, supporting dynamic environment adaptation, reusable component-based architecture, and end-to-end automated reporting with closed-loop feedback. It integrates Selenium, TestNG, Maven, and Jenkins to enable seamless embedding into CI/CD pipelines. Empirical evaluation demonstrates a reduction of manual testing effort by over 40%, a 35% improvement in defect detection rate, and a 50% decrease in script maintenance cost. These outcomes significantly enhance agility in iterative development and streamline multi-environment one-click deployment efficiency.
Traditional security testing tools deployed in CI/CD pipelines lack adaptability and struggle to effectively integrate program structure with dynamic feedback, resulting in low detection efficiency and high false-positive rates. This work presents a systematic survey of adaptive and AI-enhanced security testing approaches, introducing for the first time the notion of “structural-adaptive disconnection” to highlight the systemic misalignment between program structure representations and adaptive mechanisms. It advocates for incorporating human-in-the-loop signals into a closed-loop model refinement process. By synthesizing techniques from static and dynamic analysis, feedback-driven fuzzing, large language models, and code property graphs (CPGs), the study analyzes 55 high-quality research efforts, identifies five key open challenges, and proposes a unified research agenda for semantic-aware, feedback-driven, and multi-language-supported security testing frameworks.
This work addresses the challenge of pinpointing and tracing error sources and propagation pathways within composite AI systems comprising multiple neural network components, a task that existing robustness testing methods struggle to accomplish. To this end, the paper proposes a modular robustness testing framework that enables fine-grained fault attribution through statistical perturbation injection, component-level error tracking, and cross-module propagation inference. By moving beyond conventional end-to-end evaluation paradigms, the approach supports architecture- and modality-agnostic analysis, offering a generalized methodology for dissecting system-level robustness. The framework’s efficacy is demonstrated in a railway track inspection system, where it reveals nuanced robustness characteristics that surpass the diagnostic granularity of standard evaluation metrics.
This study addresses the limitations of experience-dependent prompt design and the lack of systematic insights into failure mechanisms in large language model (LLM) vulnerability analysis. To overcome these challenges, this work proposes a failure-driven prompt optimization paradigm that systematically analyzes recurring failure patterns—such as false positives and reasoning errors—in the DVJA dataset to reconstruct targeted prompting strategies. Furthermore, it introduces a novel evidence-based evaluation framework grounded in specific failure cases, superseding conventional comparisons based on aggregated metrics. The proposed approach is validated on the Juliet test suite, demonstrating cross-model generalizability by significantly enhancing the reliability of LLM-based vulnerability detection while distilling reusable prompt engineering design principles.
This work addresses the challenge of generating high-coverage, diverse robustness test cases for microservice APIs, where anomalous inputs can trigger cascading failures. The authors propose an automated test generation approach leveraging large language models (LLMs), integrating existing mutation taxonomies into prompt design and introducing two novel strategies: Guided and GuidedFewShot. Evaluations across three open-source LLMs (14B–70B parameters) and seven prompting strategies produced 663 test cases on mono- and multilingual microservice systems. Results demonstrate that prompting strategy exerts a greater influence on test diversity than model size; GuidedFewShot achieves the highest single-run fault coverage—detecting 5 out of 9 and 8 out of 14 failure modes in the two systems, respectively—with low cross-model similarity. Moreover, combining multiple prompting strategies with a single LLM surpasses the effectiveness of multi-model ensembles.
This study addresses the limitations of traditional statistical fault localization (SFL), which relies solely on code execution traces and often fails to accurately pinpoint root causes. To overcome this, the authors systematically incorporate execution features—such as data flow, variable values, and branch conditions—extracted via the EFDD tool from the Tests4Py dataset. They train project-specific random forest models and map feature importance back to source code lines, integrating these insights with classical SFL formulas to enhance localization accuracy. Rigorous evaluation is conducted using a confounder-adjusted mixed-effects model and paired statistical tests. Experimental results demonstrate that the proposed approach significantly improves the accuracy of reference patches while reducing inspection effort at both line and function levels, confirming its robustness and practicality across multiple dimensions.
Current endpoint detection systems exhibit insufficient robustness against adaptive code variants, making it difficult to accurately assess their adversarial resilience. This work proposes ShellForge, a novel framework that, for the first time, integrates multi-objective genetic algorithms with real-time feedback from antivirus (AV) and endpoint detection and response (EDR) systems to automatically generate functionally equivalent yet structurally diverse remote command execution payloads. By leveraging techniques such as syntactic transformation, encoding schemes, and structural rearrangement, ShellForge establishes a reproducible benchmark for evaluating the robustness of endpoint detection mechanisms. The framework exposes significant vulnerabilities in both signature-based and behavior-based detection approaches when confronted with adversarial variants, thereby providing an empirical foundation for improving defensive architectures.
研究通过构建自动驾驶车辆漏洞分类及结合大型语言模型的LLMSec-AV框架,以提高软件弱点检测效果,超越了基于规则的传统工具。
This work addresses the lack of security and verifiability in large language models for project-level code generation by proposing and evaluating an end-to-end Detect–Repair–Verify (DRV) workflow tailored for multilingual web applications. The approach generates executable code at three granularities—project, requirement, and function—integrating static and dynamic analysis, automated repair, and test-driven verification. Under unified resource constraints, the study systematically compares generative, single-round, and iterative variants of DRV. It introduces the first project-level benchmark for secure code generation that supports multiple prompting granularities, enabling a comprehensive evaluation of DRV’s efficacy. The findings reveal limitations in using vulnerability reports to guide repairs and identify common post-repair failure modes such as regressions and semantic drift. Experimental results demonstrate that the iterative DRV variant significantly enhances security while preserving functional correctness.
High-performance software systems accumulate latent reliability risks through aggressive optimizations; superficial performance metrics (e.g., high cache hit rates) mask underlying bottlenecks, leading to load amplification and cascading failures upon degradation. Current reliability engineering emphasizes reactive mitigation, lacking proactive identification and prevention of optimization-induced fragility. Method: We propose the first systematic framework for optimization-risk management, introducing a novel quantitative model and the Latent Risk Index (LRI). Our tripartite defense architecture—HYDRA (risk detection), RAVEN (perturbation-based validation), and APEX (risk-aware optimization)—integrates mathematical modeling, six categories of optimization-sensitive perturbation testing, and high-precision online monitoring. Contribution/Results: Experiments demonstrate 89.7% risk detection rate, >92.9% monitoring accuracy, 69.1% reduction in MTTR, annual cost savings of $1.44M, and a payback period of just 3.2 months.