detector failure-mode analysis and targeted hardening

Designs and implements frameworks that systematically test detectors for failure modes (including noise and other perturbations), producing empirical robustness evaluations often via automated test generation and targeted stress tests. Analyzes observed weaknesses and develops targeted hardening measures — such as preprocessing, model changes, or data-augmentation and mitigation procedures — to eliminate or reduce specific failure modes and verify improved robustness.

detectorfailure-modeanalysisand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Designing and Implementing Robust Test Automation Frameworks using Cucumber-BDD and Java.

Apr 24, 2025
SS
Srikanth Srinivas
🏛️ The University of Texas at Dallas

To address the insufficient speed, reliability, and maintainability of testing in modern software systems, this paper designs and implements a modular automated testing framework that deeply integrates Cucumber-BDD with Java. The framework introduces a novel natural-language-driven test design and engineering implementation co-development mechanism, supporting dynamic environment adaptation, reusable component-based architecture, and end-to-end automated reporting with closed-loop feedback. It integrates Selenium, TestNG, Maven, and Jenkins to enable seamless embedding into CI/CD pipelines. Empirical evaluation demonstrates a reduction of manual testing effort by over 40%, a 35% improvement in defect detection rate, and a 50% decrease in script maintenance cost. These outcomes significantly enhance agility in iterative development and streamline multi-environment one-click deployment efficiency.

Addressing test data management and CI/CD integration challengesDeveloping robust test automation frameworks for complex software systemsEnhancing communication between technical and non-technical team members

Traditional security testing tools deployed in CI/CD pipelines lack adaptability and struggle to effectively integrate program structure with dynamic feedback, resulting in low detection efficiency and high false-positive rates. This work presents a systematic survey of adaptive and AI-enhanced security testing approaches, introducing for the first time the notion of “structural-adaptive disconnection” to highlight the systemic misalignment between program structure representations and adaptive mechanisms. It advocates for incorporating human-in-the-loop signals into a closed-loop model refinement process. By synthesizing techniques from static and dynamic analysis, feedback-driven fuzzing, large language models, and code property graphs (CPGs), the study analyzes 55 high-quality research efforts, identifies five key open challenges, and proposes a unified research agenda for semantic-aware, feedback-driven, and multi-language-supported security testing frameworks.

adaptive testingCI/CDprogram analysis

This work addresses the challenge of pinpointing and tracing error sources and propagation pathways within composite AI systems comprising multiple neural network components, a task that existing robustness testing methods struggle to accomplish. To this end, the paper proposes a modular robustness testing framework that enables fine-grained fault attribution through statistical perturbation injection, component-level error tracking, and cross-module propagation inference. By moving beyond conventional end-to-end evaluation paradigms, the approach supports architecture- and modality-agnostic analysis, offering a generalized methodology for dissecting system-level robustness. The framework’s efficacy is demonstrated in a railway track inspection system, where it reveals nuanced robustness characteristics that surpass the diagnostic granularity of standard evaluation metrics.

compound AI systemserror attributionmodular analysis

This study addresses the limitations of experience-dependent prompt design and the lack of systematic insights into failure mechanisms in large language model (LLM) vulnerability analysis. To overcome these challenges, this work proposes a failure-driven prompt optimization paradigm that systematically analyzes recurring failure patterns—such as false positives and reasoning errors—in the DVJA dataset to reconstruct targeted prompting strategies. Furthermore, it introduces a novel evidence-based evaluation framework grounded in specific failure cases, superseding conventional comparisons based on aggregated metrics. The proposed approach is validated on the Juliet test suite, demonstrating cross-model generalizability by significantly enhancing the reliability of LLM-based vulnerability detection while distilling reusable prompt engineering design principles.

Failure ModesLarge Language ModelsPrompt Engineering

This work addresses the challenge of generating high-coverage, diverse robustness test cases for microservice APIs, where anomalous inputs can trigger cascading failures. The authors propose an automated test generation approach leveraging large language models (LLMs), integrating existing mutation taxonomies into prompt design and introducing two novel strategies: Guided and GuidedFewShot. Evaluations across three open-source LLMs (14B–70B parameters) and seven prompting strategies produced 663 test cases on mono- and multilingual microservice systems. Results demonstrate that prompting strategy exerts a greater influence on test diversity than model size; GuidedFewShot achieves the highest single-run fault coverage—detecting 5 out of 9 and 8 out of 14 failure modes in the two systems, respectively—with low cross-model similarity. Moreover, combining multiple prompting strategies with a single LLM surpasses the effectiveness of multi-model ensembles.

API input validationfailure-mode diversityLLM-generated tests

Latest Papers

What's happening recently
View more

This study addresses the limitations of traditional statistical fault localization (SFL), which relies solely on code execution traces and often fails to accurately pinpoint root causes. To overcome this, the authors systematically incorporate execution features—such as data flow, variable values, and branch conditions—extracted via the EFDD tool from the Tests4Py dataset. They train project-specific random forest models and map feature importance back to source code lines, integrating these insights with classical SFL formulas to enhance localization accuracy. Rigorous evaluation is conducted using a confounder-adjusted mixed-effects model and paired statistical tests. Experimental results demonstrate that the proposed approach significantly improves the accuracy of reference patches while reducing inspection effort at both line and function levels, confirming its robustness and practicality across multiple dimensions.

Developer Inspection EffortExecution FeaturesFault Localization Accuracy

Current endpoint detection systems exhibit insufficient robustness against adaptive code variants, making it difficult to accurately assess their adversarial resilience. This work proposes ShellForge, a novel framework that, for the first time, integrates multi-objective genetic algorithms with real-time feedback from antivirus (AV) and endpoint detection and response (EDR) systems to automatically generate functionally equivalent yet structurally diverse remote command execution payloads. By leveraging techniques such as syntactic transformation, encoding schemes, and structural rearrangement, ShellForge establishes a reproducible benchmark for evaluating the robustness of endpoint detection mechanisms. The framework exposes significant vulnerabilities in both signature-based and behavior-based detection approaches when confronted with adversarial variants, thereby providing an empirical foundation for improving defensive architectures.

Behavior-based DetectionCode TransformationEndpoint Detection

This work addresses the lack of security and verifiability in large language models for project-level code generation by proposing and evaluating an end-to-end Detect–Repair–Verify (DRV) workflow tailored for multilingual web applications. The approach generates executable code at three granularities—project, requirement, and function—integrating static and dynamic analysis, automated repair, and test-driven verification. Under unified resource constraints, the study systematically compares generative, single-round, and iterative variants of DRV. It introduces the first project-level benchmark for secure code generation that supports multiple prompting granularities, enabling a comprehensive evaluation of DRV’s efficacy. The findings reveal limitations in using vulnerability reports to guide repairs and identify common post-repair failure modes such as regressions and semantic drift. Experimental results demonstrate that the iterative DRV variant significantly enhances security while preserving functional correctness.

Code SecurityDetect-Repair-VerifyEmpirical Study

Detecting and Preventing Latent Risk Accumulation in High-Performance Software Systems

Oct 04, 2025
JA
Jahidul Arafat
🏛️ Auburn University | Oracle | Orange Business Development Limited | Bangladesh University of Professionals | Bangladesh Army International University of Science and Technology | Green University of Bangladesh

High-performance software systems accumulate latent reliability risks through aggressive optimizations; superficial performance metrics (e.g., high cache hit rates) mask underlying bottlenecks, leading to load amplification and cascading failures upon degradation. Current reliability engineering emphasizes reactive mitigation, lacking proactive identification and prevention of optimization-induced fragility. Method: We propose the first systematic framework for optimization-risk management, introducing a novel quantitative model and the Latent Risk Index (LRI). Our tripartite defense architecture—HYDRA (risk detection), RAVEN (perturbation-based validation), and APEX (risk-aware optimization)—integrates mathematical modeling, six categories of optimization-sensitive perturbation testing, and high-precision online monitoring. Contribution/Results: Experiments demonstrate 89.7% risk detection rate, >92.9% monitoring accuracy, 69.1% reduction in MTTR, annual cost savings of $1.44M, and a payback period of just 3.2 months.

Detecting latent risks in high-performance software systems with hidden vulnerabilitiesPreventing catastrophic fragility masked by exceptional performance optimizationsTransforming reliability engineering from reactive to proactive risk management

Hot Scholars

PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI
AC

Ahmad Chaddad

Professor @ School of Artificial Intelligence, GUET; LIVIA-ETS
Artificial intelligenceradiomic and radio-genomicsSignal & Image ProcessingElectrical & Electronic System
DM

David Martens

University of Antwerp
Data miningExplainable AIMining Behavioral DataData Science Ethics
AP

Ali Payani

Cisco Systems, Georgia Tech
Natural Language ProcessingLogic and ReasoningData Efficient AIFederated Learning