π€ AI Summary
This work addresses the challenge of pinpointing and tracing error sources and propagation pathways within composite AI systems comprising multiple neural network components, a task that existing robustness testing methods struggle to accomplish. To this end, the paper proposes a modular robustness testing framework that enables fine-grained fault attribution through statistical perturbation injection, component-level error tracking, and cross-module propagation inference. By moving beyond conventional end-to-end evaluation paradigms, the approach supports architecture- and modality-agnostic analysis, offering a generalized methodology for dissecting system-level robustness. The frameworkβs efficacy is demonstrated in a railway track inspection system, where it reveals nuanced robustness characteristics that surpass the diagnostic granularity of standard evaluation metrics.
π Abstract
Modern AI systems increasingly comprise multiple interconnected neural networks to tackle complex inference tasks. Testing such systems for robustness and safety entails significant challenges. Current state-of-the-art robustness testing techniques, whether black-box or white-box, have been proposed and implemented for single-network models and do not scale well to multi-network pipelines. We propose a modular robustness testing framework that applies a given set of perturbations to test data. Our testing framework supports (1) a component-wise system analysis to isolate errors and (2) reasoning about error propagation across the neural network modules. The testing framework is architecture and modality agnostic and can be applied across domains. We apply the framework to a real-world autonomous rail inspection system composed of multiple deep networks and successfully demonstrate how our approach enables fine-grained robustness analysis beyond conventional end-to-end metrics.