🤖 AI Summary
This study addresses the limitations of current AI auditing practices, which predominantly focus on individual models while overlooking integration risks arising from interactions among system components and between systems and their environments. Through a scoping review and reflexive thematic analysis of 58 studies, the work systematically codes existing literature to delineate, for the first time, three distinct domains of AI integration auditing: inter-component, system–environment, and multi-system. It further introduces domain-specific evaluation dimensions—compatibility, completeness, and oversight—that capture unique aspects of integrated AI systems. The findings reveal that current auditing practices remain fragmented and nascent, underscoring the critical role of accessible information and resource support in effective audit design. The paper calls for novel auditing frameworks capable of spanning components, environments, and systems to enable systematic exploration, identification, coordination, and standardization of integration-related risks.
📝 Abstract
As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments. System integration has long been central to software audits in safety-critical domains such as aerospace. However, its role in AI auditing remains underexplored. Scanning through 4,259 documents, we present a scoping review of AI audits that treat system integration as a core tenet of evaluation (n = 58). Using reflexive thematic analysis, we analyze their elements, actors, enablers, and constraints. We find that the corpus represents an emerging yet still fragmented form of AI auditing: few existing measures target integration-specific risks; large gaps remain in meeting traditional audit expectations; and access to necessary information and resources significantly influences audit design. Nonetheless, integration can be categorized across three sites (inter-component, system-environment, and multi-system), each serving the functions of risk exploration, risk determination, coordination, and procedural regularity. Deviating from other types of evaluations, these audits assess qualities specific to system integration, including compatibility, completeness, and oversight. This review calls on the AI community to prioritize system integration as a core strategy for addressing AI risk, and to develop audit practices capable of capturing failures across components, environments, and systems beyond the reach of component-level evaluation.