🤖 AI Summary
This study addresses the challenge in IT auditing where evidence from heterogeneous organizations is fragmented and compliance with security and regulatory controls must be assessed based on semantic adequacy rather than keyword matching, hindering automation. To tackle this, the work proposes the first system integrating Retrieval-Augmented Generation (RAG) with a multi-agent collaboration framework. The system orchestrates evidence retrieval, evaluation generation, adversarial challenge of assertions, and resolution of disagreements to produce interpretable audit recommendations that include citations, reasoning, gap analysis, and remediation guidance. Experimental validation under the ISO/IEC 27001 standard demonstrates that the approach effectively supports control interpretation and audit preparation, significantly improving efficiency. Nevertheless, human oversight remains necessary to calibrate judgments of evidentiary sufficiency.
📝 Abstract
IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational artifacts, and because audit conclusions depend on evidentiary sufficiency rather than keyword matching. We present IntelliAudit, a retrieval-grounded multi-agent system for IT audit evidence evaluation. Given a control and an evidence corpus, IntelliAudit retrieves relevant artifacts, generates an evidence-grounded assessment, challenges adverse findings, adjudicates disagreements, and produces an auditor-facing recommendation with cited evidence, rationale, missing-evidence analysis, and remediation guidance. We instantiate IntelliAudit on ISO/IEC 27001 and evaluate it across multiple simulated organizations using expert auditor review and audit-readiness user feedback. The evaluation shows that IntelliAudit can support control interpretation, evidence-grounded reasoning, and audit-preparation workflows, while also revealing the importance of human oversight for calibrating sufficiency judgments and correcting overly permissive recommendations. These results suggest that retrieval-grounded multi-agent systems can assist audit evidence review, but should remain decision-support tools rather than autonomous certification systems.