Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance

📅 2026-05-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of structured review in prompt specifications for multi-agent large language model (LLM) systems, which often leads to consistency defects. Conducted within the AEGIS seven-channel orchestration framework, the work performs nine rounds of iterative agent-driven audits on approximately 7,150 lines of prompt specifications, employing a checklist adapted from Weinberg and Freedman’s methodology. The paper introduces the first taxonomy of seven defect categories specific to LLM multi-agent prompt specifications, uncovers non-monotonic convergence behavior, and establishes a reproducible, finalized audit checklist. Leveraging automated auditing by Claude sub-agents, checklist-guided walkthroughs, STRIDE threat modeling, and multi-round expanded-scope reviews, the study identifies 51 consistency defects—many of which were missed in initial single-file inspections but subsequently detected in later rounds—ultimately achieving zero residual defects.
📝 Abstract
Prompt specifications for multi-agent large language model (LLM) systems carry data contracts and integration logic across many interdependent files but are rarely subjected to structured-inspection rigor. This paper reports a single-system empirical case study of iterative, agent-driven auditing applied to AEGIS (Autonomous Engineering Governance and Intelligence System), a production seven-lane orchestration pipeline whose prompt-specification surface comprises approximately 7150 lines: 6907 across seven lane PROMPT.md files and a 245-line shared Ticket Contract. Nine sequential audit rounds, executed by Claude sub-agents using a checklist-driven walkthrough adapted from Weinberg and Freedman, surfaced 51 prompt-specification consistency defects, distinct from the 51 STRIDE-categorized adversarial code findings reported in the companion preprint. Per-round counts were 15, 8, 12, 2, 8, 1, 4, 1, and 0. We report a seven-category post-hoc defect taxonomy with explicit coding rules, observed non-monotonic convergence consistent with cascading edits and audit-scope expansion, and an audit protocol distilled from the study, with the final locked checklist released as a reproducibility appendix. Single-file review missed defect classes that were surfaced only by later expanded-scope rounds in this system. The same LLM family authored and audited the specifications; replication with dissimilar models and human reviewers is required before generalization.
Problem

Research questions and friction points this paper is trying to address.

prompt specification
multi-agent LLM systems
structured inspection
consistency defects
audit convergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterative auditing
prompt engineering
multi-agent LLM systems
defect taxonomy
audit convergence
E
Elias Calboreanu
1 Swift (North) AI Lab, The Swift Group, LLC, Maryland, USA; 2 Capitol Technology University, Laurel, MD 20708, USA