agent-based auditing

Designs and implements multi-agent audit systems made of persona-grounded or generative simulated user agents that enact realistic interactions at scale to probe and measure system behavior and content exposures. Builds pipelines to create, orchestrate, and manage many synthetic accounts or agents, collect interaction traces and outputs, and analyze those traces to detect, quantify, and diagnose policy, safety, or performance issues.

agent-basedauditing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Auditable Agents

Apr 07, 2026

This study addresses the critical lack of accountability in large language model (LLM) agents following external actions, which hinders traceability and responsibility attribution. The work introduces the first systematic framework for agent auditability, articulating five core dimensions and proposing an “Auditability Card” to standardize assessment. It further develops a full-lifecycle auditing architecture integrating detection, enforcement, and recovery mechanisms. Leveraging runtime intervention, tamper-proof logging, and log-recovery techniques—validated through ecosystem-wide security evaluations and controlled experiments—the study identifies 617 security flaws across mainstream open-source LLM agent projects. Experimental results demonstrate that a pre-execution mediation layer incurs only 8.3 milliseconds of overhead and that partial reconstruction of accountability-critical information remains feasible even in the absence of complete logs.

accountabilityauditabilityevidence integrity

This study addresses the lack of structured review in prompt specifications for multi-agent large language model (LLM) systems, which often leads to consistency defects. Conducted within the AEGIS seven-channel orchestration framework, the work performs nine rounds of iterative agent-driven audits on approximately 7,150 lines of prompt specifications, employing a checklist adapted from Weinberg and Freedman’s methodology. The paper introduces the first taxonomy of seven defect categories specific to LLM multi-agent prompt specifications, uncovers non-monotonic convergence behavior, and establishes a reproducible, finalized audit checklist. Leveraging automated auditing by Claude sub-agents, checklist-guided walkthroughs, STRIDE threat modeling, and multi-round expanded-scope reviews, the study identifies 51 consistency defects—many of which were missed in initial single-file inspections but subsequently detected in later rounds—ultimately achieving zero residual defects.

audit convergenceconsistency defectsmulti-agent LLM systems

Existing black-box auditing methods struggle to disentangle user attributes from behavior while preserving behavioral authenticity, limiting causal understanding of personalization algorithms. This study introduces the first large-scale algorithmic audit using generative AI agents: leveraging real survey data, it constructs synthetic accounts with fixed personas that autonomously interact on the X platform, enabling counterfactual experiments through controlled manipulation of visible user attributes. This approach effectively decouples attributes from behavior, supporting fine-grained, scalable, and behaviorally authentic analysis. An empirical deployment of 1,120 agents reveals that the platform’s algorithms significantly amplify toxic, polarizing, and right-leaning political content, with amplification effects varying by user ideology; furthermore, the influence of demographic signals exhibits substantial heterogeneity across subpopulations.

black-box auditingcausal understandingpersonalization algorithms

This work addresses the critical yet often overlooked security risks in large language model (LLM) agent systems, which frequently stem from software stack components such as tool code, deployment configurations, and permission settings—not merely from the underlying models. To this end, we present the first dedicated security analysis framework tailored for LLM-based agent applications. Our system integrates dataflow analysis, credential detection, structured configuration parsing, and permission risk assessment to precisely identify diverse vulnerabilities across tool functions, prompts, and deployment artifacts, outputting results in the standardized SARIF format. Evaluated on 22 real-world samples containing 42 annotated vulnerabilities, our approach successfully detects 40 true positives with only 6 false positives, substantially outperforming general-purpose static application security testing (SAST) tools while completing each scan in under one second.

credential exposureLLM agentprivilege misconfiguration

This work addresses the lack of an efficient, systematic, and comprehensive evaluation framework for modern AI agents, which struggles to keep pace with rapidly evolving deployment infrastructures. To this end, we propose A²E, an end-to-end agent auditing engine that introduces the novel Agent Task Protocol to decouple tasks from underlying frameworks, enabling unified integration of diverse evaluation scenarios. Leveraging automated instrumentation and standardized execution traces, A²E establishes a multidimensional assessment methodology encompassing execution efficiency, tool utilization, task planning, and error recovery. Experimental results reveal substantial performance variations across different model–framework combinations on various tasks, with no single configuration consistently dominating others. These findings underscore the necessity of systematic benchmarking and provide empirical grounding for the co-optimization of models and agent frameworks.

agent evaluationexecution tracesharness ecosystem

Latest Papers

What's happening recently
View more

Existing collaborative auditing methods for large language model (LLM) agents struggle to simultaneously capture pending tasks, responsibility attribution, and evidence of state transitions, leading to verification blind spots. This work proposes iCORE, a novel framework that unifies the modeling of cooperative structure, obligation allocation, and verifiable evidence through a coupled representation of a cooperation graph \(G\), an obligation graph \(Q\), and an audit mapping \(\Pi\), thereby enabling formal verification of collaborative processes. iCORE introduces local-to-global rationality reasoning and regret-bound analysis for obligation assignment, allowing rigorous certification of task rationality and agent stability. Experimental results demonstrate that iCORE-Audit improves trajectory quality by 11.5% and 26.4%, and boosts end-task performance by 15.1% and 31.0%, in controlled and real-world LLM environments, respectively.

agent-assignment stabilityauditability gapcooperation-obligation coupling

This study addresses the lack of systematic analysis and actionable controls linking large language model (LLM) agent security threats to real-world financial regulations. It presents the first mapping of six categories of LLM agent risks to regulatory obligations in the U.S. and EU, and proposes four scalable compliance architectures centered on auditability, authorization, and boundary enforcement. Key technical contributions include agent-to-agent (A2A) compliance orchestration, audit-driven Grounded-RAG, case ID propagation, and reasoning-boundary de-identification proxies. Empirical evaluation demonstrates that the framework reduces manual processing from multiple days to same-day handling, automates approximately 80% of use cases, and uncovers two types of control failures detectable only through internal audit, as well as one category of legitimate applicants erroneously rejected.

agent securityauditabilitycompliance

Hot Scholars

KL

Kristian Lum

University of Chicago
Algorithmic FairnessBayesian StatisticsPopulation EstimationMicrosimulation
RH

Ruqi Huang

Tsinghua Shenzhen International Graduate School
3D Computer VisionShape AnalysisGeometry Processing
MH

Martin Hirzel

IBM Research
Programming LanguagesData ManagementAI
CT

Chenhao Tan

University of Chicago
Human-centered AICommunication & IntelligenceScientific DiscoveryAI alignment