๐ค AI Summary
AI-generated code is precipitating crises of trust, fairness, and sustainability in code review. Employing a mixed-methods approach comprising surveys and qualitative scenario analyses with 239 practitioners across 31 countries, this study investigates how reviewers govern AI-generated pull requests. We introduce the novel concept of โauditability debtโ to reframe AI code governance around the visibility of human judgment, elucidating how accountability and auditability shape the allocation of review resources. Our findings demonstrate that AI authorship does not constitute an absolute ground for rejection. Furthermore, we delineate the essential characteristics of accountability, verifiability, and ownership required for valid contributions. This work offers theoretical guidance for AI-assisted software development by reconceptualizing governance mechanisms in collaborative coding environments.
๐ Abstract
AI coding agents are moving from local code assistance into pull-based workflows, where generated contributions must be reviewed, explained, and maintained within existing project norms. Although recent work has begun to characterize AI-authored pull requests (AIPRs), less is known about how reviewers govern their entry into review, how AI authorship reshapes credibility and fairness, and what intake mechanisms protect review sustainability. We report a mixed-method questionnaire survey of 239 practitioners from 31 countries with code-review experience and varying exposure to AIPRs. In the scenarios and self-reports elicited by the survey, AI authorship was not a categorical rejection signal. Instead, respondents described review effort as conditional on whether an AIPR arrived as an accountable contribution, with bounded scope, project-grounded rationale, validation beyond Continuous Integration (CI), contributor responsiveness, and identifiable post-merge ownership. This conditional logic extended to newcomer AIPRs, where respondents emphasized visible participation in the current review process over profile-level reputation alone. Qualitative responses further described shifts in mentoring, scrutiny, deferral, and routing when human stewardship was difficult to observe. We conceptualize missing rationale, validation, and ownership as reviewability debt, the work reviewers must absorb when generated code lacks sufficient human grounding. These findings reframe AIPR governance around making human judgment observable before generated contributions consume scarce reviewer attention.