DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of realistic benchmarks and quality validation for AI agent-generated user documentation by constructing the first benchmark tailored to real-world maintenance scenarios. Leveraging documentation update tasks triggered by code changes across 292 open-source projects, it introduces a fine-grained, multi-dimensional automated evaluation metric incorporating multi-agent trajectory analysis, an abstention mechanism, and maintainer verification to systematically assess agentsโ€™ capabilities in updating or appropriately abandoning documentation. Experiments reveal that the best-performing agent achieves only 47.3 points, exposing critical deficiencies including neglecting the readerโ€™s perspective, lacking evidential support, and overlooking impact scope. This work fills a significant benchmark gap in the field and provides essential empirical foundations for improving the quality of AI-generated documentation.
๐Ÿ“ Abstract
We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).
Problem

Research questions and friction points this paper is trying to address.

Documentation Generation
AI Agents
Benchmark
Software Documentation
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Documentation Generation Benchmark
AI Agents
Software Documentation
Patch Evaluation
Failure Mode Analysis
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
F
Frances Liu
M
Manny Silva
P
Paige Calvert
A
Ayu Adiati
S
Sarah Sanders