DoGBench

Research paper

DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

  • Frances LiuPromptless
  • Manny SilvaPromptless / Doc Detective
  • Paige CalvertHelm
  • Ayu AdiatiMautic
  • Sarah SandersPostHog

Abstract

We introduce DoGBench (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).

1 Introduction

User-facing documentation is the main public description of what a software product is, what capabilities it offers, and when and how to use them. For closed-source products in particular, documentation may be the only structured source an AI agent can use to understand and operate the product. People increasingly rely on AI agents to find, choose, and use products. Documentation therefore affects whether an agent uses a product correctly and whether the agent considers the product for the user’s task at all.

Producing this content requires more than translating implementation details into prose. The work demands judgment about where information belongs, what readers are trying to accomplish, and what they already know. Teams increasingly use agents to write documentation, but no one has yet systematically evaluated the user-facing documentation that agents produce.

Prior work evaluates code-facing documentation rather than user-facing documentation (Section 2.1). That work covers function-level docstrings, repository-level code summaries, and internal developer documentation, and the work assumes that a human already decided that the documentation needs an update. No existing documentation benchmark tests whether an agent can tell when to leave the documentation alone.

We introduce DoGBench, a benchmark of 292 items from open source projects. Each item gives an agent a pre-change repository and a trigger, which is either a pull request or a reported documentation gap. The agent first decides whether the trigger calls for a documentation change. For 205 items, the correct response is a patch. For the other 87, the correct response is to abstain. Task-specific rubrics, validated with project maintainers, score each patch. The rubrics do not reward similarity to the documentation that humans merged. Depending on the task, a correct patch may revise existing guidance, create and register a new page, move content, or remove stale or redundant content. Every item starts from an existing product and its then-current documentation. The benchmark therefore does not cover writing a product’s documentation from scratch or redesigning an information architecture without constraints.

We use a random, stratified 117-item held-out split for primary evaluation. The other 175 items form a public development split. We use the public split and the full 292 items only for robustness analyses.

We evaluated seven agent lanes. The highest-scoring agent reached 47.3 out of 100 on the held-out split. Only 6.1% of submissions contained fabricated content. The more common failure was a plausible patch that left the reader unable to finish the task. Of 1,267 submissions, 45.5% had a task-completion gap. Section 7 traces these failures to how agents investigate. In 36.0% of submissions, agents explained product interfaces without checking how readers use them. In 33.1%, agents stopped before finding decisive evidence. In 30.1%, agents edited the first plausible documentation surface and missed other surfaces the change affected.

2.1 Documentation generation benchmarks

Existing documentation benchmarks focus on code-facing documentation. CodeSearchNet supplied a corpus of paired functions and documentation for semantic code search, and CodeXGLUE used CodeSearchNet-derived data for code summarization (Husain et al., 2019; Lu et al., 2021). More recent benchmarks study docstring updates after code changes (CoDocBench) and repository-level internal documentation (CodeWikiBench) (Pai et al., 2025; Nguyen Hoang et al., 2026). Like DoGBench, SWD-Bench builds its tasks from pull requests, but it scores repository-level documentation by how well a model can use that documentation to answer questions about the repository’s functionality (Wang et al., 2026). None of these benchmarks asks the model to decide whether documentation needs an update.

2.2 Scoring open-ended edits

SWE-bench, which also draws from open source repositories, is widely used to evaluate code generation (Jimenez et al., 2024). OpenAI has since questioned the validity of its Verified subset because narrow tests can reject correct alternative solutions (OpenAI, 2026a). That risk is larger for documentation because many different edits can satisfy the same reader need. DoGBench therefore scores each patch against requirements drawn from the triggering change and the pre-change repository rather than against similarity to the merged human patch.

3 Benchmark Design

Each item in the benchmark gives the agent a pre-change repository and a trigger. The trigger is either a code pull request or a user-reported documentation gap, often a GitHub issue. This setup mirrors how maintainers work. A maintainer either ships documentation changes with a feature or updates the documentation in response to a community issue. The agent must either return a patch that edits the documentation or abstain from making documentation changes when no user-facing change is needed. Figure 1 summarizes the item-construction and evaluation pipeline.

Overview of the DogBench task and evaluation pipeline
Figure 1. Overview of DoGBench. Each item begins with a real trigger event and a frozen, identity-masked snapshot of the pre-change repository. A documentation agent either abstains or emits a patch, and task-specific rubrics score the patch. DoGBench reports decision correctness and documentation quality separately. Its composite score combines delivered patch quality with abstention recall.

3.1 Dataset construction

The benchmark contains 292 items: 205 require a documentation change, and 87 require abstention. We built them from three pools. The first holds 90 changes initially sampled as likely abstention cases. The second holds 136 code-triggered documentation updates, and the third holds 66 explicit user-reported documentation gaps. During final adjudication, we reclassified three of the 90 likely abstention cases as requiring documentation, which left 87 abstention items. The final set therefore has 139 code-triggered documentation items, 66 items triggered by reported documentation gaps, and 87 abstention items. Project maintainers reviewed 42 items from repositories such as Helm, Doc Detective, Mautic, and PostHog. Appendix [app:dataset-details] in the full PDF gives additional selection details and two case studies.

Constructing the trigger

For a pull request that ships code and documentation together, we remove the documentation changes and give the agent only the code change. For a pull request that changes only documentation, we build the trigger from linked issues and discussions. We add sanitized versions of the source pull request’s title and description. Both methods keep the documentation need and the reason for it but hide how the maintainer wrote the documentation.

We selected open source repositories with English-language documentation across diverse ecosystems. We exclude the following:

  • reverts, release-only version bumps, and merge or sync pull requests

  • pure refactors

  • changes where the connection between trigger and the documentation need is unclear

  • items that need context unavailable in the public repository or trigger

We also exclude bot-authored pull requests are excluded unless a human maintainer reviewed and revised the change before merging. Appendix [app:dataset-details] in the full PDF gives the full rules and known limitations.

3.2 Contamination controls

Because the source events are public, contamination can happen if a model has seen the merged documentation during training or if an agent finds it while running. We probe training-time exposure with source-event-date and repository-footprint ablations (Section 6.3) (Sainz et al., 2023). To prevent execution-time exposure, agents work in fresh Docker containers with identity-masked repositories and no network access except to the model provider. We also inspect agent trajectories for attempts to reach the merged human patch. A conservative overlap detector flags suspicious similarity to the merged documentation for manual review. We found no confirmed case of copying. We reviewed every retained detector alert and judged each one a false positive.

4 Evaluation Protocol

For each item that requires a documentation update, we score the submitted patch against a task-specific rubric. Across the 205 documentation-needed items, the rubrics contain 3,273 criteria, including 798 P0 criteria.

We design each criterion to test one observable review decision. For example, suppose a change lets a Helm values file to contain multiple YAML documents. Separate criteria can then check the documentation for three statements: documents are processed in order, later values take precedence, and nested maps are merged recursively. A broad criterion such as “explains multi-document values well” would not be testable enough.

4.1 Criterion types and priorities

Each criterion use one of three scoring types:

  • A requirement always applies and receives a binary Pass or Fail verdict.

  • A conditional criterion applies only when the patch meets its condition. We call a conditional criterion triggered when it applies.

  • A deduction-only guardrail prohibits content such as a fabricated command or an unsafe recovery step. Avoiding the prohibited content earns no credit, and introducing it costs a deduction.

Requirements and triggered conditional criteria receive only Pass or Fail verdicts, with no partial credit. Untriggered conditional criteria are left out of scoring.

Each criterion also has a priority from P0 to P3, which is separate from its scoring type. P0 is reserved for a defect that blocks the patch on its own: failing that criterion alone would require revision under the benchmark’s standard. Typical P0 defects include the following:

  • materially misstating product behavior

  • omitting information the reader needs to complete the central reader task

  • giving an unsafe or destructive instruction

  • inventing a public interface

  • leaving an essential maintained documentation surface contradictory or unusable

Optional examples, secondary edge cases, stylistic preferences, and exact wording do not qualify as P0 merely because they would improve the patch. A patch is P0-clean when no applicable P0 criterion failed. P0-clean diagnoses only critical defects and does not a prediction whether a maintainer would merge the patch unchanged.

4.2 Score construction

Let RR contain all requirements and triggered conditional criteria, and let GG contain all violated deduction-only guardrails. Each criterion in RR contributes one point if it passes, and each violated guardrail deducts one point. For |R|>0|R|>0, the uncapped patch score is q̃=100max⁡(0,∑i∈R𝟏[i passes]−|G|)|R|.\widetilde q = 100\,\frac{\max\!\left(0,\sum_{i\in R}\mathbf{1}[i\text{ passes}]-|G|\right)} {|R|}. An untriggered conditional criterion counts in neither the numerator nor the denominator. A guardrail that the patch does not violate is also left out, so a patch earns no points merely for avoiding an optional risk. If RR is empty, the score is zero.

Let B=1B=1 when a P0 requirement or triggered conditional criterion fails, or when a P0 guardrail is violated. Otherwise, let B=0B=0. The reported patch score is q={min⁡(q̃,60),B=1,q̃,B=0.q = \begin{cases} \min(\widetilde q,60), & B=1,\\ \widetilde q, & B=0. \end{cases} On a documentation-needed item, an empty patch or an abstention also scores zero. We set the 60-point ceiling as an evaluation policy. We did not estimate it from maintainer editing time or acceptance decisions. The ceiling prevents success on many secondary criteria from averaging away a critical defect.

4.3 Rubric construction

Research and synthesis

We build each rubric through repository research and an LLM-council process. The merged human patch is not included among the candidate patches supplied to the rubric agents. A rubric research agent inspects the triggering change and the repository to identify what users need to know and the evidence that supports each requirement. The research agent has internet access and may encounter the merged human patch, but every criterion must have independent supporting evidence; the human patch alone cannot justify a criterion. A synthesis stage turns these findings into criteria with explicit passing and failing conditions. (Mei et al., 2026) also explore this research-to-criteria approach.

Differential review

The differential stage compares anonymous candidate patches side by side to find editorial choices that the draft rubric does not yet cover. This stage follows the observation that inspecting model outputs can help refine evaluation criteria (Shankar et al., 2024). For each uncovered difference, the rubric agent checks the research findings and gathers more evidence where needed. It then decides whether the difference matters to the reader. A difference between candidates only raises a question and does not establish what is correct. Trivial or neutral differences do not become criteria. The rubric agent also records supported documentation needs that no candidate meets. These findings are then used to revise the draft rubric.

Audit and debate

A second model, from a different model family, audits the revised criteria and their priorities. When the rubric author and the auditor disagree, they revisit the evidence and exchange arguments. Together they revise or remove any criterion that cannot be justified. The debate ends when the auditor accepts the revised rubric or the exchange reaches its configured limit. The longest saved debate we inspected ran 28 messages after the opening audit.

4.4 Human validation

We compare the resulting rubrics with independently collected maintainer criteria on a reviewed subset. Separately, we compare the Pass or Fail verdicts of a scoring model with human judgments. The first check asks whether the rubric captures the requirements that maintainers consider important. The second asks whether the scoring model applies those requirements correctly.

On 42 maintainer-reviewed items, the automatic documentation-need gate agrees with the maintainers on 41, with one false positive. On 21 documentation-needed items with completed criterion alignment, mean priority-weighted recall against maintainer criteria is 0.892. In a separate study of 20 items and 330 criteria, the scoring model agrees with the post-adjudication human reference on 310 criteria (93.9%). One paper author resolved disagreements in the human reference after review, so this figure is not blinded agreement between two humans. Appendix [app:evaluation-details] in the full PDF describes both validation studies.

5 Experimental Setup

We evaluate seven agent lanes:

  • GLM 5.2, Qwen3.8 Max, and Kimi K2.7 Code with OpenCode

  • Claude Opus 4.8 and Claude Sonnet 4.6 with Claude Code

  • GPT-5.5 and GPT-5.6 Sol with Codex

Every agent runs each item in a fresh Docker environment with the same inputs. The inputs are the pre-change code and documentation repository, plus the trigger: a code diff or a reported documentation gap. Each agent–item pair runs once and must either return a patch or abstain. Agents can use command-line tools to inspect and edit the repository, but the environment has no network access except to the model providers.

Why the environment is sealed

Internet access can be valuable for documentation agents in ordinary use, but DoGBench blocks it to prevent contamination. Because the source events and merged documentation are public, a connected agent could retrieve the answer instead of solving the task from the supplied evidence. In an audit of an earlier version of the benchmark, we found that 31% of runs retrieved the source pull request despite identity masking. A further 6% copied directly from other agents’ earlier trajectories because the agents shared a host.

Data splits

The release has two splits. The development split holds 175 items: 123 documentation-needed and 52 abstention. The held-out split holds 117 items: 82 documentation-needed and 35 abstention. The random split draw comes from a procedure stratified by class, source pool, task type, and repository. The split was drawn without access to measured scores. Development items include inputs, rubrics, and reference artifacts. Held-out rubrics, references, and item-level scores remain private. We report only the seven reproducible agents in this paper.

6 Results

Table 1 reports the composite score on the 117-item held-out split. The composite score combines two capabilities: delivering useful patches when documentation is needed and correctly refraining from editing when it is not. The table reports the following measures:

Delivered patch quality, DD, averages rubric scores across the 82 held-out items that require an update. A missed or empty patch scores zero. Conditional quality averages only valid emitted patches.

Abstention recall, NN, measures correct abstention across the 35 held-out items that need no update.

We combine DD and NN with the harmonic mean, C=2DN/(D+N)C=2DN/(D+N). The harmonic mean treats useful patches and correct abstention as jointly necessary. It also keeps the score independent of the benchmark’s constructed class proportions.

Decision accuracy covers all 117 items.

P0-clean delivery is the share of the 82 documentation-needed items that received a valid P0-clean patch (Section 4).

Qwen3.8 Max+OpenCode has the highest composite score (47.3), followed by GPT-5.6 Sol+Codex (46.2) and GLM 5.2+OpenCode (44.2). GPT-5.6 Sol+Codex has the highest P0-clean delivery (39.0). Appendix [app:analysis-details] in the full PDF reports results on all 292 items separately.

Agent Score Accuracy Patch recall Abstention recall P0-clean delivery Delivered quality Conditional quality
Qwen3.8 Max +OpenCode 47.3 79.5 80.5 77.1 26.8 34.1 42.3
GPT-5.6 Sol +Codex 46.2 81.2 95.1 48.6 39.0 44.1 46.3
GLM 5.2 +OpenCode 44.2 83.8 84.1 82.9 23.2 30.1 35.8
Kimi K2.7 Code +OpenCode 43.9 82.1 91.5 60.0 24.4 34.6 37.8
Claude Opus 4.8 +Claude Code 41.2 81.2 79.3 85.7 24.4 27.1 34.2
Claude Sonnet 4.6 +Claude Code 40.0 82.9 93.9 57.1 20.7 30.8 32.8
GPT-5.5 +Codex 34.0 76.1 96.3 28.6 36.6 42.0 43.6
Table 1. Primary results on the 117-item held-out split (%). Rows are ordered by the unrounded composite score.

6.1 Patch or abstention decision performance

The held-out decision task contains 82 documentation-needed items (70.1%) and 35 abstention items (29.9%). Table 2 treats patch as the positive class.

Agent Accuracy Patch recall Abstention recall
GLM 5.2+OpenCode 83.8 (98/117) 84.1 (69/82) 82.9 (29/35)
Claude Sonnet 4.6+Claude Code 82.9 (97/117) 93.9 (77/82) 57.1 (20/35)
Kimi K2.7 Code+OpenCode 82.1 (96/117) 91.5 (75/82) 60.0 (21/35)
GPT-5.6 Sol+Codex 81.2 (95/117) 95.1 (78/82) 48.6 (17/35)
Claude Opus 4.8+Claude Code 81.2 (95/117) 79.3 (65/82) 85.7 (30/35)
Qwen3.8 Max+OpenCode 79.5 (93/117) 80.5 (66/82) 77.1 (27/35)
GPT-5.5+Codex 76.1 (89/117) 96.3 (79/82) 28.6 (10/35)
Table 2. Patch or abstention decisions for seven lanes on the 117-item held-out split (%). Parentheses give the numerator and denominator. Patch recall measures recovery of required updates, and abstention recall measures correct abstention.

Abstention recall ranges from 28.6% to 85.7%. GLM 5.2+OpenCode has the highest held-out decision accuracy (83.8%). GPT-5.5+Codex recovers 96.3% of required patches, and GPT-5.6 Sol+Codex recovers 95.1%. They make 25 and 18 incorrect decisions, respectively, on the 35 abstention items.

Unnecessary edits have real cost for both the documentation reader and the maintainers. These edits can add implementation details that users neither need nor can act on, which bloats the documentation and makes relevant guidance harder to find. Every unnecessary patch also needs maintainer attention during triage, review, and ongoing maintenance, and it adds work to downstream tasks such as translation and versioning. Abstaining from an unwarranted edit is therefore a documentation-quality and governance requirement, and we think the benchmark should measure it.

6.2 Documentation quality results

Table 3 reports quality conditional on a correct patch decision. GPT-5.6 Sol+Codex (46.3) and GPT-5.5+Codex (43.6) have the highest means.

Agent Mean 95% CI Median Correct patches (nn)
GPT-5.6 Sol+Codex 46.3 [40.8, 51.3] 50.0 78
GPT-5.5+Codex 43.6 [37.8, 48.7] 42.9 79
Qwen3.8 Max+OpenCode 42.3 [35.0, 48.7] 42.8 66
Kimi K2.7 Code+OpenCode 37.8 [31.9, 43.3] 35.7 75
GLM 5.2+OpenCode 35.8 [29.7, 41.8] 35.7 69
Claude Opus 4.8+Claude Code 34.2 [28.5, 40.1] 36.4 65
Claude Sonnet 4.6+Claude Code 32.8 [27.7, 38.1] 33.3 77
Table 3. Scored documentation quality on the 82 documentation-needed held-out items, conditional on a correct patch decision. Empty outputs, abstentions, and wrong decisions receive no quality score here. Their cost appears in delivered patch quality and the composite score. Intervals are percentile 95% intervals from 20,000 bootstrap resamples of documentation repositories.

6.3 Ablations

Source-event-date ablation We use source-event dates to probe training-time exposure: pull-request merge dates and issue creation dates. For each agent whose model has a provider-published knowledge cutoff, we divide all 205 documentation-needed items at the cutoff. All four agents score lower on post-cutoff items, but every repository-clustered interval includes zero, so the comparison provides no statistically conclusive evidence of a cutoff effect.

Agent Pre nn Pre score Post nn Post score Δ\Delta Repo 95% CI
Claude Opus 4.8+Claude Code 86 29.4 119 27.1 −2.2-2.2 [−9.5-9.5, 4.8]
Claude Sonnet 4.6+Claude Code 50 34.6 155 30.9 −3.7-3.7 [−12.8-12.8, 5.3]
GPT-5.5+Codex 69 42.9 136 40.9 −2.0-2.0 [−9.0-9.0, 7.1]
GPT-5.6 Sol+Codex 94 44.3 111 44.0 −0.3-0.3 [−6.6-6.6, 6.7]
Table 4. Mean documentation-quality score before and after each model’s published knowledge cutoff, on all 205 documentation-needed items. Event dates are pull-request merge dates or issue creation dates. Δ\Delta is the post-cutoff score minus the pre-cutoff score, and intervals resample repositories.

Repository-footprint ablation We also compare 82 items from low-footprint repositories (fewer than 1,000 GitHub stars) with 91 items from popular repositories (over 5,000 GitHub stars). The comparison shows no consistent advantage for popular repositories. Table [tab:popularity-ablation] in the full PDF in the appendix gives per-agent values.

Documentation-scale ablation We also tested whether larger documentation sets affect scores. Across repository-level means, doubling the page count changes the score by −0.98-0.98 points (95% CI −2.81-2.81 to +0.69+0.69). The interval includes zero, so the data show no clear relationship between documentation size and score.

7 Failure Analysis

7.1 Patch/abstention decision failures

Agents make decision failures by making edits when no change is needed (over-editing) and by failing to edit when change is needed (under-editing).

Over-editing often starts with a misleading cue. The cue is an internal symbol with a user-facing-sounding name, or an existing documentation page that mentions the affected component. The agent treats that cue as proof that the change alters documented user behavior. Typically, the agent documents one of the following, none of which needs documentation:

  • an internal refactor, rename, or mechanically regenerated type whose public contract is unchanged

  • a performance optimization, test, CI, or dependency change with no observable user effect

  • generated files or reference artifact that should not be manually updated

These patches often look defensible because the terms and destination are topically related to the diff. The missing step is establishing that the change creates behavior a user can invoke, observe, or act on. Once an agent predicts that it should edit, it tends to search for somewhere to put text. It does not go back to ask whether a documentation obligation exists, even when it later finds evidence for the opposite decision.

Under-editing happens when agents treat implementation location, the absence of existing coverage, or weak keyword overlap as evidence that no documentation change is needed. These proxies cause agents to miss user-facing changes. Common forms include the following:

  • treating a real public interface, such as a SQL keyword, API function, configuration option, or credential setting, as an implementation detail because it lives in a parser, registry, or configuration module

  • reading a new data source, integration, or landing-page capability as internal pipeline plumbing rather than a change to the reader’s available workflow

  • treating a failed search for existing coverage as evidence that a subject is intentionally undocumented, when the missing coverage may be the documentation gap the task exposes

  • declining documentation-only or issue-driven work because there is no implementation diff, even when the task identifies a verified gap in existing behavior

The absence-as-evidence case reinforces itself. In incorrect-abstention trajectories, an agent searches the existing documentation for the feature, interface, or workflow that the task affects. It finds no coverage and reads that silence as evidence that the subject is intentionally out of scope. From the agent’s perspective, deliberate exclusion and an unfilled gap produce the same search result. Once the agent treats absence as policy, the omission propagates.

7.2 Documentation quality failures

Technical-writing failure mode Effect on the documentation Rate
Task-completion gap Omits a decision point, procedure, verification step, or recovery path needed to complete the reader’s task. 45.5%
Technical inaccuracy Misstates an interface or behavior, including its scope, default, lifecycle, or compatibility. 36.6%
Incomplete conceptual or reference coverage Omits the core concept, capability, or contract the reader needs to understand and use the change. 32.5%
Missing supporting information Covers the main task but omits material rationale, boundaries, examples, operational detail, or secondary cases. 26.5%
Missing prerequisites Omits permissions, dependencies, credentials, versions, resources, enablement, or other setup conditions. 21.0%
Information-architecture or findability failure Places content outside the reader’s likely path or omits navigation, cross-references, and findable terminology. 20.7%
Missing audience and purpose framing Does not establish who the content is for, why it matters, or when to use it. 19.9%
Cross-surface content inconsistency Updates one surface while leaving an authoritative, mirrored, generated, or linked surface stale. 13.3%
Table 5. Share of submissions with each patch-level problem, across the 1,267 submissions audited before the reruns. The table lists labels assigned to at least 10% of submissions, and Table [tab:full-patch-failure-taxonomy] in the full PDF in the appendix lists the rest. One patch can have several problems. The labels describe defects in the patch itself. Table 6 covers the causes in the agents’ trajectories. Only 1.3% of submissions were labeled “no material defect” among the selected top labels.

Hallucinations often distort real behavior

Our audit labeled only 6.1% of submissions as containing fabricated content, a category that includes invented classes, flags, and endpoints. Technical inaccuracies appeared in 36.6% of submissions and involved scope, defaults, lifecycle, or compatibility. The failure taxonomy classifies invented interfaces or capabilities as fabricated content and false descriptions of real interfaces as technical inaccuracies. In practice, the boundary is not always clear.

One agent wrote that Helm 4 uses Server-Side Apply by default when installing or upgrading releases.1 That default applies to new installations. Releases created with Helm 3 continue using client-side apply after upgrading unless the user explicitly switches them. The agent explained this distinction later in the patch, but its opening statement still gave readers the wrong default. We classify this error as a technical inaccuracy because the feature exists. Extending the feature’s behavior beyond its supported conditions could also reasonably count as hallucination. In many audited cases, the agent took behavior that holds under narrow conditions and presented it as true in general. These claims are unfounded, like hallucinations, but they distort real functionality instead of inventing it.

Agent patches often lack a model of the reader’s task

Agents can often describe a feature’s behavior. They are less able to write for a reader who came to the documentation to decide something or reach a goal. Audience and purpose framing is missing in 19.9% of audited submissions. These patches explain what a feature does but not who should use it, why it is useful, or when to choose it. For example, one Strawberry GraphQL patch correctly documented how to select an older Apollo Federation version. It did not explain why a reader might need to: to upgrade Strawberry while staying compatible with an older Apollo Router or Gateway.2 The patch documented the setting but omitted the decision it was designed to support.

This limitation extends beyond explaining when or why to use a feature. Much technical documentation guides readers through a task, and the reader’s goal is to complete that task. Doing so may require prerequisites, intermediate decisions, procedural steps, verification, and recovery guidance. Task-completion gaps appear in 45.5% of audited submissions, and missing prerequisites in 21.0%. These failures suggest that agents treat a change as one piece of information to convey rather than as one part of a larger user journey. A patch may therefore describe the behavior accurately and still leave the reader unable to accomplish the task that brought them to the documentation.

This narrow view also affects how agents treat the documentation as a whole. Readers, both humans and agents, reach a page through search, navigation, related guides, and examples. Agents may add accurate information to a page that the intended reader is unlikely to visit, create a page without linking it from the relevant workflow, or update one surface and leave another surface on the same topic stale. Information-architecture or findability failures appear in 20.7% of audited submissions, and cross-surface inconsistencies in 13.3%. A patch can therefore be accurate in isolation and still fail within the larger documentation system.

7.3 Trajectory analysis of failure root causes

Patch-level labels describe what is wrong with the resulting documentation, but not why the agent produced it. We therefore inspected the trajectory behind each of the 1,267 audited submissions. For each material problem, we assigned one or more causes that the trace supports. Table 6 reports submission-level rates across this full population.

Trajectory-level root cause Observable reasoning failure Submissions Rate
Stopped at explaining the interface without examining how it is used in practice The agent described changed fields, settings, callbacks, or lifecycle mechanics, but did not test the explanation against the reader’s setup, decision, execution, verification, or recovery path. 456 36.0%
Did not inspect decisive evidence and filled the gap with a plausible assumption The agent found related material but stopped before the controlling implementation, schema, test, or public contract, then completed the explanation with a convention that sounded reasonable. 420 33.1%
Stopped searching after finding the first plausible documentation surface The agent found a reasonable page to edit and did not continue checking other maintained, generated, mirrored, migration, or workflow surfaces affected by the same change. 382 30.1%
Inspected relevant evidence but did not convert it into a complete coverage checklist The agent reached evidence bearing on the requirement but began drafting without tracking the claims, setup, boundaries, examples, and reader actions that needed to survive into the final patch. 349 27.5%
Committed too early to a narrow interpretation of the task Before completing the investigation, the agent declared the task to be a rename, reference update, single-page edit, or similarly narrow deliverable and ignored evidence outside that frame. 346 27.3%
Overgeneralized or misinterpreted partial evidence The agent inspected relevant evidence but converted one branch, example, implementation detail, or deployment pattern into a broader or different public rule. 333 26.3%
Table 6. The six most common trajectory-level root causes across the 1,267 pre-rerun trajectories, which were frozen separately. Multiple causes may apply, so rates do not sum to 100%. Table [tab:trajectory-root-causes-remaining] in the full PDF in the appendix lists the less common causes.

Agents often lack a reliable test for whether they have gathered enough evidence. Sometimes agents stop researching before they reach the decisive evidence. Other times they begin drafting from partial information without recognizing that their evidence is incomplete. This pattern suggests a failure to recognize uncertainty. Prior work reports similar findings (Liu et al., 2025; Shao et al., 2025; Liu et al., 2025). Models struggle to identify the source of uncertainty. Information-seeking agents often answer before the available evidence is sufficient, and they do not reliably recognize when more information gathering has value.

Premature closure, shifting from investigation to drafting too soon, cuts across many of the root causes. Once an agent finds a plausible interpretation or a reasonable page to edit, it often starts drafting. It may stop before reaching key evidence, and it may also stop before checking every affected documentation surface, which leaves parts of the documentation stale. Insufficient search is only part of the problem. The larger part is that agents lack a reliable stopping rule. Such a rule would tell an agent when it understands the task, the evidence, and the documentation impact well enough to begin writing.

We found no clear relationship between the assigned root cause and trajectory length, whether measured by turns or by token use. Effort also did not rise with task difficulty. We defined an item’s difficulty from the mergeability of the other six agents’ patches on that item. Within each agent, the rank correlation between difficulty and effort was +0.034+0.034 for processed tokens, +0.034+0.034 for trace-event count, and +0.030+0.030 for tool actions. All task-clustered 95% confidence intervals include zero. Premature closure therefore does not necessarily produce a short trajectory. An agent may stop investigating early and then spend substantial effort drafting, revising, or elaborating an incomplete account.

8 Discussion

8.1 Missing context about the reader

Some failures that we attribute to a missing model of the reader may instead reflect missing context about how the software is used. Agents cannot always infer from parametric knowledge alone what readers are trying to accomplish or which details they need. That inference is especially hard when user motivations and the surrounding workflow context are implicit rather than stated.

8.2 Coarse training rewards

Premature-closure failures may be related to coarse reward signals during post-training. RAGEN finds that trajectory-level rewards do not reliably teach agents how to reason through multi-turn tasks. Without fine-grained, reasoning-aware feedback, agents may learn shallow strategies or produce reasoning that is not grounded in the environment (Wang et al., 2025). Kim et al. report a similar pattern (Kim et al., 2026). In their experiments, outcome-only reinforcement learning improved final accuracy but made intermediate reasoning less accurate and less internally consistent. Models learned shortcuts rather than reliable reasoning procedures. These results offer possible explanations for our findings. The agents seemed to infer scope from early cues, such as the location of a code change or the name of a feature. They then began drafting within that narrow frame and filled evidence gaps with plausible assumptions that their environment did not support.

8.3 Knowing when to stop investigating

Another explanation is that deciding when the evidence is sufficient is itself a difficult capability. SeekBench reports that search agents trained with reinforcement learning answered before gathering sufficient evidence in 76.5% of the evaluated trajectories (Shao et al., 2025). CaRT shows that models may rely on superficial stopping rules, such as the number of turns, instead of checking for a decisive fact (Liu et al., 2025). Related studies report that language models struggle to retract an earlier inference when new evidence contradicts it and tend to seek examples that confirm an initial hypothesis rather than examples that might disprove it (Wilie et al., 2024; Jhaveri et al., 2026). These findings match the patterns in our trajectory analysis.

8.4 Additional guidance and scaffolding

Several changes could plausibly address the observed failures: explicit instructions in the prompts that ask agents to consider the reader’s goal, skills that emphasize task completion and findability, broader tools, scratch notes, and explicit verification. We explored these approaches informally but did not systematically compare them against a baseline, so we cannot conclude whether they improved documentation quality. Future work should test their effects through controlled comparisons.

9 Limitations

9.1 Measurement

Task-specific rubrics and LLM judges (Section 4) let us score open-ended documentation patches at scale. Because human validation covers only part of the evaluation, automated rubric generation and scoring may still introduce errors. We check commands and examples against the available evidence instead of running every documented procedure in its repository’s native build and runtime environment. As a result, the benchmark has no deterministic checks for code samples and links.

We find that generated rubrics tend to contain more criteria than maintainer-authored rubrics. Many of these additional criteria identify valid documentation improvements, but maintainers may consider them less important. During human calibration, reviewers prioritized the noncritical P1–P3 criteria differently. Depending on repository norms, some emphasized style, while others placed less weight on it. Some preferred comprehensive documentation, whereas others favored a simple, easy-to-follow user path over broader coverage. These preferences do not support a single universal weighting of P1–P3 criteria. The score therefore weights all criteria equally in the mean and handles P0 failures separately through the score cap. The published dataset retains the P0–P3 labels, and the accompanying scoring code lets practitioners apply other weights.

Manual review found that some rubric criteria overlap and are not fully independent. Some overlap is warranted, because a single documentation failure can cause several related problems. In a later audit, we tried to merge overlapping criteria. Merging sometimes lost important distinctions, so we kept the overlapping criteria.

9.2 No internet access

We evaluated agents without internet access (Sections 3.2 and 5). To assess how this restriction affected performance, we reviewed 683 trajectories from items on which no agent produced a mergeable patch. We found blocked network requests in 73 runs (10.7%). In 63 of these runs, network access was not necessary to produce a correct patch. The requests mostly involved setup or validation, such as installing dependencies, building documentation, running formatters or linters, and parsing YAML or JSON configuration files. Only 10 runs (1.5% of all audited trajectories) tried to retrieve external evidence that a correct patch required and the supplied inputs lacked.

Internet access might therefore have helped on a small number of items. However, in an earlier web-enabled pilot, 26 of 85 runs (30.6%) retrieved the item’s upstream pull request and its merged documentation. Given this contamination risk, we kept the reported evaluation sealed.

9.3 Asymmetric label construction

The evidence for the abstention and patch labels (Section 3.1) is not equally strong. Documentation-needed items often have direct evidence. Some abstention labels, by contrast, rely only on the absence of a related documentation change within a 90-day audit window. That absence does not necessarily show that documentation was unnecessary. An update may have been forgotten, or it may have happened after 90 days without a link to the code pull request. As a result, some items that needed documentation may be mislabeled as abstention items. This asymmetric label noise could distort the measured decision performance.

9.4 Single-run evaluation

For budget and time reasons, we run each agent on each item once (Section 5). We therefore do not report pass@k, pass^k, best-of-kk performance, or within-item run-to-run variance. The results characterize one sampled trajectory per agent and item. They do not measure the probability that an agent reliably produces the same decision or documentation quality across repeated attempts.

9.5 Low-information prose may be under-penalized

The rubric-based scoring (Section 4) and the patch-level failure rates (Section 7.2) may miss low-information prose. A general instruction to identify coherent but low-information prose flagged 3.2% of submissions. A second prompt asked the judge to flag submissions with two or more specific patterns. The patterns included redundant paraphrases, unnecessary explanations, excessive bulleted lists, formulaic contrasts (not X, but Y), three-part constructions, and heavy use of em dashes. This prompt flagged 7.6% of submissions. The increase suggests that LLM judges may miss low-information prose during scoring.

Disclosure

Two authors are affiliated with Promptless, a company that builds documentation agents.

10 Conclusion

DoGBench shows that even frontier models paired with frontier coding harnesses cannot yet reliably produce expert-level user-facing documentation in one attempt. The highest composite score on the 117-item held-out split is 47.3 out of 100. Agents still misjudge whether documentation is needed, and when they do edit, their patches may be factually correct but miss what readers need. Fluent prose and capable repository tooling do not yet close this gap.

DoGBench measures this gap with 292 real items, drawn from software changes and reported documentation gaps, and it shows where decisions and patches fail. We release the evaluation harness, item schema, dataset card, and development examples so that others can build on the benchmark. Benchmark materials and release information are available at https://dogbench.ai.

References

  1. Choudhury, S. Process reward models for LLM agents: Practical framework and directions. arXiv preprint arXiv:2502.10325, 2025.
  2. Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. ICML, 2023.
  3. Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436, 2019.
  4. Jimenez, C. E., Yang, J., Wettig, A., et al. SWE-bench: Can language models resolve real-world GitHub issues? ICLR, 2024.
  5. Jhaveri, A. R., GX-Chen, A., Sucholutsky, I., and Choi, E. Failing to falsify: Evaluating and mitigating confirmation bias in language models. arXiv preprint arXiv:2604.02485, 2026.
  6. Kim, K., Wang, K., Xie, Y., et al. Correct answers from sound reasoning: Verifiable process supervision for language models. COLM, 2026.
  7. Lightman, H., Kosaraju, V., Burda, Y., et al. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  8. Liu, G., Qu, Y., Schneider, J., Singh, A., and Kumar, A. CaRT: Teaching LLM agents to know when they know enough. arXiv preprint arXiv:2510.08517, 2025.
  9. OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/, 2026.
  10. Lu, S., Guo, D., Ren, S., et al. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. NeurIPS Datasets and Benchmarks, 2021.
  11. Liu, J., Peng, J., Wu, X., et al. Do not abstain! Identify and solve the uncertainty. ACL, pages 17177–17197, 2025.
  12. Nguyen Hoang, A., Le-Anh, M., Le, B., and Bui, N. D. Q. CodeWiki: Evaluating AI’s ability to generate holistic documentation for large-scale codebases. Findings of ACL, pages 5812–5827, 2026.
  13. Mei, W., Gu, Z., Bai, Z., et al. Deep Research as Rubric for Reinforcement Learning. arXiv preprint arXiv:2606.01091, 2026.
  14. Panickssery, A., Bowman, S. R., and Feng, S. LLM evaluators recognize and favor their own generations. NeurIPS, 2024.
  15. Pai, K., Devanbu, P., and Ahmed, T. CoDocBench: A dataset for code-documentation alignment in software maintenance. arXiv preprint arXiv:2502.00519, 2025.
  16. Sainz, O., Campos, J. A., García-Ferrero, I., et al. NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. EMNLP Findings, 2023.
  17. Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., et al. Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. UIST, 2024.
  18. Shao, J., Lin, Y., Lohani, M. P., Miao, Y., and Luo, B. Do LLM agents know how to ground, recover, and assess? A benchmark for epistemic competence in information-seeking agents. arXiv preprint arXiv:2509.22391, 2025.
  19. Wang, X., Hu, R., Gao, C., Gao, P., and Peng, C. Evaluating repository-level software documentation via question answering and feature-driven development. arXiv preprint arXiv:2604.06793, 2026.
  20. Wang, Z., Wang, K., Wang, Q., et al. RAGEN: Understanding self-evolution in LLM agents via multi-turn reinforcement learning. arXiv preprint arXiv:2504.20073, 2025.
  21. Wilie, B., Cahyawijaya, S., Ishii, E., He, J., and Fung, P. Belief revision: The adaptability of large language models reasoning. EMNLP, pages 10480–10496, 2024.
  22. Ye, Z., Shi, W., Liu, Y., et al. Look before you leap: Autonomous exploration for LLM agents. arXiv preprint arXiv:2605.16143, 2026.
  23. Zhou, H., Huang, H., Long, Y., et al. Mitigating the bias of large language model evaluation. Proceedings of the 23rd Chinese National Conference on Computational Linguistics, pages 1310–1319, 2024.
  24. Zhou, J., Zhang, Q., Wang, Y., et al. RubricBench: Aligning model-generated rubrics with human standards. ACL, pages 31179–31200, 2026.

  1. Item helm-helm-www-pr1926, Claude Sonnet 4.6+Claude Code.↩︎

  2. Item strawberry-graphql-strawberry-pr4045, GPT-5.5+Codex.↩︎