1 Introduction
User-facing documentation is the main public description of what a software product is, what capabilities it offers, and when and how to use them. For closed-source products in particular, documentation may be the only structured source an AI agent can use to understand and operate the product. People increasingly rely on AI agents to find, choose, and use products. Documentation therefore affects whether an agent uses a product correctly and whether the agent considers the product for the userâs task at all.
Producing this content requires more than translating implementation details into prose. The work demands judgment about where information belongs, what readers are trying to accomplish, and what they already know. Teams increasingly use agents to write documentation, but no one has yet systematically evaluated the user-facing documentation that agents produce.
Prior work evaluates code-facing documentation rather than user-facing documentation (Section 2.1). That work covers function-level docstrings, repository-level code summaries, and internal developer documentation, and the work assumes that a human already decided that the documentation needs an update. No existing documentation benchmark tests whether an agent can tell when to leave the documentation alone.
We introduce DoGBench, a benchmark of 292 items from open source projects. Each item gives an agent a pre-change repository and a trigger, which is either a pull request or a reported documentation gap. The agent first decides whether the trigger calls for a documentation change. For 205 items, the correct response is a patch. For the other 87, the correct response is to abstain. Task-specific rubrics, validated with project maintainers, score each patch. The rubrics do not reward similarity to the documentation that humans merged. Depending on the task, a correct patch may revise existing guidance, create and register a new page, move content, or remove stale or redundant content. Every item starts from an existing product and its then-current documentation. The benchmark therefore does not cover writing a productâs documentation from scratch or redesigning an information architecture without constraints.
We use a random, stratified 117-item held-out split for primary evaluation. The other 175 items form a public development split. We use the public split and the full 292 items only for robustness analyses.
We evaluated seven agent lanes. The highest-scoring agent reached 47.3 out of 100 on the held-out split. Only 6.1% of submissions contained fabricated content. The more common failure was a plausible patch that left the reader unable to finish the task. Of 1,267 submissions, 45.5% had a task-completion gap. Section 7 traces these failures to how agents investigate. In 36.0% of submissions, agents explained product interfaces without checking how readers use them. In 33.1%, agents stopped before finding decisive evidence. In 30.1%, agents edited the first plausible documentation surface and missed other surfaces the change affected.
2 Related Work
2.1 Documentation generation benchmarks
Existing documentation benchmarks focus on code-facing documentation. CodeSearchNet supplied a corpus of paired functions and documentation for semantic code search, and CodeXGLUE used CodeSearchNet-derived data for code summarization (Husain et al., 2019; Lu et al., 2021). More recent benchmarks study docstring updates after code changes (CoDocBench) and repository-level internal documentation (CodeWikiBench) (Pai et al., 2025; Nguyen Hoang et al., 2026). Like DoGBench, SWD-Bench builds its tasks from pull requests, but it scores repository-level documentation by how well a model can use that documentation to answer questions about the repositoryâs functionality (Wang et al., 2026). None of these benchmarks asks the model to decide whether documentation needs an update.
2.2 Scoring open-ended edits
SWE-bench, which also draws from open source repositories, is widely used to evaluate code generation (Jimenez et al., 2024). OpenAI has since questioned the validity of its Verified subset because narrow tests can reject correct alternative solutions (OpenAI, 2026a). That risk is larger for documentation because many different edits can satisfy the same reader need. DoGBench therefore scores each patch against requirements drawn from the triggering change and the pre-change repository rather than against similarity to the merged human patch.
3 Benchmark Design
Each item in the benchmark gives the agent a pre-change repository and a trigger. The trigger is either a code pull request or a user-reported documentation gap, often a GitHub issue. This setup mirrors how maintainers work. A maintainer either ships documentation changes with a feature or updates the documentation in response to a community issue. The agent must either return a patch that edits the documentation or abstain from making documentation changes when no user-facing change is needed. Figure 1 summarizes the item-construction and evaluation pipeline.
3.1 Dataset construction
The benchmark contains 292 items: 205 require a documentation change, and 87 require abstention. We built them from three pools. The first holds 90 changes initially sampled as likely abstention cases. The second holds 136 code-triggered documentation updates, and the third holds 66 explicit user-reported documentation gaps. During final adjudication, we reclassified three of the 90 likely abstention cases as requiring documentation, which left 87 abstention items. The final set therefore has 139 code-triggered documentation items, 66 items triggered by reported documentation gaps, and 87 abstention items. Project maintainers reviewed 42 items from repositories such as Helm, Doc Detective, Mautic, and PostHog. Appendix [app:dataset-details] in the full PDF gives additional selection details and two case studies.
Constructing the trigger
For a pull request that ships code and documentation together, we remove the documentation changes and give the agent only the code change. For a pull request that changes only documentation, we build the trigger from linked issues and discussions. We add sanitized versions of the source pull requestâs title and description. Both methods keep the documentation need and the reason for it but hide how the maintainer wrote the documentation.
We selected open source repositories with English-language documentation across diverse ecosystems. We exclude the following:
reverts, release-only version bumps, and merge or sync pull requests
pure refactors
changes where the connection between trigger and the documentation need is unclear
items that need context unavailable in the public repository or trigger
We also exclude bot-authored pull requests are excluded unless a human maintainer reviewed and revised the change before merging. Appendix [app:dataset-details] in the full PDF gives the full rules and known limitations.
3.2 Contamination controls
Because the source events are public, contamination can happen if a model has seen the merged documentation during training or if an agent finds it while running. We probe training-time exposure with source-event-date and repository-footprint ablations (Section 6.3) (Sainz et al., 2023). To prevent execution-time exposure, agents work in fresh Docker containers with identity-masked repositories and no network access except to the model provider. We also inspect agent trajectories for attempts to reach the merged human patch. A conservative overlap detector flags suspicious similarity to the merged documentation for manual review. We found no confirmed case of copying. We reviewed every retained detector alert and judged each one a false positive.
4 Evaluation Protocol
For each item that requires a documentation update, we score the submitted patch against a task-specific rubric. Across the 205 documentation-needed items, the rubrics contain 3,273 criteria, including 798 P0 criteria.
We design each criterion to test one observable review decision. For example, suppose a change lets a Helm values file to contain multiple YAML documents. Separate criteria can then check the documentation for three statements: documents are processed in order, later values take precedence, and nested maps are merged recursively. A broad criterion such as âexplains multi-document values wellâ would not be testable enough.
4.1 Criterion types and priorities
Each criterion use one of three scoring types:
A requirement always applies and receives a binary Pass or Fail verdict.
A conditional criterion applies only when the patch meets its condition. We call a conditional criterion triggered when it applies.
A deduction-only guardrail prohibits content such as a fabricated command or an unsafe recovery step. Avoiding the prohibited content earns no credit, and introducing it costs a deduction.
Requirements and triggered conditional criteria receive only Pass or Fail verdicts, with no partial credit. Untriggered conditional criteria are left out of scoring.
Each criterion also has a priority from P0 to P3, which is separate from its scoring type. P0 is reserved for a defect that blocks the patch on its own: failing that criterion alone would require revision under the benchmarkâs standard. Typical P0 defects include the following:
materially misstating product behavior
omitting information the reader needs to complete the central reader task
giving an unsafe or destructive instruction
inventing a public interface
leaving an essential maintained documentation surface contradictory or unusable
Optional examples, secondary edge cases, stylistic preferences, and exact wording do not qualify as P0 merely because they would improve the patch. A patch is P0-clean when no applicable P0 criterion failed. P0-clean diagnoses only critical defects and does not a prediction whether a maintainer would merge the patch unchanged.
4.2 Score construction
Let contain all requirements and triggered conditional criteria, and let contain all violated deduction-only guardrails. Each criterion in contributes one point if it passes, and each violated guardrail deducts one point. For , the uncapped patch score is An untriggered conditional criterion counts in neither the numerator nor the denominator. A guardrail that the patch does not violate is also left out, so a patch earns no points merely for avoiding an optional risk. If is empty, the score is zero.
Let when a P0 requirement or triggered conditional criterion fails, or when a P0 guardrail is violated. Otherwise, let . The reported patch score is On a documentation-needed item, an empty patch or an abstention also scores zero. We set the 60-point ceiling as an evaluation policy. We did not estimate it from maintainer editing time or acceptance decisions. The ceiling prevents success on many secondary criteria from averaging away a critical defect.
4.3 Rubric construction
Research and synthesis
We build each rubric through repository research and an LLM-council process. The merged human patch is not included among the candidate patches supplied to the rubric agents. A rubric research agent inspects the triggering change and the repository to identify what users need to know and the evidence that supports each requirement. The research agent has internet access and may encounter the merged human patch, but every criterion must have independent supporting evidence; the human patch alone cannot justify a criterion. A synthesis stage turns these findings into criteria with explicit passing and failing conditions. (Mei et al., 2026) also explore this research-to-criteria approach.
Differential review
The differential stage compares anonymous candidate patches side by side to find editorial choices that the draft rubric does not yet cover. This stage follows the observation that inspecting model outputs can help refine evaluation criteria (Shankar et al., 2024). For each uncovered difference, the rubric agent checks the research findings and gathers more evidence where needed. It then decides whether the difference matters to the reader. A difference between candidates only raises a question and does not establish what is correct. Trivial or neutral differences do not become criteria. The rubric agent also records supported documentation needs that no candidate meets. These findings are then used to revise the draft rubric.
Audit and debate
A second model, from a different model family, audits the revised criteria and their priorities. When the rubric author and the auditor disagree, they revisit the evidence and exchange arguments. Together they revise or remove any criterion that cannot be justified. The debate ends when the auditor accepts the revised rubric or the exchange reaches its configured limit. The longest saved debate we inspected ran 28 messages after the opening audit.
4.4 Human validation
We compare the resulting rubrics with independently collected maintainer criteria on a reviewed subset. Separately, we compare the Pass or Fail verdicts of a scoring model with human judgments. The first check asks whether the rubric captures the requirements that maintainers consider important. The second asks whether the scoring model applies those requirements correctly.
On 42 maintainer-reviewed items, the automatic documentation-need gate agrees with the maintainers on 41, with one false positive. On 21 documentation-needed items with completed criterion alignment, mean priority-weighted recall against maintainer criteria is 0.892. In a separate study of 20 items and 330 criteria, the scoring model agrees with the post-adjudication human reference on 310 criteria (93.9%). One paper author resolved disagreements in the human reference after review, so this figure is not blinded agreement between two humans. Appendix [app:evaluation-details] in the full PDF describes both validation studies.
5 Experimental Setup
We evaluate seven agent lanes:
GLM 5.2, Qwen3.8 Max, and Kimi K2.7 Code with OpenCode
Claude Opus 4.8 and Claude Sonnet 4.6 with Claude Code
GPT-5.5 and GPT-5.6 Sol with Codex
Every agent runs each item in a fresh Docker environment with the same inputs. The inputs are the pre-change code and documentation repository, plus the trigger: a code diff or a reported documentation gap. Each agentâitem pair runs once and must either return a patch or abstain. Agents can use command-line tools to inspect and edit the repository, but the environment has no network access except to the model providers.
Why the environment is sealed
Internet access can be valuable for documentation agents in ordinary use, but DoGBench blocks it to prevent contamination. Because the source events and merged documentation are public, a connected agent could retrieve the answer instead of solving the task from the supplied evidence. In an audit of an earlier version of the benchmark, we found that 31% of runs retrieved the source pull request despite identity masking. A further 6% copied directly from other agentsâ earlier trajectories because the agents shared a host.
Data splits
The release has two splits. The development split holds 175 items: 123 documentation-needed and 52 abstention. The held-out split holds 117 items: 82 documentation-needed and 35 abstention. The random split draw comes from a procedure stratified by class, source pool, task type, and repository. The split was drawn without access to measured scores. Development items include inputs, rubrics, and reference artifacts. Held-out rubrics, references, and item-level scores remain private. We report only the seven reproducible agents in this paper.
6 Results
Table 1 reports the composite score on the 117-item held-out split. The composite score combines two capabilities: delivering useful patches when documentation is needed and correctly refraining from editing when it is not. The table reports the following measures:
Delivered patch quality, , averages rubric scores across the 82 held-out items that require an update. A missed or empty patch scores zero. Conditional quality averages only valid emitted patches.
Abstention recall, , measures correct abstention across the 35 held-out items that need no update.
We combine and with the harmonic mean, . The harmonic mean treats useful patches and correct abstention as jointly necessary. It also keeps the score independent of the benchmarkâs constructed class proportions.
Decision accuracy covers all 117 items.
P0-clean delivery is the share of the 82 documentation-needed items that received a valid P0-clean patch (Section 4).
Qwen3.8 Max+OpenCode has the highest composite score (47.3), followed by GPT-5.6 Sol+Codex (46.2) and GLM 5.2+OpenCode (44.2). GPT-5.6 Sol+Codex has the highest P0-clean delivery (39.0). Appendix [app:analysis-details] in the full PDF reports results on all 292 items separately.
| Agent | Score | Accuracy | Patch recall | Abstention recall | P0-clean delivery | Delivered quality | Conditional quality |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Max +OpenCode | 47.3 | 79.5 | 80.5 | 77.1 | 26.8 | 34.1 | 42.3 |
| GPT-5.6 Sol +Codex | 46.2 | 81.2 | 95.1 | 48.6 | 39.0 | 44.1 | 46.3 |
| GLM 5.2 +OpenCode | 44.2 | 83.8 | 84.1 | 82.9 | 23.2 | 30.1 | 35.8 |
| Kimi K2.7 Code +OpenCode | 43.9 | 82.1 | 91.5 | 60.0 | 24.4 | 34.6 | 37.8 |
| Claude Opus 4.8 +Claude Code | 41.2 | 81.2 | 79.3 | 85.7 | 24.4 | 27.1 | 34.2 |
| Claude Sonnet 4.6 +Claude Code | 40.0 | 82.9 | 93.9 | 57.1 | 20.7 | 30.8 | 32.8 |
| GPT-5.5 +Codex | 34.0 | 76.1 | 96.3 | 28.6 | 36.6 | 42.0 | 43.6 |
6.1 Patch or abstention decision performance
The held-out decision task contains 82 documentation-needed items
(70.1%) and 35 abstention items (29.9%). Table 2 treats
patch as the positive class.
| Agent | Accuracy | Patch recall | Abstention recall |
|---|---|---|---|
| GLM 5.2+OpenCode | 83.8 (98/117) | 84.1 (69/82) | 82.9 (29/35) |
| Claude Sonnet 4.6+Claude Code | 82.9 (97/117) | 93.9 (77/82) | 57.1 (20/35) |
| Kimi K2.7 Code+OpenCode | 82.1 (96/117) | 91.5 (75/82) | 60.0 (21/35) |
| GPT-5.6 Sol+Codex | 81.2 (95/117) | 95.1 (78/82) | 48.6 (17/35) |
| Claude Opus 4.8+Claude Code | 81.2 (95/117) | 79.3 (65/82) | 85.7 (30/35) |
| Qwen3.8 Max+OpenCode | 79.5 (93/117) | 80.5 (66/82) | 77.1 (27/35) |
| GPT-5.5+Codex | 76.1 (89/117) | 96.3 (79/82) | 28.6 (10/35) |
Abstention recall ranges from 28.6% to 85.7%. GLM 5.2+OpenCode has the highest held-out decision accuracy (83.8%). GPT-5.5+Codex recovers 96.3% of required patches, and GPT-5.6 Sol+Codex recovers 95.1%. They make 25 and 18 incorrect decisions, respectively, on the 35 abstention items.
Unnecessary edits have real cost for both the documentation reader and the maintainers. These edits can add implementation details that users neither need nor can act on, which bloats the documentation and makes relevant guidance harder to find. Every unnecessary patch also needs maintainer attention during triage, review, and ongoing maintenance, and it adds work to downstream tasks such as translation and versioning. Abstaining from an unwarranted edit is therefore a documentation-quality and governance requirement, and we think the benchmark should measure it.
6.2 Documentation quality results
Table 3 reports quality conditional on a correct patch decision. GPT-5.6 Sol+Codex (46.3) and GPT-5.5+Codex (43.6) have the highest means.
| Agent | Mean | 95% CI | Median | Correct patches () |
|---|---|---|---|---|
| GPT-5.6 Sol+Codex | 46.3 | [40.8, 51.3] | 50.0 | 78 |
| GPT-5.5+Codex | 43.6 | [37.8, 48.7] | 42.9 | 79 |
| Qwen3.8 Max+OpenCode | 42.3 | [35.0, 48.7] | 42.8 | 66 |
| Kimi K2.7 Code+OpenCode | 37.8 | [31.9, 43.3] | 35.7 | 75 |
| GLM 5.2+OpenCode | 35.8 | [29.7, 41.8] | 35.7 | 69 |
| Claude Opus 4.8+Claude Code | 34.2 | [28.5, 40.1] | 36.4 | 65 |
| Claude Sonnet 4.6+Claude Code | 32.8 | [27.7, 38.1] | 33.3 | 77 |
6.3 Ablations
Source-event-date ablation We use source-event dates to probe training-time exposure: pull-request merge dates and issue creation dates. For each agent whose model has a provider-published knowledge cutoff, we divide all 205 documentation-needed items at the cutoff. All four agents score lower on post-cutoff items, but every repository-clustered interval includes zero, so the comparison provides no statistically conclusive evidence of a cutoff effect.
| Agent | Pre | Pre score | Post | Post score | Repo 95% CI | |
|---|---|---|---|---|---|---|
| Claude Opus 4.8+Claude Code | 86 | 29.4 | 119 | 27.1 | [, 4.8] | |
| Claude Sonnet 4.6+Claude Code | 50 | 34.6 | 155 | 30.9 | [, 5.3] | |
| GPT-5.5+Codex | 69 | 42.9 | 136 | 40.9 | [, 7.1] | |
| GPT-5.6 Sol+Codex | 94 | 44.3 | 111 | 44.0 | [, 6.7] |
Repository-footprint ablation We also compare 82 items from low-footprint repositories (fewer than 1,000 GitHub stars) with 91 items from popular repositories (over 5,000 GitHub stars). The comparison shows no consistent advantage for popular repositories. Table [tab:popularity-ablation] in the full PDF in the appendix gives per-agent values.
Documentation-scale ablation We also tested whether larger documentation sets affect scores. Across repository-level means, doubling the page count changes the score by points (95% CI to ). The interval includes zero, so the data show no clear relationship between documentation size and score.
7 Failure Analysis
7.1 Patch/abstention decision failures
Agents make decision failures by making edits when no change is needed (over-editing) and by failing to edit when change is needed (under-editing).
Over-editing often starts with a misleading cue. The cue is an internal symbol with a user-facing-sounding name, or an existing documentation page that mentions the affected component. The agent treats that cue as proof that the change alters documented user behavior. Typically, the agent documents one of the following, none of which needs documentation:
an internal refactor, rename, or mechanically regenerated type whose public contract is unchanged
a performance optimization, test, CI, or dependency change with no observable user effect
generated files or reference artifact that should not be manually updated
These patches often look defensible because the terms and destination are topically related to the diff. The missing step is establishing that the change creates behavior a user can invoke, observe, or act on. Once an agent predicts that it should edit, it tends to search for somewhere to put text. It does not go back to ask whether a documentation obligation exists, even when it later finds evidence for the opposite decision.
Under-editing happens when agents treat implementation location, the absence of existing coverage, or weak keyword overlap as evidence that no documentation change is needed. These proxies cause agents to miss user-facing changes. Common forms include the following:
treating a real public interface, such as a SQL keyword, API function, configuration option, or credential setting, as an implementation detail because it lives in a parser, registry, or configuration module
reading a new data source, integration, or landing-page capability as internal pipeline plumbing rather than a change to the readerâs available workflow
treating a failed search for existing coverage as evidence that a subject is intentionally undocumented, when the missing coverage may be the documentation gap the task exposes
declining documentation-only or issue-driven work because there is no implementation diff, even when the task identifies a verified gap in existing behavior
The absence-as-evidence case reinforces itself. In incorrect-abstention trajectories, an agent searches the existing documentation for the feature, interface, or workflow that the task affects. It finds no coverage and reads that silence as evidence that the subject is intentionally out of scope. From the agentâs perspective, deliberate exclusion and an unfilled gap produce the same search result. Once the agent treats absence as policy, the omission propagates.
7.2 Documentation quality failures
| Technical-writing failure mode | Effect on the documentation | Rate |
|---|---|---|
| Task-completion gap | Omits a decision point, procedure, verification step, or recovery path needed to complete the readerâs task. | 45.5% |
| Technical inaccuracy | Misstates an interface or behavior, including its scope, default, lifecycle, or compatibility. | 36.6% |
| Incomplete conceptual or reference coverage | Omits the core concept, capability, or contract the reader needs to understand and use the change. | 32.5% |
| Missing supporting information | Covers the main task but omits material rationale, boundaries, examples, operational detail, or secondary cases. | 26.5% |
| Missing prerequisites | Omits permissions, dependencies, credentials, versions, resources, enablement, or other setup conditions. | 21.0% |
| Information-architecture or findability failure | Places content outside the readerâs likely path or omits navigation, cross-references, and findable terminology. | 20.7% |
| Missing audience and purpose framing | Does not establish who the content is for, why it matters, or when to use it. | 19.9% |
| Cross-surface content inconsistency | Updates one surface while leaving an authoritative, mirrored, generated, or linked surface stale. | 13.3% |
Hallucinations often distort real behavior
Our audit labeled only 6.1% of submissions as containing fabricated content, a category that includes invented classes, flags, and endpoints. Technical inaccuracies appeared in 36.6% of submissions and involved scope, defaults, lifecycle, or compatibility. The failure taxonomy classifies invented interfaces or capabilities as fabricated content and false descriptions of real interfaces as technical inaccuracies. In practice, the boundary is not always clear.
One agent wrote that Helm 4 uses Server-Side Apply by default when installing or upgrading releases.1 That default applies to new installations. Releases created with Helm 3 continue using client-side apply after upgrading unless the user explicitly switches them. The agent explained this distinction later in the patch, but its opening statement still gave readers the wrong default. We classify this error as a technical inaccuracy because the feature exists. Extending the featureâs behavior beyond its supported conditions could also reasonably count as hallucination. In many audited cases, the agent took behavior that holds under narrow conditions and presented it as true in general. These claims are unfounded, like hallucinations, but they distort real functionality instead of inventing it.
Agent patches often lack a model of the readerâs task
Agents can often describe a featureâs behavior. They are less able to write for a reader who came to the documentation to decide something or reach a goal. Audience and purpose framing is missing in 19.9% of audited submissions. These patches explain what a feature does but not who should use it, why it is useful, or when to choose it. For example, one Strawberry GraphQL patch correctly documented how to select an older Apollo Federation version. It did not explain why a reader might need to: to upgrade Strawberry while staying compatible with an older Apollo Router or Gateway.2 The patch documented the setting but omitted the decision it was designed to support.
This limitation extends beyond explaining when or why to use a feature. Much technical documentation guides readers through a task, and the readerâs goal is to complete that task. Doing so may require prerequisites, intermediate decisions, procedural steps, verification, and recovery guidance. Task-completion gaps appear in 45.5% of audited submissions, and missing prerequisites in 21.0%. These failures suggest that agents treat a change as one piece of information to convey rather than as one part of a larger user journey. A patch may therefore describe the behavior accurately and still leave the reader unable to accomplish the task that brought them to the documentation.
This narrow view also affects how agents treat the documentation as a whole. Readers, both humans and agents, reach a page through search, navigation, related guides, and examples. Agents may add accurate information to a page that the intended reader is unlikely to visit, create a page without linking it from the relevant workflow, or update one surface and leave another surface on the same topic stale. Information-architecture or findability failures appear in 20.7% of audited submissions, and cross-surface inconsistencies in 13.3%. A patch can therefore be accurate in isolation and still fail within the larger documentation system.
7.3 Trajectory analysis of failure root causes
Patch-level labels describe what is wrong with the resulting documentation, but not why the agent produced it. We therefore inspected the trajectory behind each of the 1,267 audited submissions. For each material problem, we assigned one or more causes that the trace supports. Table 6 reports submission-level rates across this full population.
| Trajectory-level root cause | Observable reasoning failure | Submissions | Rate |
|---|---|---|---|
| Stopped at explaining the interface without examining how it is used in practice | The agent described changed fields, settings, callbacks, or lifecycle mechanics, but did not test the explanation against the readerâs setup, decision, execution, verification, or recovery path. | 456 | 36.0% |
| Did not inspect decisive evidence and filled the gap with a plausible assumption | The agent found related material but stopped before the controlling implementation, schema, test, or public contract, then completed the explanation with a convention that sounded reasonable. | 420 | 33.1% |
| Stopped searching after finding the first plausible documentation surface | The agent found a reasonable page to edit and did not continue checking other maintained, generated, mirrored, migration, or workflow surfaces affected by the same change. | 382 | 30.1% |
| Inspected relevant evidence but did not convert it into a complete coverage checklist | The agent reached evidence bearing on the requirement but began drafting without tracking the claims, setup, boundaries, examples, and reader actions that needed to survive into the final patch. | 349 | 27.5% |
| Committed too early to a narrow interpretation of the task | Before completing the investigation, the agent declared the task to be a rename, reference update, single-page edit, or similarly narrow deliverable and ignored evidence outside that frame. | 346 | 27.3% |
| Overgeneralized or misinterpreted partial evidence | The agent inspected relevant evidence but converted one branch, example, implementation detail, or deployment pattern into a broader or different public rule. | 333 | 26.3% |
Agents often lack a reliable test for whether they have gathered enough evidence. Sometimes agents stop researching before they reach the decisive evidence. Other times they begin drafting from partial information without recognizing that their evidence is incomplete. This pattern suggests a failure to recognize uncertainty. Prior work reports similar findings (Liu et al., 2025; Shao et al., 2025; Liu et al., 2025). Models struggle to identify the source of uncertainty. Information-seeking agents often answer before the available evidence is sufficient, and they do not reliably recognize when more information gathering has value.
Premature closure, shifting from investigation to drafting too soon, cuts across many of the root causes. Once an agent finds a plausible interpretation or a reasonable page to edit, it often starts drafting. It may stop before reaching key evidence, and it may also stop before checking every affected documentation surface, which leaves parts of the documentation stale. Insufficient search is only part of the problem. The larger part is that agents lack a reliable stopping rule. Such a rule would tell an agent when it understands the task, the evidence, and the documentation impact well enough to begin writing.
We found no clear relationship between the assigned root cause and trajectory length, whether measured by turns or by token use. Effort also did not rise with task difficulty. We defined an itemâs difficulty from the mergeability of the other six agentsâ patches on that item. Within each agent, the rank correlation between difficulty and effort was for processed tokens, for trace-event count, and for tool actions. All task-clustered 95% confidence intervals include zero. Premature closure therefore does not necessarily produce a short trajectory. An agent may stop investigating early and then spend substantial effort drafting, revising, or elaborating an incomplete account.
8 Discussion
8.1 Missing context about the reader
Some failures that we attribute to a missing model of the reader may instead reflect missing context about how the software is used. Agents cannot always infer from parametric knowledge alone what readers are trying to accomplish or which details they need. That inference is especially hard when user motivations and the surrounding workflow context are implicit rather than stated.
8.2 Coarse training rewards
Premature-closure failures may be related to coarse reward signals during post-training. RAGEN finds that trajectory-level rewards do not reliably teach agents how to reason through multi-turn tasks. Without fine-grained, reasoning-aware feedback, agents may learn shallow strategies or produce reasoning that is not grounded in the environment (Wang et al., 2025). Kim et al. report a similar pattern (Kim et al., 2026). In their experiments, outcome-only reinforcement learning improved final accuracy but made intermediate reasoning less accurate and less internally consistent. Models learned shortcuts rather than reliable reasoning procedures. These results offer possible explanations for our findings. The agents seemed to infer scope from early cues, such as the location of a code change or the name of a feature. They then began drafting within that narrow frame and filled evidence gaps with plausible assumptions that their environment did not support.
8.3 Knowing when to stop investigating
Another explanation is that deciding when the evidence is sufficient is itself a difficult capability. SeekBench reports that search agents trained with reinforcement learning answered before gathering sufficient evidence in 76.5% of the evaluated trajectories (Shao et al., 2025). CaRT shows that models may rely on superficial stopping rules, such as the number of turns, instead of checking for a decisive fact (Liu et al., 2025). Related studies report that language models struggle to retract an earlier inference when new evidence contradicts it and tend to seek examples that confirm an initial hypothesis rather than examples that might disprove it (Wilie et al., 2024; Jhaveri et al., 2026). These findings match the patterns in our trajectory analysis.
8.4 Additional guidance and scaffolding
Several changes could plausibly address the observed failures: explicit instructions in the prompts that ask agents to consider the readerâs goal, skills that emphasize task completion and findability, broader tools, scratch notes, and explicit verification. We explored these approaches informally but did not systematically compare them against a baseline, so we cannot conclude whether they improved documentation quality. Future work should test their effects through controlled comparisons.
9 Limitations
9.1 Measurement
Task-specific rubrics and LLM judges (Section 4) let us score open-ended documentation patches at scale. Because human validation covers only part of the evaluation, automated rubric generation and scoring may still introduce errors. We check commands and examples against the available evidence instead of running every documented procedure in its repositoryâs native build and runtime environment. As a result, the benchmark has no deterministic checks for code samples and links.
We find that generated rubrics tend to contain more criteria than maintainer-authored rubrics. Many of these additional criteria identify valid documentation improvements, but maintainers may consider them less important. During human calibration, reviewers prioritized the noncritical P1âP3 criteria differently. Depending on repository norms, some emphasized style, while others placed less weight on it. Some preferred comprehensive documentation, whereas others favored a simple, easy-to-follow user path over broader coverage. These preferences do not support a single universal weighting of P1âP3 criteria. The score therefore weights all criteria equally in the mean and handles P0 failures separately through the score cap. The published dataset retains the P0âP3 labels, and the accompanying scoring code lets practitioners apply other weights.
Manual review found that some rubric criteria overlap and are not fully independent. Some overlap is warranted, because a single documentation failure can cause several related problems. In a later audit, we tried to merge overlapping criteria. Merging sometimes lost important distinctions, so we kept the overlapping criteria.
9.2 No internet access
We evaluated agents without internet access (Sections 3.2 and 5). To assess how this restriction affected performance, we reviewed 683 trajectories from items on which no agent produced a mergeable patch. We found blocked network requests in 73 runs (10.7%). In 63 of these runs, network access was not necessary to produce a correct patch. The requests mostly involved setup or validation, such as installing dependencies, building documentation, running formatters or linters, and parsing YAML or JSON configuration files. Only 10 runs (1.5% of all audited trajectories) tried to retrieve external evidence that a correct patch required and the supplied inputs lacked.
Internet access might therefore have helped on a small number of items. However, in an earlier web-enabled pilot, 26 of 85 runs (30.6%) retrieved the itemâs upstream pull request and its merged documentation. Given this contamination risk, we kept the reported evaluation sealed.
9.3 Asymmetric label construction
The evidence for the abstention and patch labels (Section 3.1) is not equally strong. Documentation-needed items often have direct evidence. Some abstention labels, by contrast, rely only on the absence of a related documentation change within a 90-day audit window. That absence does not necessarily show that documentation was unnecessary. An update may have been forgotten, or it may have happened after 90 days without a link to the code pull request. As a result, some items that needed documentation may be mislabeled as abstention items. This asymmetric label noise could distort the measured decision performance.
9.4 Single-run evaluation
For budget and time reasons, we run each agent on each item once
(Section 5). We therefore do not report
pass@k, pass^k,
best-of-
performance, or within-item run-to-run variance. The results
characterize one sampled trajectory per agent and item. They do not
measure the probability that an agent reliably produces the same
decision or documentation quality across repeated attempts.
9.5 Low-information prose may be under-penalized
The rubric-based scoring (Section 4) and the patch-level failure rates (Section 7.2) may miss low-information prose. A general instruction to identify coherent but low-information prose flagged 3.2% of submissions. A second prompt asked the judge to flag submissions with two or more specific patterns. The patterns included redundant paraphrases, unnecessary explanations, excessive bulleted lists, formulaic contrasts (not X, but Y), three-part constructions, and heavy use of em dashes. This prompt flagged 7.6% of submissions. The increase suggests that LLM judges may miss low-information prose during scoring.
Disclosure
Two authors are affiliated with Promptless, a company that builds documentation agents.
10 Conclusion
DoGBench shows that even frontier models paired with frontier coding harnesses cannot yet reliably produce expert-level user-facing documentation in one attempt. The highest composite score on the 117-item held-out split is 47.3 out of 100. Agents still misjudge whether documentation is needed, and when they do edit, their patches may be factually correct but miss what readers need. Fluent prose and capable repository tooling do not yet close this gap.
DoGBench measures this gap with 292 real items, drawn from software changes and reported documentation gaps, and it shows where decisions and patches fail. We release the evaluation harness, item schema, dataset card, and development examples so that others can build on the benchmark. Benchmark materials and release information are available at https://dogbench.ai.
References
- Choudhury, S. Process reward models for LLM agents: Practical framework and directions. arXiv preprint arXiv:2502.10325, 2025.
- Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. ICML, 2023.
- Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436, 2019.
- Jimenez, C. E., Yang, J., Wettig, A., et al. SWE-bench: Can language models resolve real-world GitHub issues? ICLR, 2024.
- Jhaveri, A. R., GX-Chen, A., Sucholutsky, I., and Choi, E. Failing to falsify: Evaluating and mitigating confirmation bias in language models. arXiv preprint arXiv:2604.02485, 2026.
- Kim, K., Wang, K., Xie, Y., et al. Correct answers from sound reasoning: Verifiable process supervision for language models. COLM, 2026.
- Lightman, H., Kosaraju, V., Burda, Y., et al. Letâs verify step by step. arXiv preprint arXiv:2305.20050, 2023.
- Liu, G., Qu, Y., Schneider, J., Singh, A., and Kumar, A. CaRT: Teaching LLM agents to know when they know enough. arXiv preprint arXiv:2510.08517, 2025.
- OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/, 2026.
- Lu, S., Guo, D., Ren, S., et al. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. NeurIPS Datasets and Benchmarks, 2021.
- Liu, J., Peng, J., Wu, X., et al. Do not abstain! Identify and solve the uncertainty. ACL, pages 17177â17197, 2025.
- Nguyen Hoang, A., Le-Anh, M., Le, B., and Bui, N. D. Q. CodeWiki: Evaluating AIâs ability to generate holistic documentation for large-scale codebases. Findings of ACL, pages 5812â5827, 2026.
- Mei, W., Gu, Z., Bai, Z., et al. Deep Research as Rubric for Reinforcement Learning. arXiv preprint arXiv:2606.01091, 2026.
- Panickssery, A., Bowman, S. R., and Feng, S. LLM evaluators recognize and favor their own generations. NeurIPS, 2024.
- Pai, K., Devanbu, P., and Ahmed, T. CoDocBench: A dataset for code-documentation alignment in software maintenance. arXiv preprint arXiv:2502.00519, 2025.
- Sainz, O., Campos, J. A., GarcĂa-Ferrero, I., et al. NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. EMNLP Findings, 2023.
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., et al. Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. UIST, 2024.
- Shao, J., Lin, Y., Lohani, M. P., Miao, Y., and Luo, B. Do LLM agents know how to ground, recover, and assess? A benchmark for epistemic competence in information-seeking agents. arXiv preprint arXiv:2509.22391, 2025.
- Wang, X., Hu, R., Gao, C., Gao, P., and Peng, C. Evaluating repository-level software documentation via question answering and feature-driven development. arXiv preprint arXiv:2604.06793, 2026.
- Wang, Z., Wang, K., Wang, Q., et al. RAGEN: Understanding self-evolution in LLM agents via multi-turn reinforcement learning. arXiv preprint arXiv:2504.20073, 2025.
- Wilie, B., Cahyawijaya, S., Ishii, E., He, J., and Fung, P. Belief revision: The adaptability of large language models reasoning. EMNLP, pages 10480â10496, 2024.
- Ye, Z., Shi, W., Liu, Y., et al. Look before you leap: Autonomous exploration for LLM agents. arXiv preprint arXiv:2605.16143, 2026.
- Zhou, H., Huang, H., Long, Y., et al. Mitigating the bias of large language model evaluation. Proceedings of the 23rd Chinese National Conference on Computational Linguistics, pages 1310â1319, 2024.
- Zhou, J., Zhang, Q., Wang, Y., et al. RubricBench: Aligning model-generated rubrics with human standards. ACL, pages 31179â31200, 2026.
DoGBench