Research readout
DoGBench: Can agents meet expert standards for user-facing documentation?
A benchmark of real repository events, scored against the criteria documentation maintainers actually apply in review.
01 · Provenance
Built by expert technical writers from widely used open source projects
Real events
Items are pull requests or reported documentation gaps from open-source projects. Rubrics reward actual content instead of similarity to the docs that humans merged.
Validated with maintainers
Maintainers reviewed 42 items. The rubrics recover 90.9% of the criteria maintainers wrote themselves, and the scoring model agrees with human judgments on 93.9% of criteria.
02 · The task
What the benchmark asks an agent to do
Step 1 — Decide
Does this change need docs?
The agent sees a pull request or a user-reported gap in a repository, and must return either a patch or an explicit no-op.
Step 2 — Write
Is the patch good?
Scored 0–100 on accuracy, completeness, placement and reader guidance. One failed P0 requirement caps the score at 60.
03 · Results
The best agents barely cleared 50 out of 100
| # | Logo | System | Composite | Delivered quality |
Patch quality |
Abstention recall |
|---|
Composite combines two things: Delivered quality, the average rubric score on the items that needed a patch (a missed patch scores zero), and Abstention recall, how often the agent left docs alone on the items that needed no change. An agent has to do well on both. Cloud agents were told not to use the internet. We verified that Promptless did not browse the internet to find contaminating information, but we could not verify this for other cloud agents.
04 · Results
Humans are better at judging what matters and making sure we get them right
05 · Results
Agents don’t know when not to write
Abstention recall (%): of the items where no documentation change was needed, how often the agent correctly wrote nothing.
Why does it matter?
Internal detail leaks out
Implementation details users can't act on, published as product.
Reviewer burden
Each one still has to be read and rejected.
Maintenance burden
If merged, it must be versioned, translated and kept true.
Wrong precedent
The next agent run reads it as a pattern to follow.
06 · Failure analysis
What is actually wrong with the patches
Task-completion gap
Omits a decision point, procedure, verification step, or recovery path the reader needs.
Technical inaccuracy
Misstates an interface or behavior, including its scope, default, lifecycle, or compatibility.
Incomplete conceptual or reference coverage
Omits the core concept, capability, or contract the reader needs to understand and use the change.
Missing supporting information
Covers the main task but omits rationale, boundaries, examples, or secondary cases.
Missing prerequisites
Omits permissions, dependencies, credentials, versions, or other setup conditions.
Hard to find
Puts content on a page readers won't reach, or never links to it from where they are.
Missing audience and purpose framing
Does not establish who the content is for, why it matters, or when to use it.
Out-of-date related pages
Updates one page but leaves other pages that describe the same thing out of date.
Share of 1,267 audited submissions. Labels selected for at least 10% of submissions; multiple modes may apply to one patch.
07 · Failure analysis
The dominant failure mode: accurate text that still leaves the reader stuck
Agents treat a change as one piece of information to convey, rather than one step in a journey that needs prerequisites, decisions, verification and recovery.
Example · Strawberry GraphQL
The patch correctly documented how to select an older Apollo Federation version, but not why a reader would want to: to upgrade Strawberry while staying compatible with an older Apollo Router or Gateway.
It documented the setting and omitted the decision the setting exists to support.
The same narrowness at corpus scale
Agents can add accurate text to a page the reader is unlikely to visit, or update one page and leave others that cover the same thing out of date.
A patch can be accurate in isolation and still fail inside the documentation system.
08 · Failure analysis
Why it happens: agents have no model of how the product is actually used
Trace-supported root causes across all 1,267 audited submissions. One or more per material problem.
Explained the interface without checking how it is used
Assumed instead of reading the decisive evidence
Stopped at the first plausible page
Read the evidence, didn't use it
Committed early to a narrow reading of the task
Generalized one example into a public rule
09 · Failure analysis
AI is not very good at detecting AI-slop, and hallucination has gotten subtler: less invention, more misstated facts
had low information density: repetition, extra structure, long prose. The judges may under-count this — a looser prompt flagged only 3.2%.
contained fabricated content: invented classes, flags, endpoints.
had technical inaccuracies. The feature is real, but the patch overstates its scope, default or compatibility.
10 · Implications
What this means for how we work
Review the decision, not just the prose
Sometimes, there shouldn't be any documentation for a particular code change at all. Writer still needs to triage.
Read for the reader's whole task
Check prerequisites, decision points, verification and recovery. Agents still rely on humans to keep the full user journey in mind.
Check scope and conditions, not just facts
Defaults, versions, lifecycle and compatibility are where accurate sentences go wrong.
Look beyond the page that was edited
Navigation, related pages and cross-references sit outside an agent's stopping point.
DoGBench