Research readout slides

For documentation owners who want to use this research with their teams or leadership. Present the slides from this page, or download them as a PDF.

Download PDF PDF with speaker notes

Open speaker notes shows the notes and the next slide in a second window. Keep it on your screen, then select Present. Both windows follow the arrow keys.

Promptless

Research readout

DoGBench: Can agents meet expert standards for user-facing documentation?

A benchmark of real repository events, scored against the criteria documentation maintainers actually apply in review.

systems evaluated: local lanes, cloud agents dogbench.ai

01 · Provenance

Built by expert technical writers from widely used open source projects

Helm PostHog Mautic

Real events

Items are pull requests or reported documentation gaps from open-source projects. Rubrics reward actual content instead of similarity to the docs that humans merged.

Validated with maintainers

Maintainers reviewed 42 items. The rubrics recover 90.9% of the criteria maintainers wrote themselves, and the scoring model agrees with human judgments on 93.9% of criteria.

02 · The task

What the benchmark asks an agent to do

Step 1 — Decide

Does this change need docs?

The agent sees a pull request or a user-reported gap in a repository, and must return either a patch or an explicit no-op.

Step 2 — Write

Is the patch good?

Scored 0–100 on accuracy, completeness, placement and reader guidance. One failed P0 requirement caps the score at 60.

real repository events
/ need a patch / need no change

03 · Results

The best agents barely cleared 50 out of 100

# Logo System Composite Delivered
quality
Patch
quality
Abstention
recall

Composite combines two things: Delivered quality, the average rubric score on the items that needed a patch (a missed patch scores zero), and Abstention recall, how often the agent left docs alone on the items that needed no change. An agent has to do well on both. Cloud agents were told not to use the internet. We verified that Promptless did not browse the internet to find contaminating information, but we could not verify this for other cloud agents.

04 · Results

Humans are better at judging what matters and making sure we get them right

05 · Results

Agents don’t know when not to write

Abstention recall (%): of the items where no documentation change was needed, how often the agent correctly wrote nothing.

Why does it matter?

Internal detail leaks out

Implementation details users can't act on, published as product.

Reviewer burden

Each one still has to be read and rejected.

Maintenance burden

If merged, it must be versioned, translated and kept true.

Wrong precedent

The next agent run reads it as a pattern to follow.

06 · Failure analysis

What is actually wrong with the patches

Task-completion gap

Omits a decision point, procedure, verification step, or recovery path the reader needs.

45.5%

Technical inaccuracy

Misstates an interface or behavior, including its scope, default, lifecycle, or compatibility.

36.6%

Incomplete conceptual or reference coverage

Omits the core concept, capability, or contract the reader needs to understand and use the change.

32.5%

Missing supporting information

Covers the main task but omits rationale, boundaries, examples, or secondary cases.

26.5%

Missing prerequisites

Omits permissions, dependencies, credentials, versions, or other setup conditions.

21.0%

Hard to find

Puts content on a page readers won't reach, or never links to it from where they are.

20.7%

Missing audience and purpose framing

Does not establish who the content is for, why it matters, or when to use it.

19.9%

Out-of-date related pages

Updates one page but leaves other pages that describe the same thing out of date.

13.3%

Share of 1,267 audited submissions. Labels selected for at least 10% of submissions; multiple modes may apply to one patch.

07 · Failure analysis

The dominant failure mode: accurate text that still leaves the reader stuck

Agents treat a change as one piece of information to convey, rather than one step in a journey that needs prerequisites, decisions, verification and recovery.

Example · Strawberry GraphQL

The patch correctly documented how to select an older Apollo Federation version, but not why a reader would want to: to upgrade Strawberry while staying compatible with an older Apollo Router or Gateway.

It documented the setting and omitted the decision the setting exists to support.

The same narrowness at corpus scale

Agents can add accurate text to a page the reader is unlikely to visit, or update one page and leave others that cover the same thing out of date.

20.7%hard to find
13.3%related pages left stale

A patch can be accurate in isolation and still fail inside the documentation system.

08 · Failure analysis

Why it happens: agents have no model of how the product is actually used

Trace-supported root causes across all 1,267 audited submissions. One or more per material problem.

36.0%

Explained the interface without checking how it is used

33.1%

Assumed instead of reading the decisive evidence

30.1%

Stopped at the first plausible page

27.5%

Read the evidence, didn't use it

27.3%

Committed early to a narrow reading of the task

26.3%

Generalized one example into a public rule

09 · Failure analysis

AI is not very good at detecting AI-slop, and hallucination has gotten subtler: less invention, more misstated facts

7.6%

had low information density: repetition, extra structure, long prose. The judges may under-count this — a looser prompt flagged only 3.2%.

6.1%

contained fabricated content: invented classes, flags, endpoints.

36.6%

had technical inaccuracies. The feature is real, but the patch overstates its scope, default or compatibility.

10 · Implications

What this means for how we work

Review the decision, not just the prose

Sometimes, there shouldn't be any documentation for a particular code change at all. Writer still needs to triage.

Read for the reader's whole task

Check prerequisites, decision points, verification and recovery. Agents still rely on humans to keep the full user journey in mind.

Check scope and conditions, not just facts

Defaults, versions, lifecycle and compatibility are where accurate sentences go wrong.

Look beyond the page that was edited

Navigation, related pages and cross-references sit outside an agent's stopping point.