Skip to content

thought-leadership · AI

Where AI helps technical writers - the evidence

Vlad Kuzin
On this page

AI is most useful to technical writers around the writing, not inside it. Writers already draft with it: 76% of documentation practitioners use AI for creation, according to the State of Docs 2026 report. Yet the same report found technical writers get the smallest time savings of any role. The report's own diagnosis: "The bottleneck in documentation has never really been the writing. It was always the information gathering, the change detection, the verification."

That sentence is the whole argument. A plan that expects AI to write the docs speeds up the one step that was never slow. A plan that points AI at change detection, verification, and gap finding speeds up the steps that were.

This page collects the evidence for both halves, with a link for every claim. It is meant to be forwarded.

If you need one line for a manager: AI drafts fast, but documentation was never limited by drafting speed. It is limited by knowing what changed and proving the page still matches the product.

Two ways to aim AI at documentation

AI writes the docsAI detects and verifies
Step it speeds upDrafting, where writers report the smallest savingsChange detection, testing, and gap finding, the steps State of Docs names as the bottleneck
Input it needsProduct truth the model has no access toCode diffs, test runs, and reader questions it can read directly
Typical failureA fluent, plausible error that ships to readersA false alarm the writer dismisses
Who checks the outputEngineers asked to review draftsThe writer, before anything publishes
AccountabilityUnclear when nobody verified the pageStays with the writer who acts on the finding

Why "AI writes the docs" aims at the wrong step

1. The model can only write what its inputs contain. Fabrizio Ferri-Benedetti calls the missing ingredient product truth: the support tickets, customer feedback, and release tension behind every mature feature, which no model can see unless someone hands it over. Tom Johnson gives the concrete failure in his principles for the cyborg technical writer: feed a model a product requirements document that describes the full vision while you ship version 1, and "the AI will likely hallucinate features that aren't there yet."

2. Writers are not told about every change. When Johnson had AI compare API definition files between releases, the diffs showed changes the writers had never heard about. His conclusion: "Engineers aren't telling writers" about them. An AI drafter waiting for input inherits the same blind spot. Detecting the change is the scarce step.

3. Drafting got instant, so verification became the job. Johnson again: "AI generates content almost instantly, making verification the slowdown point." Readers feel the result. In the Stack Overflow 2025 developer survey, 66% of developers named "AI solutions that are almost right, but not quite" as their biggest frustration, and 46% actively distrust AI accuracy against 33% who trust it.

4. AI makes wrong answers more convincing. In the BCG and Harvard field experiment with 758 consultants, published in Organization Science, participants using AI on a task chosen to sit beyond its abilities were 19% less likely to reach a correct solution. Ethan Mollick's summary adds the detail that matters for documentation: the AI users wrote up their wrong answers better than the control group. Fluent prose around a wrong fact is the exact failure a docs team exists to prevent.

5. Models invent technical specifics with confidence. A study of code-generating models found that at least 5.2% of package recommendations from commercial models, and 21.7% from open-source models, named packages that do not exist. Commands, parameters, and settings in documentation carry the same risk.

6. People misjudge their own speed with AI. In METR's early-2025 randomized trial, experienced developers took 19% longer with AI tools while believing they had been 20% faster. Treat the slowdown as dated: METR's February 2026 update says developers are likely faster with today's tools. The perception gap is the durable lesson. A team that feels faster has not measured anything.

7. The company owns every wrong answer. In Moffatt v. Air Canada, a Canadian tribunal held the airline liable for its chatbot's wrong refund advice: "It makes no difference whether the information comes from a static page or a chatbot." Cursor's support bot invented a login policy and triggered cancellation threats. Deloitte Australia partially refunded A$440,000 for a government report with AI-generated errors, including a fabricated court quote.

8. Unverified drafts move the work to someone else. BetterUp Labs and the Stanford Social Media Lab coined "workslop": AI output "that looks good, but lacks substance." In their survey of 1,150 US desk workers, 40% had received workslop in the past month, at about 2 hours to resolve each case. In documentation, that cleanup lands on the engineers asked to review the draft.

9. Documentation is now what the AI reads. Product chatbots, coding agents, and answer engines retrieve from published docs. Ferri-Benedetti puts it bluntly in his letter to teams that cut writers: "even your favorite AI must RTFM." Sarah O'Keefe of Scriptorium makes the same point from the content-strategy side: if the content is not accurate, current, and consistent, "the AI will not produce good output." Generating the source with the same models that will read it removes the one independent check in the loop.

Where AI pays off for technical writers

Each item below is in production use somewhere, with a public source. Vendor pages describe what the vendor claims; independent evaluations of output quality are still rare.

1. Docs as tests. Run the documented steps against the live product and fail when a step breaks. Doc Detective, the open-source tool behind the Docs as Tests approach, "performs your instructions step-by-step, just like your users would." Google Cloud uses Gemini to turn quickstart steps into Playwright scripts that execute in real environments.

2. Screenshot drift. Recapture screenshots on every run and compare them with the stored version. Doc Detective's screenshot action passes a step only when the new image stays within a set variation of the reference.

3. Change detection from code. Have AI read the diff between two releases and list the user-facing changes, as Johnson describes. Commercial tools now watch merged pull requests and propose doc updates for a writer to review.

4. Style and terminology checks. Encode the style guide in Vale, which ships Microsoft, Google, and Red Hat rule packages, and run it in CI. Elastic's docs team layers an LLM on top: "A deterministic Vale pass runs first. The agent adds other checks on top."

5. Stale-content and duplicate sweeps. Elastic runs scheduled agents that flag old pages, screenshots that fell behind their pages, references to versions past end of life, and pages that duplicate or contradict the rest of the corpus, per Ferri-Benedetti's write-up.

6. Gap finding from reader questions. Answer tools can cluster the questions they failed to answer. kapa.ai's Coverage Gaps does this, and its own docs add that the suggestions "require human review." The writer decides what the gap means.

7. Getting up to speed. Johnson reports reaching a working understanding of a new product after an hour of having AI distill more than 100 pages of source documents. Summarizing engineering material for the writer is lower risk than summarizing for the reader, because the writer checks it.

8. Structural transformations. Converting markup, building tables from unstructured notes, and suggesting metadata tags are mechanical jobs with checkable output. Google Cloud's writers use Gemini for "translating between markup languages" and "generating formatted tables from unstructured content."

9. Code-sample testing. Build, lint, and run every sample before it ships. Google Cloud's pipeline does this for AI-generated samples; rustdoc has run documentation examples as tests for years.

10. Machine translation with human post-editing. Autodesk measured translator productivity on post-edited machine translation of software documentation back in 2012. The pattern is mature; the new models raise the quality of the first pass.

11. Serving docs to AI tools. Publish an llms.txt index and Markdown versions of pages, or a docs server for AI agents, as Microsoft did with its Learn MCP Server. The writer's output becomes the context every downstream assistant answers from.

What the evidence does not show

A brief like this is only useful if it survives a skeptical reader, so here are its limits.

  • The METR slowdown is dated. METR itself says current tools likely speed developers up. Cite the perception gap, not "AI makes you slower."
  • The BCG result applies to tasks beyond AI's reach. Inside that boundary, the same study found large gains. The lesson is that nobody can see the boundary from the output.
  • Three of the sources are surveys. State of Docs, Stack Overflow, and the workslop study rely on self-reported answers, and a vendor runs the workslop study.
  • Vendor capability claims are not quality evidence. Tools that draft updates from pull requests are new, and no independent evaluation of their accuracy has been published.
  • None of this says writers should avoid drafting with AI. Most already do. The claim is narrower: drafting is the step where AI saves writers the least, so a plan built around it buys the smallest gain.

How to pilot it

Pick one bottleneck and one number, then measure before and after.

  1. Test your ten most-visited procedures. Run them as docs tests against the current release and count the steps that fail. Every failure is a support ticket that has not happened yet.
  2. Diff one release. Have AI read the code diff between your last two releases, list the user-facing changes, and count the ones that never reached the writers.
  3. Mine one month of unanswered questions. Cluster what readers searched for and did not find, and rank the clusters by volume.

The best first pilot is the procedure test, because every failure it finds is a defect a reader would otherwise hit, and the count needs no interpretation.

Each pilot runs with open-source tools: Doc Detective, Vale, a CI runner, and a link checker such as lychee. I build Topicary, which bundles pull-request watching, Doc Detective result ingestion, and a ranked gap worklist in one tool, but the measurement matters more than the tool. A number from your own docs will move a planning meeting further than any vendor demo, including mine.

FAQ

Frequently asked

Will AI replace technical writers?

The evidence points the other way. In the State of Docs 2026 survey, 76% of documentation practitioners already use AI for creation, yet technical writers report the smallest time savings of any role. The bottleneck is gathering information, detecting changes, and verifying accuracy, and those steps need someone accountable for product truth.

What is the best use of AI in technical writing?

The highest-value uses sit around the writing, not inside it: flagging doc impact when code changes, running documented procedures against the real product, catching outdated screenshots, sweeping for stale pages, and finding gaps from questions readers could not get answered.

Why is AI-generated documentation risky?

A model can only write what its inputs contain, and it writes wrong answers as fluently as right ones. In a study of 758 BCG consultants, participants using AI on a task beyond its abilities were 19% less likely to produce a correct solution. A tribunal also held Air Canada liable for its chatbot's wrong answer.

What are docs as tests?

Docs as tests means running the steps in your documentation against the live product, the way a user would, and failing the check when a step no longer works. Doc Detective is the main open-source tool; Google Cloud uses Gemini to turn quickstart steps into Playwright scripts that run in real environments.

How should a documentation team pilot AI?

Pick one bottleneck and one number. Run your ten most-visited procedures as tests and count the failures, or have AI read the code diff between two releases and count the user-facing changes nobody told the writers about. A measured result persuades faster than a drafting demo.

Ready to try Topicary?

Start free. No credit card required.