What happens when you let an AI agent patrol your codebase for documentation drift — and it catches a bug you didn't know existed.


Most documentation dies the same way: slowly, silently, and without anyone noticing until a new team member tries to onboard and everything is wrong.

I know this because I built 13 production AI services and watched their documentation rot in real time. Not because I'm lazy — because documentation staleness is a structural problem, not a discipline problem. And I got tired of pretending otherwise.

So I built Doc-Steward: an autonomous documentation agent that scans my entire codebase daily, detects when docs have drifted from reality, rewrites them, and submits pull requests for my review. It runs at 6:35 every morning. Most days I wake up and everything is fine. Some days there's a PR waiting. And one night, it found a bug in its own infrastructure that had been silently failing for hours.

Here's how it works, what it took to get right, and why the hardest part wasn't the AI.


The Documentation Problem Everyone Ignores

Documentation staleness isn't a content problem. It's a detection problem.

Your API endpoint changes. Your service adds a new capability. Your architecture evolves in a dozen small ways over a month. The code knows about all of this. The documentation doesn't. And nobody notices until someone — a new hire, a future version of you, an AI agent trying to integrate — reads the docs and does the wrong thing.

I measured this across my 13 services. When I first ran Doc-Steward's scanning logic, it reported 75% documentation coverage. That sounded reasonable until I looked closer: 7 of 13 services had stale registry entries pointing to documentation that was either outdated or full of TODO placeholders. Four services had no knowledge graph representation at all, meaning they were invisible to every other service in the ecosystem. The system was reporting "documented" for services that were, functionally, undocumented.

The gap between reported coverage and actual coverage is where documentation goes to die.


Why Docs Go Stale (It's Not Laziness — It's Structural)

The standard solution to documentation drift is "developers should update docs when they change code." This is the engineering equivalent of "people should eat more vegetables." Correct, unhelpful, and routinely ignored under pressure.

The structural problem is that documentation and code exist in different feedback loops. Code has tests, CI, type checkers — if you break something, the system yells at you within minutes. Documentation has nothing. You can let docs rot for months and nothing fails, nothing alerts, nothing even looks slightly different in your deployment pipeline.

For a solo operator running 13 services, this gap compounds fast. Every service I touch might invalidate documentation in two other services. The interconnection overhead that the knowledge graph handles for runtime discovery creates an equivalent documentation maintenance burden that nothing was handling.

I needed the same thing for documentation that I'd built for service discovery: an autonomous system that detects problems and fixes them without requiring me to remember.


How Doc-Steward Detects Drift

Doc-Steward's core is a scanning pipeline that runs against every service in the ecosystem. The detection layer does three things:

Existence checking. Does the service have a CLAUDE.md file at its canonical path? This sounds trivial, but the first scan revealed that 5 services had documentation files on disk that the services registry didn't know about. The files existed — the system just couldn't see them. A NULL claude_md_path in the registry made a service invisible to both the documentation drafter and the gap detector, creating a permanent blind spot.

Freshness tracking. Every documented service gets a checksum and a last-verified timestamp. If the CLAUDE.md hasn't been verified against the actual service state in 30 days, it gets flagged. The initial scan found zero checksum mismatches — when docs existed, they were at least internally consistent. The problem was always existence and completeness, not corruption.

Content depth analysis. This was the hardest lesson. Early versions only checked whether a file existed and whether it was fresh. But four services had CLAUDE.md files that were 80%+ TODO placeholders. The steward was reporting them as "documented." They were not. Adding a TODO-line ratio threshold to distinguish real documentation from scaffolding took a later iteration, but it eliminated the single biggest source of false confidence in the coverage numbers.

The detection pipeline feeds into a knowledge graph, creating documentation_gap nodes for every issue it finds. These nodes are typed, prioritized, and searchable — so any agent in the ecosystem can query "what documentation needs attention?" and get a structured answer.


The Self-Healing Pipeline: Detect → Register → Link → Retire

Detection without remediation is just a fancier way of ignoring the problem. The breakthrough was closing the loop — making Doc-Steward not just find gaps, but fix them.

This shipped in PR #103 as roughly 106 lines of Python in kg_doc_steward.py. The core addition was two new functions:

_register_filesystem_claude_mds walks the services table, checks the filesystem for CLAUDE.md files at canonical paths, and auto-registers any it finds as knowledge graph nodes. This was the critical piece. Before this, if you created a CLAUDE.md for a service, nothing connected it to the service's knowledge graph representation. The file sat there, invisible, while the gap detector kept flagging the service as undocumented.

_resolve_stale_gaps marks documentation gap nodes as inactive when their underlying conditions clear. Before this, gaps accumulated forever. The system could detect 10 gaps, you could fix all 10, and it would still report 10 open gaps because nothing retired the old nodes.

Both functions are gated behind an auto_fix=True parameter — no surprise behavior changes for anything using the default scan mode.

The result after deploying: open gaps dropped from 10 to 0 in a single scan. Documents edges went from 6 to 10. Coverage went from 75% to over 90%. The backlog of accumulated ghost gaps from months of compounding drift was cleared in one pass.


The Night It Fixed Its Own Bug

PR #104 shipped the same night as #103 — a four-item overnight build that included a daily launchd job so Doc-Steward would run automatically at 6:35 AM, a new service_update MCP tool so agents could repair the services registry without direct database access, and a fix for a coverage percentage math bug.

While verifying the new daily job, the system hit a 401 error from Memory Archive. Investigation revealed something unsettling: the Personal Assistant service's environment file contained MEMORY_ARCHIVE_TOKEN=your_memory_archive_api_token — the placeholder value. It had never been replaced with a real token.

This meant the existing poll-critical job — a separate scheduled task that checks for urgent documentation proposals — had been silently failing every 30 minutes for hours. It was hitting a 401, catching the error, and reporting "no new critical proposals" as if everything were fine. The exact silent-failure pattern its own docstring warned against.

This is the kind of bug that a human would never catch through code review. The code was correct. The configuration was wrong. And the failure mode was designed to look like success. Doc-Steward's infrastructure verification during deployment surfaced it because the system was testing its own integration paths, not just its logic.

The fix took 30 seconds — replace the placeholder with the real token from the credentials store. The bug had been invisible for an unknown period. Without the self-healing infrastructure forcing a full integration test during deployment, it would have stayed invisible indefinitely.


Three Iterations to Get Confidence Right

The gap detection pipeline needed a confidence model: how sure is the system that a gap is real versus a false positive? Getting this wrong in either direction is expensive. False positives waste the solo operator's review time. False negatives let drift compound.

Iteration one was binary: documented or not. This is what reported 75% coverage while reality was closer to 40% usable. It couldn't distinguish a complete CLAUDE.md from a file that was 80% TODO placeholders.

Iteration two added freshness and checksum tracking. This caught stale documents but still couldn't distinguish shallow documentation from deep documentation. A file with three lines and twelve TODOs scored the same as a comprehensive architectural overview.

Iteration three added content-depth heuristics: TODO-line ratios, section completeness checks, and a comparison between documented capabilities and the service's actual registered capabilities in the knowledge graph. This is what finally brought reported coverage in line with reality.

The key insight across all three iterations: confidence scoring isn't about the AI's ability to assess documentation quality. It's about defining what "documented" means precisely enough that a machine can measure it. The hardest part was the definition, not the implementation.


Why the Human-in-the-Loop Gate Matters More Than the AI

Doc-Steward can detect drift, draft corrections, and submit pull requests. It cannot merge its own PRs. This is deliberate.

The pipeline has three gates: proposal approval on a dashboard, PR creation, and PR merge. Each gate requires human action. The system generates the work and presents it for review — it doesn't execute autonomously.

This might sound like I'm hedging, adding friction to a system that could be fully automated. I'm not. The human gate is the architectural decision I'm most confident about, for three reasons.

First, documentation serves humans. An AI-generated correction that's technically accurate but confusingly written, or that reorganizes information in a way that breaks someone's mental model, is worse than the original stale version. Human review catches tone and structure problems that no automated check can.

Second, the review gate is where I learn. Every PR from Doc-Steward is a signal about what changed in my system that I might not have noticed. Scanning the diff takes 30 seconds. Writing the fix from scratch would take 30 minutes. The system is a force multiplier for my attention, not a replacement for it.

Third, and most practically: the system got its confidence scoring wrong twice before getting it right. During those iterations, the human gate prevented bad automated corrections from shipping to production documentation. The gate isn't just quality control — it's the safety net that makes aggressive iteration on the detection pipeline possible.


The Results, One Month In

Before Doc-Steward's self-healing pipeline: 75% reported coverage, ~40% actual usable documentation, 10 accumulated gap nodes that nobody was resolving, and documentation drift that compounded with every commit.

After one month of daily autonomous scanning: 92.3% coverage across 13 services, with the single remaining gap being a service that's embedded by design and doesn't need its own documentation. Zero accumulated stale gaps. Daily Pushover notifications when anything needs attention — which most days, it doesn't.

The daily scan runs at 6:35 AM. It walks the filesystem, checks for new or moved documentation files, verifies registry linkage, retires any gaps whose conditions have cleared, and sends a notification summarizing what it found. On a quiet day, the entire process takes under a minute and I never see it. On an active day, there's a PR waiting by the time I open my laptop.

The biggest surprise: the maintenance cost of the self-healing system itself is nearly zero. It's 106 lines of core logic, a launchd job, and a daily scan call. The ROI isn't measured in hours saved per week — it's measured in architectural confidence. I can make changes to any service knowing that documentation will either stay current or I'll be told exactly where it drifted.


The Takeaway for Builders

Documentation staleness is a detection problem disguised as a discipline problem. You don't fix it by trying harder — you fix it by building systems that notice drift and close the loop.

The specific implementation doesn't matter. What matters is the pattern: scan for gaps automatically, score confidence honestly, fix what you can programmatically, and put a human gate on everything that ships. The AI handles the tedious part (scanning 13 services daily for drift). The human handles the judgment part (does this correction actually make the documentation better).

If you're running production services and documentation debt keeps growing despite your best intentions, stop blaming your discipline. Start blaming your feedback loop. Code has CI. Documentation should too.


See what I've been up to: coreyscherrer.com

See what fun I've been up to: coreyiscorey.com


Disclaimer: These articles were drafted with AI assistance (Claude) and reviewed by a human. All projects, systems, and technical details described are real — sourced directly from production sessions captured in a PostgreSQL database. Questions? I'd love to talk shop — reach out anytime.