The 502 error that took down my knowledge graph had nothing to do with AI. That's the point.


Last month, my Memory Archive — the central knowledge graph connecting 13 production AI services — started throwing sustained 502 errors. Users of Career Bot got blank responses. Doc-Steward couldn't check documentation drift. The Blog App couldn't query for related content.

The cause? Connection pool exhaustion. Not a hallucination. Not a prompt injection. Not a model capability gap. A PostgreSQL connection pool running dry because knowledge graph queries were holding connections open too long under load.

This is the reality of production AI that the demo-to-deployment pipeline doesn't prepare you for. The failures that actually take your system down are almost never AI failures. They're distributed systems failures wearing an AI costume.


The Gap Between Demo and Production

There's a statistic that gets cited in every enterprise AI report: roughly 90% of AI projects fail to reach production. The number varies by source, but the direction is consistent. Most AI doesn't ship.

The conventional explanation is that the models aren't good enough, or the data isn't clean enough, or the use case wasn't viable. Sometimes that's true. But in my experience building and operating 13 production services, the failures cluster somewhere more mundane: integration, operations, and the infrastructure nobody thinks about until it breaks.

Industry surveys consistently show that integration challenges — not model performance — are the number one obstacle teams cite when moving AI from prototype to production. This tracks perfectly with my experience. The AI part of my system works remarkably well. The parts that break are the same parts that break in any distributed system.


The Actual Stack Nobody Talks About

Here's what my production AI infrastructure actually depends on, ordered by how often each layer has caused an outage:

Connection pool management. Every one of my 13 Flask services shares a PostgreSQL backend. Each service runs its own Gunicorn workers, each worker maintains its own connection pool. When I tuned these wrong — pool_size too small, no overflow, no pre_ping — services would intermittently lose database connectivity. The fix was QueuePool with pool_size=5, max_overflow=10, pool_pre_ping=True. Not glamorous. Completely essential.

Health check infrastructure. 11 service ports need monitoring. I run a site-wide health sweep that hits every port and reports status. When one service's health endpoint returned 404 instead of 200 (because I forgot to register the route), the monitoring assumed the service was down and I spent 30 minutes debugging a non-problem.

Secret management. During an overnight deployment, Doc-Steward discovered that the Personal Assistant service had been running with a placeholder API token for an unknown period. The service was silently catching 401 errors and reporting "no new critical proposals" as if nothing were wrong. The failure mode was designed to look like success.

Token refresh and authentication. Services authenticate to each other through Memory Archive's API. Token expiration, rotation, and refresh logic sounds trivial until it's 2 AM and you're trying to figure out why Career Bot can't reach the knowledge graph despite both services being healthy.

Graceful degradation. What happens when Memory Archive is down? In a memory-first architecture, a central outage means total service discovery failure. That's a deliberate tradeoff — correlated failure in exchange for operational simplicity.


Why Distributed Systems Skills Matter More Than Prompt Engineering

This is the contrarian take: for production AI, being good at distributed systems is more valuable than being good at prompts.

The gap between a prototype that demos well and a system that runs reliably in production is filled almost entirely with distributed systems concerns: connection management, retry logic, circuit breakers, timeout tuning, error propagation, graceful degradation, and monitoring.


A Production-Readiness Checklist (The Boring Version)

Based on two years of operating 13 services:

Connection management. Do your database connections pool correctly? Do they time out? Do they recover after a connection is dropped?

Health endpoints. Does every service expose a health check that actually verifies dependencies?

Secret hygiene. Are any of your environment files still using placeholder values?

Failure modes. What happens when each dependency is unavailable?

Monitoring and alerting. When something breaks, how long until you know?

Deployment verification. After every deploy, do you verify end-to-end?


What the 10% Do Differently

The projects that reach production and stay there don't have better models. They have better infrastructure. Production AI is 20% model and 80% plumbing. The plumbing isn't exciting. But it's the difference between a demo and a system.


See what I've been up to: coreyscherrer.com

See what fun I've been up to: coreyiscorey.com


Disclaimer: These articles were drafted with AI assistance (Claude) and reviewed by a human. All projects, systems, and technical details described are real — sourced directly from production sessions captured in a PostgreSQL database. Questions? I'd love to talk shop — reach out anytime.