How a knowledge graph with 200+ nodes, pgvector embeddings, and a time-weighted decay algorithm keeps 13 AI services from drowning in their own context.


Every AI agent I've built hits the same wall around month three: it forgets everything.

Not literally — the code still works, the database still has data. But the context that makes the system useful, the accumulated knowledge about what works, what broke, and how things connect, lives nowhere persistent. Each new session starts from scratch. Each new integration requires re-explaining things the system "knew" last week.

This is the memory problem, and it's the single biggest gap between AI systems that demo well and AI systems that compound in value over time. I solved it by building Memory Archive: a knowledge graph with semantic search that serves as the shared brain for all 13 of my production AI services.

Here's the architecture, the algorithm that keeps it from becoming noise, and what I've learned about making AI memory actually work.


Why RAG Alone Isn't Enough

The default answer to AI memory is retrieval-augmented generation: embed your documents, store them in a vector database, retrieve relevant chunks at query time. RAG works well for document-grounded Q&A. It does not work well as a general-purpose memory system.

The problem is structural. RAG treats all information as flat documents. But the knowledge an AI system accumulates isn't flat — it's a graph of relationships. Service A depends on Service B. Decision X was made because of Learning Y. Pattern Z applies to Projects W, V, and U. Flattening these relationships into document chunks loses the structure that makes the knowledge useful.

Memory Archive combines two approaches: a typed knowledge graph for structured relationships, and pgvector embeddings for semantic search. The graph captures what things are and how they relate. The embeddings capture what things mean. Together, they let any service in the ecosystem ask both structured queries ("which services depend on the blog database?") and semantic queries ("what do we know about connection pool tuning?") against the same knowledge base.


The Knowledge Graph Design

Every piece of knowledge in Memory Archive is a node with a type, a maturity level, and an importance score.

Node types include memories (things the system has learned), decisions (choices that were made and why), patterns (recurring approaches), plans (active work in progress), recovery points (bookmarks for resuming context), and service nodes (the 13 services themselves).

Maturity levels track how well-established a piece of knowledge is: draft, canonical, or deprecated. A fresh observation starts as draft. Once verified or reinforced, it graduates to canonical. When superseded, it moves to deprecated.

Edges connect nodes with typed relationships: belongs_to, produced_by, tagged_with, documents, decided_at_version. These edges are what make the knowledge graph a graph rather than a flat collection.

The current graph has over 200 nodes and hundreds of edges across all 13 services, plus cross-cutting knowledge about the architecture itself.


Hybrid Search: Semantic + Keyword

Querying the knowledge graph uses a hybrid approach that combines semantic similarity with keyword matching.

Semantic search uses pgvector with OpenRouter's text-embedding-3-small model (1,536 dimensions). Every memory's content is embedded and stored alongside the node. When a service queries "what do we know about documentation drift?", the system finds the most semantically similar memories regardless of exact keyword matches.

Keyword search catches cases where semantic similarity misses: exact service names, specific PR numbers, error codes, or technical terms that embedding models don't reliably cluster. The hybrid approach runs both searches and merges results with configurable weighting.

One lesson learned the hard way: embedding model input limits matter more than you'd expect. OpenRouter's text-embedding-3-small has an 8,191-token limit, but ChatGPT conversation exports and code-heavy content compress at roughly 2-3 characters per token instead of the typical 4. I was truncating at 30,000 characters and getting API rejections on batches with dense JSON content. Dropping to 8,000 characters eliminated all embedding failures. The full 23,465-message Discotheque corpus embedded cleanly in 25.8 minutes at that threshold.


The Decay Algorithm: Forgetting What Doesn't Matter

A knowledge graph without pruning becomes a junk drawer within weeks. Without a mechanism for the system to forget, search quality degrades as noise overwhelms signal.

Memory Archive uses a time-weighted importance decay algorithm that mimics how human memory works: recent and frequently-accessed information stays prominent, while stale knowledge gradually fades.

Every node has an importance score between 0 and 1. Over time, that score decays. When a memory gets accessed, its score bumps back up. The system never deletes anything; it deprioritizes.

Getting the decay curve right took three iterations:

Too aggressive. The first version decayed scores on a daily basis. Within two weeks, foundational architectural decisions had faded below the search threshold.

Too gentle. The second version barely decayed at all. After a month, search results for common queries were polluted with one-off observations from early sessions.

The working version. The current algorithm combines time-weighted decay with access-frequency boosting and type-based floors. Architectural decisions decay slower than one-off observations. The combination has been stable for months.


Integration Patterns: How Services Use the Memory Layer

Service discovery. Services discover each other through meaning, not configuration.

Context enrichment. Doc-Steward queries Memory Archive for recent architectural decisions that might affect documentation.

Cross-service learning. When one service learns something, that learning is available to every other service.

Session continuity. Recovery point nodes bookmark the state of a complex task for resumption in future sessions.


The Counterintuitive Takeaway

Most AI memory systems are designed to maximize recall. That's the wrong optimization target.

The right optimization target is signal-to-noise ratio over time. A memory system that remembers 1,000 things with declining search quality is less valuable than one that effectively surfaces the 50 things that matter right now. The decay algorithm isn't a compromise — it's the core feature.

If you're building AI systems that need persistent context, start with three things: a typed knowledge graph for structured relationships, vector embeddings for semantic search, and a decay mechanism that ensures your system forgets gracefully. Everything else is optimization.


See what I've been up to: coreyscherrer.com

See what fun I've been up to: coreyiscorey.com


Disclaimer: These articles were drafted with AI assistance (Claude) and reviewed by a human. All projects, systems, and technical details described are real — sourced directly from production sessions captured in a PostgreSQL database. Questions? I'd love to talk shop — reach out anytime.