The One That Actually Hurt

I've shipped a lot of stuff. Most projects go like this: confused, then productive, then a day of debugging, then it's out the door. Fine. Next one.

Memory Archive MCP was not that. I lost an entire weekend to problems I didn't know existed, and spent part of a Sunday night seriously considering whether I should crawl back to a nice safe PM job where the worst thing that happens all week is someone moving a standup.

On paper it's nothing special. An MCP server that gives Claude a persistent memory layer: store conversations, pull them back by semantic search, surface the right context so I'm not re-explaining my life every new chat. Should've been a weekend. A little tool for me.

It was not a weekend.

Where It All Went Sideways

Nobody warns you that memory systems fail differently than a normal CRUD app. A broken API throws. You see the error and fix it. A broken memory layer fails quietly. It hands you back something, just not the thing you wanted. And Claude, helpful little assistant that it is, will cheerfully build a whole response on top of the wrong context and never once flag that the retrieval was garbage.

I spent hours convinced my embeddings were broken. They weren't. Retrieval was technically doing its job, returning the top-k most similar memories. The problem is that "most similar" by cosine distance is not the same thing as "actually relevant to what I'm asking right now." I'd ask about a project from last month and get five memories back about a completely different project that happened to share vocabulary. No error. No warning. Just confidently wrong.

Then the chunking. I'd assumed (wrongly, and at significant cost to my Saturday) that I could dump entire conversations into the store and let semantic search figure it out. A 4000-token conversation embedded as a single vector is basically meaningless. The embedding averages everything into mush. Two conversations about totally unrelated topics end up looking similar because they both contain "the" and "I think" and "actually."

Fix was smaller chunks with overlap, plus metadata I could filter on before semantic search even ran. Slower to ingest. Worked.

The Part That Actually Made It Hard

The code wasn't the hard part. MCP has decent docs. Embedding libraries are fine. The hard part is that I couldn't tell when the system was wrong.

Normal software has tests. Assert 2 + 2 = 4, done. With memory retrieval, the output is "here are five memories I think are relevant." How do you assert on that? What's the unit test for "this is the right context for this question"?

I tried eval sets. Picked twenty real questions I actually wanted the system to handle, hand-labeled which memories should come back, scored the retrievals. Helped some. But every time I added new memories the distribution shifted and the evals went stale. I haven't solved this. I've just gotten better at noticing when the system is lying to me.

That's the thing I keep coming back to. The hardest part of building anything with an LLM in the loop isn't the LLM. It isn't the framework. It's that your system can fail in ways that look fine, and the only defense is a paranoid, slightly grumpy human staring at the output going "wait, that's not what I asked."

I shipped it. I use it every day. It's not done and probably never will be. The next version involves me ripping out the retrieval scoring and rebuilding it, because I'm now pretty sure I was wrong about something foundational. Fine. That's how it goes.

If you've built anything memory-adjacent for an LLM and you think it's working great, I've got one question for you: how would you know?