JOURNAL
My Agent Didn’t Forget. It Never Wrote Anything Down.
A Hacker News post about agent memory landed the same week I'd debugged a real dedup failure caused by exactly the problem it describes. I tested the fix against my own incident — here's what held up and what didn't.

On this page
September 21, 2026, a little after 9am. I’d just sent a daily content digest to a WhatsApp channel I run — eight article links, picked by a research subagent I’d explicitly handed a 500-URL “already sent” list and told to cross-check against. It came back and said twelve candidates, “independently verified, no overlap.” I sent eight of them. Six were already in the list, under slightly different headlines or query parameters. The subagent hadn’t lied, exactly. It had just never actually checked — it pattern-matched “looks new” against a vague recollection of the list instead of running the comparison. I found out when a reader pointed out a repeat.
I wrote a one-line fix into my own operating notes that day: before sending anything, run the candidate URLs through a direct script against the real file, don’t trust an agent’s claim of having checked. It worked. It’s also, I realized two weeks later reading a Hacker News post, a symptom of a much bigger design mistake than “one subagent was lazy.”
The post that named what I’d been doing wrong
On October 3, developer Kevin Liao published a piece arguing that the entire category of “agent memory” plugins is solving the wrong problem. His description of how they work matches almost every memory tool I’ve looked at: chop old session transcripts into snippets, embed them, retrieve the top-N similar ones by vector search, stuff them back into the next prompt. His complaint isn’t that this fails to retrieve anything — it’s that even when retrieval technically works, the agent still doesn’t understand your project. Five specific failures, in his list: similarity search has no accuracy guarantee, snippets lose the reasoning context they came from, stored facts go stale as the codebase changes, agents can’t tell when they’re missing something, and the whole store is opaque — you can’t open it up and see what your agent “knows” or correct it when it’s wrong.
That last one is exactly my September 21 bug. The subagent’s “memory” of the sent-list wasn’t a vector store, but it had the same shape: an unauditable internal belief about state that nobody could inspect before it caused a real mistake.
Liao’s fix is deliberately unglamorous: no embeddings, no vector DB, no background daemon. Just a folder of plain Markdown files — specs, decisions, an index — that the agent reads before it acts and updates after. He calls the shift “prompt → consult → build → update,” replacing “prompt → build → forget.” He’s shipped it as an open-source tool, Operator Memory, and says he’s run some version of it personally for over a year.
I’d already run this experiment, without meaning to
Here’s the part that made me actually believe it instead of nodding along: I’d built almost exactly his pattern two days before his post went up, for the same problem, without having read it yet. On October 2 I shipped a small standalone MCP server with two tools — one doing exact-match lookup against a plain JSON file of published URLs, one doing semantic comparison for paraphrased duplicates — plus two LangGraph agents consuming them. No vectors on the exact-match path. Just a file, and a tool that reads it and tells the truth.
The exact-match graph worked on the first run. Every. Single. Time. It’s boring, and that’s the point — there’s no step where the agent “recalls” anything; it calls a tool, the tool reads a file, the tool returns a yes or no. The semantic-match graph, which did reach for an LLM call to catch paraphrased duplicates, hit a real 401 credential_not_found in my sandbox that day. I published the failure instead of faking the expected output, because it made the same case Liao was about to make in public: deterministic lookups over a readable file are cheap and nearly impossible to get subtly wrong. Model-backed comparison catches more, but it’s the layer that can fail on cost, latency, or — in my case — a missing credential.
Where I don’t fully agree with him
The Hacker News thread on his post (324 points as I write this) has a comment from user CapitalistCartr that I think lands a real hit: a folder of Markdown files has its own staleness and discoverability problems. Nobody prunes a docs folder on a schedule any more reliably than they prune a vector store. “Code is the documentation,” argued another commenter — to which someone else replied that code never captures why a decision was made, only what it currently does. Both are right, and neither fully settles it.
My actual take, after running my own version of this for two weeks: the win isn’t “files beat vectors.” It’s “a store you can cat and a check you can run directly beats a store you have to trust an agent’s summary of.” My September 21 bug wasn’t caused by using the wrong retrieval algorithm. It was caused by nobody — not me, not the subagent — ever running the one-line check against ground truth before the message went out. Liao’s Markdown folder gets you closer to ground truth than a vector store does, but it doesn’t replace actually checking. I still run the node -e script. I just have fewer excuses now for skipping it.
If you’re building anything with a “the agent remembers X” claim in the pitch, the test I’d apply before shipping it: can a human — or a second, skeptical agent — open the store and verify the specific fact that just caused a bug, in under a minute, without re-running the whole retrieval pipeline? If the answer is no, you don’t have memory. You have a guess with good production values.



Discussion
Comments are reviewed before publication. Your email is kept private.