Blind model comparison
Model comparison: Agents Don't Need Memory
A blind comparison between two summaries of Kevin Liao's essay about documentation-based agent memory.
Read before revealing
The model names are hidden. Read both versions, choose which one you prefer, then reveal who wrote each. Nothing on this page is sent anywhere.
Version A
Qwen3.8-27B · Q4_K_M · local llama.cpp
Kevin Liao’s argument is that the entire agent-memory plugin category is built on the wrong diagnosis. The standard product — scan transcripts, generate snippets, store them in a vector database, retrieve the top five on every prompt — is RAG aimed at your own conversation history. And like all RAG, it fails in the same predictable ways.
Why recall-based memory breaks down
- Retrieval ranks by embedding similarity — it can’t tell you what’s current, correct, or missing
- Snippets are stored without context: the motivation, the lesson, the environment all get cut
- The past is treated as truth, but the codebase changes daily — 500 stale snippets about authentication quietly rot
- An agent can’t search for what it doesn’t know; it has no signal for when to reach for the search tool
- The store is unauditable: thousands of embeddings, no way to see which are stale, unused, or wrong
The whole thesis — “agents forget, so remember better” — is how no one else handles knowledge. Nobody rewatches a team meeting from three years ago to recall a constraint. People write things down.
The alternative: a document-based workspace
Liao’s system (he’s run it for over a year, now open-sourced as Operator Memory) replaces the vector store with a plain-Markdown “brain”: instructions, specs, decisions, research, and indexes. The agent consults the relevant documents before working and updates them after. The loop changes from prompt → build → forget to prompt → consult → build → update.
- No vector database, no embeddings, no background daemons — no “dreamers” rewriting memories overnight
- Everything is a file you can read, edit, commit, and share with your team
- It generalizes
AGENTS.md: one file is the floor, not the ceiling — the project needs a whole structured workspace
The timing point lands hard: AI is producing features at a speed where nobody reads the code, which makes documentation more important, not less. A memory plugin is a black box you can’t audit; a documentation workspace is just your repo doing something useful.
Version B
GPT-5.6 · OpenAI Codex
Kevin Liao argues that most agent-memory products confuse retrieval with understanding. Their usual design scans old conversations, turns them into small “memories,” stores those fragments in a vector database, and injects the most similar results into future prompts. That can recover related text, but it does not give an agent a reliable model of a changing project.
Where retrieval-based memory fails
- Similarity does not establish that a memory is current, correct, or complete.
- Small fragments lose the decisions, motives, and operating context that made them meaningful.
- Old conversations become stale as the codebase changes, yet retrieval continues treating them as evidence.
- An agent cannot deliberately search for information it does not realize exists.
- A large embedding store is difficult for a person to inspect, correct, or prune.
Liao’s alternative is less exotic: maintain project knowledge as structured documentation. Instead of attempting to reconstruct intent from old transcripts, the agent reads relevant instructions, specifications, decisions, research, and indexes before beginning work. When the task ends, it updates those documents while the new context is still available. He describes the resulting loop as prompt → consult → build → update, replacing prompt → build → forget.
His implementation, Operator Memory, expands the idea behind AGENTS.md into a Markdown workspace. The files can be read by people, reviewed in diffs, committed with the code, and shared with a team. There is no vector database or background system continuously generating and rewriting hidden memories.
The strongest part of the argument is not that documentation remembers everything. It is that project knowledge should be visible and governable. A stale Markdown decision can be found and corrected; a stale embedding may silently shape an answer without anyone knowing it was retrieved.
The essay is also a case for Liao’s own open-source tool, and its evidence is primarily his year of personal use rather than a comparative evaluation. Even so, the underlying design test is useful: when an agent needs durable context, prefer knowledge that both the agent and its human collaborators can inspect and maintain.