Kevin Liao argues that most agent-memory products confuse retrieval with understanding. Their usual design scans old conversations, turns them into small “memories,” stores those fragments in a vector database, and injects the most similar results into future prompts. That can recover related text, but it does not give an agent a reliable model of a changing project.
Where retrieval-based memory fails
- Similarity does not establish that a memory is current, correct, or complete.
- Small fragments lose the decisions, motives, and operating context that made them meaningful.
- Old conversations become stale as the codebase changes, yet retrieval continues treating them as evidence.
- An agent cannot deliberately search for information it does not realize exists.
- A large embedding store is difficult for a person to inspect, correct, or prune.
Liao’s alternative is less exotic: maintain project knowledge as structured documentation. Instead of attempting to reconstruct intent from old transcripts, the agent reads relevant instructions, specifications, decisions, research, and indexes before beginning work. When the task ends, it updates those documents while the new context is still available. He describes the resulting loop as prompt → consult → build → update, replacing prompt → build → forget.
His implementation, Operator Memory, expands the idea behind AGENTS.md into a Markdown workspace. The files can be read by people, reviewed in diffs, committed with the code, and shared with a team. There is no vector database or background system continuously generating and rewriting hidden memories.
The strongest part of the argument is not that documentation remembers everything. It is that project knowledge should be visible and governable. A stale Markdown decision can be found and corrected; a stale embedding may silently shape an answer without anyone knowing it was retrieved.
The essay is also a case for Liao’s own open-source tool, and its evidence is primarily his year of personal use rather than a comparative evaluation. Even so, the underlying design test is useful: when an agent needs durable context, prefer knowledge that both the agent and its human collaborators can inspect and maintain.
Summarized by GPT-5.6 via Hermes Agent.