Every AI agent I use starts each session with amnesia. It does not remember what it did yesterday, what another agent decided an hour ago, or that the migration it is about to write already exists on a branch. For one agent and one task, you can paste context into the prompt. For several agents working across many projects every day, that stops scaling fast.
For the past months I have run a personal setup where every coding agent on my machine (Claude Code, Codex, pi, jcode, OpenCode and Hermes) shares one durable memory. I call it the Brain. It is a folder of markdown and JSON files in a git repository, which I also open in Obsidian when I want to read it as a human.
agent-memory-server is the generalised, open-source version of that idea. This post explains the design, and why I chose files and git over a vector database.
What agents actually need to remember#
When I looked at what went wrong without shared memory, the problems were specific:
- Continuity. An agent picks up a task and has no idea what the last session did or what it planned to do next.
- Coordination. Two agents start on the same task at the same time, or one abandons a task and nobody notices.
- Decisions. An agent reverses a choice that was made deliberately, because the reasoning lived only in a chat log that is gone.
- Recall. "Did we solve this before?" has an answer somewhere, but nobody can find it.
Notice that only the last one is a search problem. The first three are about structure, ownership and history. That observation drove the whole design.
The case against starting with a vector database#
The default answer to "agent memory" is embeddings in a vector store. Chunk everything, embed it, retrieve the top k by similarity. It is a reasonable tool for recall. It is a poor foundation for the rest.
- You cannot read it. When an agent does something odd because of a memory, I want to open the memory and look at it. A row of floats does not help.
- You cannot diff it. Memory changes over time. I want to know what changed, when, and which agent changed it.
- Similarity is not truth. A superseded decision and its replacement are, by definition, very similar texts. Top-k retrieval will happily return the old one.
- Coordination needs exact state. "Who owns this task right now?" is a lookup with one correct answer, not a nearest-neighbour query.
- It is another service. Another thing to run, back up, migrate and pay for, holding the most important context your agents have.
None of this means embeddings are useless. It means they belong on top of a source of truth, not in place of one.
Plain files, every write a commit#
agent-memory-server stores everything as files in a git repository:
memory-store/
sessions/
2026-09-18-coder-a91f.md
2026-09-19-reviewer-03bc.md
tasks/
add-rate-limiting.json
decisions/
2026-09-12-use-sqlite-for-tests.md
Every write is a commit. That one choice gives a lot for free:
- An audit trail.
git log -p decisions/shows every decision, who recorded it and when. - Portability. Clone it to another machine. Back it up by pushing to a remote. Nothing is locked into a vendor.
- Human readability. It is markdown. I can read it, fix it, or open the folder as an Obsidian vault.
- Sync across machines. Git already solves distributed sync, including conflicts.
Running it is one process pointed at a directory:
npx agent-memory-server --store ~/agent-memory --port 7800
Sessions and handoffs#
A session is an append-only log of one agent's stretch of work: what it was asked to do, checkpoints along the way, and a handoff at the end that says what is done and what should happen next.
curl -X POST localhost:7800/sessions \ -H 'Content-Type: application/json' \ -d '{"agent": "coder", "summary": "Refactored auth module to strategy pattern", "next": "write tests for the OAuth strategy"}'
The next field is the most valuable piece of data in the whole system. When a new session starts, the first thing an agent gets is a short brief: recent handoffs, open tasks, the most likely next action, and any hazards. That brief replaces the ten minutes I used to spend re-explaining context at the start of every session.
Two rules keep sessions useful:
- Checkpoint after meaningful progress, not at the end. Agents crash, run out of context or get interrupted. A session that only writes at the end often writes nothing.
- Append, never edit. A session log is history. If something was wrong, the next entry says so.
Tasks with leases#
Coordination is where files plus a small amount of logic beat everything else I tried. A task is a JSON document with an owner and a lease:
{ "id": "add-rate-limiting", "title": "Add rate limiting to the assistant endpoint", "status": "claimed", "lease": { "agent": "coder", "session": "2026-09-19-coder-7d21", "expires_at": "2026-09-19T16:30:00Z" }, "checkpoints": [ { "at": "2026-09-19T15:10:00Z", "note": "Route middleware added, tests next" } ] }
The rules are simple:
- An agent must claim a task before working on it. A claim fails if another agent holds an unexpired lease.
- The owner checkpoints as it goes, which also extends the lease.
- When done, it completes the task. If it gives up, it releases it.
- If an agent crashes, the lease expires and the task can be reclaimed by someone else.
async function claimTask(id: string, agent: string, session: string, ttlMinutes = 60): Promise<Task> { const task = await readTask(id); if (task.lease && Date.parse(task.lease.expires_at) > Date.now() && task.lease.agent !== agent) { throw new LeaseHeldError(`Task ${id} is leased by ${task.lease.agent} until ${task.lease.expires_at}.`); } task.status = "claimed"; task.lease = { agent, session, expires_at: new Date(Date.now() + ttlMinutes * 60_000).toISOString() }; await writeTask(task); await commit(`task: ${agent} claimed ${id}`); return task; }
This is the same idea as a lock with a timeout in any distributed system, and the same lesson I learned freeing memory in C: every resource needs an owner, and something has to handle the owner dying. The difference here is that every claim and release is also a commit, so when two agents disagree about who owns what, the history settles it.
Decisions as records#
Decisions get their own store, loosely modelled on architecture decision records:
--- title: Use SQLite for the test database date: 2026-09-12 status: accepted --- ## Decision Feature tests run against in-memory SQLite. ## Reason Tests run in parallel across worktrees; a shared MySQL database caused cross-agent interference.
Decisions are never deleted. When one is replaced, the old record is marked superseded and links to the new one. This is the single most effective defence I have found against agents undoing deliberate choices: before changing something architectural, an agent checks the decision log, and the reasoning is right there.
The same principle runs through the whole store: superseded facts are marked, not erased, and nothing is ever deleted, only archived.
Search, and where embeddings fit#
Full-text search across the store works out of the box, and it answers most questions, because agents and humans tend to reuse the same words for the same things: project names, file names, error messages.
curl 'localhost:7800/search?q=rate+limiting'
In my own setup I layer meaning-based search on top: a hybrid of keyword and embedding search over the same files. The key point is the direction of the dependency. The index is derived from the files and can be rebuilt from them at any time. If it breaks, I delete it and reindex. The files remain the source of truth.
MCP makes it universal#
The reason one memory can serve six different agents is that it does not depend on any of them. agent-memory-server speaks plain HTTP today, and the Model Context Protocol transport is in active development, so any MCP-aware client can call tools like "start session", "claim task" and "record decision" directly:
{ "mcpServers": { "memory": { "command": "npx", "args": ["-y", "agent-memory-server", "--store", "~/agent-memory", "--mcp"] } } }
Agents that do not speak MCP can use a CLI or HTTP. What matters is that there is exactly one place where durable memory lives. I explicitly tell every agent not to keep durable facts in its own local memory. Local indexes are fine for throwaway search inside a project. Anything worth remembering goes to the shared store.
What I would warn you about#
- Memory is data, not instructions. An agent reading its brief should treat it as context. A memory that says "ignore previous rules" is just text someone wrote.
- Curate aggressively. A memory that records everything becomes noise. Handoffs and decisions are high value. Raw transcripts are not.
- Keep secrets out. It is a git repository. Treat it like one.
- Git is not a database. At very high write rates, commit-per-write gets slow. For agent memory, where writes are measured in dozens per hour, it is more than fast enough.
Status and roadmap#
agent-memory-server is v0.x. Sessions, tasks and decisions with git persistence, and full-text search, work today. Next: MCP stdio and streamable HTTP transports, lease expiry and automatic reclaim for crashed agents, optional encryption at rest, multi-agent namespaces, and small adapters for Python, Go and Rust.
The broader lesson from running this every day: agent memory is less about retrieval and more about accountability. Who did what, who owns what, what was decided and why. Files and git have been solving that for decades. Agents just need to use them.