An agent that closes its session forgets everything it learned in it. agent-memory is the runtime that fixes that, for any agent — not only coding ones. Markdown files in one store are the single source of truth, the SQLite index beside them is a cache you can delete at any time, and Claude Code, Codex CLI, Muse Code, and anything else that can run a shell command share that store.
Retrieval is local and ranked, and it answers with paths rather than pasted text — the agent opens each hit only as deep as the task needs. Writes do not wait for the agent to remember to make them: they fire at conversation boundaries. A sleep-time pass then consolidates and forgets by value, on its own clock. None of it needs an API key.
Agent memory has grown along two architectural lines. One builds a retrieval engine —
embeddings, a knowledge graph, a ranking pipeline — which finds the right thing, but hands the
agent an opaque chunk it cannot inspect and a store it cannot migrate off. The other hands the
agent a filesystem — markdown it reads directly, browsable with ls and grep — which is
legible and costs nothing to run, but does not rank, and stops scaling the moment the tree
outgrows a listing.
agent-memory is the two of them in one store: the retrieval engine indexes a filesystem the agent can also just read. Relations live as links inside the memories, a local index ranks them, and every hit resolves to a whole markdown file on disk. Recall gains the precision of a graph and a vector search without giving up a plain directory an agent can walk — and it stays fast, because nothing in the read path calls a model or crosses a network.
Recall does not paste text into your context. It answers with an L0 list — one-line abstract, file path, anchor, score — and the agent opens what it wants at the depth the task needs:
mem recall "why files instead of a database" # L0 list, 8 entries by default
mem recall "why files instead of a database" --limit 20 # more Memory candidates
mem read <name> --level outline # headings only; or abstract, or full
mem context "why files instead of a database" # both in one call, top few expanded in full
mem trace <name> # cited raw messages, when neededIndex line → abstract → full file → raw material: each rung costs an order of magnitude more than the last, and each is a place to stop. Long files add two free rungs — the anchor that matched, and an outline computed at read time.
- Three read tracks, so a miss on one is not a miss. Deterministic
MEMORY.mdinjection at session start; BM25 recall over an FTS5 index, with a vector plugin fused in by RRF when you want one; and the plain directory tree, reachable withlsandgrepwhen both fail. Same-directory memories are a free neighbourhood, andlinksin the frontmatter carry the graph without a graph database under them. - Write coverage is the system's job, not the agent's judgement. Distillation is triggered at boundaries and runs without holding up the task; the full trace is copied first, so "missed by the distiller" never means "lost by the system".
- Files are the truth; every index is a rebuildable cache.
rm -rf .index/ && mem rebuildloses zero knowledge — enforced by a test, not promised in a doc. Your memory stays greppable, git-able, and portable off this system. - A real Manage layer, on its own clock. Sleep-time consolidation with authority tiers: an
unattended pass may add and update, deletion only ever arrives as a proposal you confirm.
Every competitor either has no M, or buries it in the write path. Supersede leaves the chain
intact and
recall --as-ofanswers as of a date, so updating never destroys. - No LLM client inside the library. Zero keys to install and no billing surface: judgement is borrowed from the host agent's own CLI, which keeps every write visible in your transcript.
$AGENT_MEMORY_STORE/
├── MEMORY.md root index, one line per memory — the only resident injection
├── config.toml every tunable; an unknown knob is refused at load
├── schemas/ one file per type: its key fields, the field it groups by, write mode
├── decision/ memories live at <type>/<group>/<name>.md, placed by the schema
│ └── agent-memory/ …/markdown-files-are-the-single-source-of-truth.md
├── archive/ append-only, out of the retrieval surface by default
│ ├── provenance/ distillation evidence, kept forever
│ └── sessions/ full trace copies, in case the host prunes its own
├── dream-reports/ one per sleep: what moved, what was proposed, evidence pointers
├── .index/ fully rebuildable: content-hash manifest, FTS5, access log
└── .state/ runtime state that is not: distillation watermark, write lock
One memory is one file. valid_from and optional invalid_at define its validity interval;
replaced and deleted files stay in the store for recall --as-of and trace. Current recall,
MEMORY.md, BM25, and Vector use files without invalid_at. Frontmatter carries the stable
name, a one-sentence abstract, the type and its schema fields, timestamps, links, weight,
and provenance; the body is free markdown. Existing files with status load without a reset.
Explicit links must name distinct active memories in the same store. correct --link
replaces the full list; correct --clear-links removes every link. MCP memory_correct
uses links: [...] and links: [] for the same operations. Omitting links preserves
historical relationships during unrelated correction. See the
operation boundary design.
Agents can use mem correct <name> to revise a memory, mem record --supersedes <old>
to create a successor, mem supersede <old> <new> to use an existing successor,
mem merge <first> <second> --abstract ... --body ... to combine memories atomically,
and mem delete <name> to end current validity. Splitting uses new record calls
followed by delete of the original. Core checks file scope, relationships, provenance,
and concurrent writes; Sleep keeps its proposal and cap limits for unattended actions.
Measured on LongMemEval-S with a bounded haystack, 120 episodes, claude -p (Haiku 4.5) as
host, one calibrated Sonnet 5 judge, two exam replays per arm.
| arm | pooled accuracy | paired vs agent-memory |
|---|---|---|
| agent-memory W2 | 127/240 = 52.9% | — |
| MemCore W2 | 86/240 = 35.8% | +37/−17, p=0.009 · +35/−14, p=0.004 |
| no memory | 7/120 = 5.8% | +61/−4 · +60/−4, p<0.001 |
Absolute numbers are not comparable to published LongMemEval scores — the haystack is bounded to 12 sessions per episode, which makes this a write-strategy study rather than a corpus-size one. The system-to-system row differs in write and read together, so it is an end-to-end comparison and licenses no attribution to either half.
One store, three hosts: all 9 ordered writer/reader pairs across Claude Code, Codex CLI, and Hermes pass — what one host's shell writes, another's finds, specifics intact. Pooled net contribution over no memory: 2/36 → 13/36, p=0.0074.
The protocol that decides whether a measurement counts as a result, the full ledger, and the
raw run records live in docs/experiments.md and experiments/ in the working tree. They ship
with the source, not with git history.
Requires Python 3.12 or higher and uv. There is no release on PyPI yet, so install from a checkout:
git clone https://github.com/tigerless-labs/agent-memory.git
cd agent-memory
uv sync --all-packagesThat builds mem, mem-mcp, and mem-hook into .venv/bin. Inside the checkout uv run mem
reaches them; put the directory on your PATH so your shell can too:
export PATH="$PWD/.venv/bin:$PATH"mem initThe store defaults to ~/agent-memory-store; export AGENT_MEMORY_STORE only to put it
somewhere else, and export it everywhere your agents run, not just in this shell.
Write one memory, find it again, then throw the index away and prove nothing was lost:
mem record --type decision --field project=agent-memory \
--abstract "Markdown files are the single source of truth" \
--body "Indexes are rebuildable caches."
mem --json recall "source of truth"
rm -rf ~/agent-memory-store/.index && mem rebuildmem setup --host claude-code
mem setup --host codex
mem setup --host muse-code --provider openrouterAll three commands use the same pipeline: probe the host, initialize or check the Store, safely
merge host settings, install lifecycle hooks and the skill, perform provider-specific setup,
then run preflight. A successful command ends with status: READY; a written configuration is
not by itself considered ready. Re-running setup is safe and does not duplicate managed hooks,
skills, provider routing, or MCP entries. Existing unrelated host settings are retained, and a
conflicting setting fails closed instead of being replaced.
Use a non-default Store in the usual way; setup pins its absolute path into every installed hook:
AGENT_MEMORY_STORE=/absolute/path/to/store mem setup --host codexRun the same checks later, or get a precise failure layer after setup fails:
mem doctor --host claude-code
mem doctor --host codex
mem doctor --host muse-code --provider openrouterDiagnostics distinguish host/config/Store/hook/reasoner failures and, for Muse, credential,
proxy, authentication, model-catalog, model-availability, and live-request failures. Codex also
reports disabled hooks, a read-only default sandbox, and the one-time hook trust review. Codex
user hooks remain subject to Codex's trust prompt. Setup and doctor perform a minimal live
reasoner request by default; --no-live is available for offline inspection but deliberately
reports FAILED because readiness was not proven.
When the Store is outside the Codex workspace, SessionStart injection and boundary hooks still
run, but agent-initiated CLI recall needs write access for the Store access log. Doctor reports
CODEX_STORE_ACCESS_REVIEW until the Store is covered by
sandbox_workspace_write.writable_roots, or Codex is launched with --add-dir <store>.
Setup does not silently expand sandbox permissions.
Pass --mcp to setup when the host should also receive the agent-memory stdio MCP server.
Agents that speak MCP get the same core calls through mem-mcp (memory_recall,
memory_read, memory_trace, memory_record, memory_correct, memory_supersede,
memory_merge, memory_delete, memory_feedback). Anything that can run a
shell command needs neither: the CLI is the universal fallback, and it is the wider surface —
context, sleep, and the proposal ledger have no MCP tool yet.
Codex keeps its normal MCP approval boundary; the first tool call can still require explicit
user approval even after setup has configured and verified the server.
SessionStart injects, Stop and SessionEnd distil where supported, and PreCompact evicts.
Distillation reasons through the host that fired the boundary, using its existing login. Point
[executor] in Store config.toml at a model endpoint instead to use an endpoint reasoner.
Endpoint use is opt-in and never carries a Tigerless project or credential. A failed host,
credential, request, or endpoint response is reported and leaves the archived backlog eligible
for retry. Detached boundary failures are recorded in the Store's .state/hooks.log.
To use Vertex AI, enable it in a Google Cloud project you control, grant the calling identity
prediction access, authenticate the gcloud CLI, and configure that project explicitly:
[executor]
reasoner = "endpoint"
project = "your-google-cloud-project"
location = "global"
model = "google/gemini-3.7-flash"GOOGLE_CLOUD_PROJECT and VERTEX_LOCATION override the Store values. Alternatively, an
OpenAI-compatible endpoint can be selected with endpoint and a GEMINI_API_KEY supplied in
the environment. Store configuration contains no secret, and service-account keys must not be
committed to the repository.
Install Muse Code, make sure muse is on PATH, and make
an OpenRouter credential available either as OPENROUTER_API_KEY or in Muse's credential store.
Setup recognizes an already provisioned credential and never prints or replaces it:
export OPENROUTER_API_KEY="..." # omit when Muse auth is already provisioned
mem setup --host muse-code --provider openrouter
museThe provider step configures Muse's meta transport to the local pproxy bridge, selects the
Muse model, starts or reuses pproxy, validates OpenRouter authentication and model availability,
then asks Muse to complete a minimal real request. It merges SessionStart, PreCompact,
Stop, and SessionEnd into $XDG_CONFIG_HOME/muse/settings.json (or
~/.config/muse/settings.json) without replacing unrelated settings. Managed hooks pin that
exact settings path as well as the Store and Muse data directory, so background distillation
does not depend on Muse preserving the launching shell's XDG environment.
pproxy remains a separate, pinned external dependency. Setup never silently installs system
software. If it is missing, the FAILED report prints the exact uv tool install command; run it
and repeat setup. pproxy is launched as a user process and recorded under
$XDG_STATE_HOME/agent-memory (or ~/.local/state/agent-memory). After a reboot, repeat setup
or run doctor if the proxy is no longer reachable. Use --mcp to merge the MCP server too:
mem setup --host muse-code --provider openrouter --mcpMuse native memory and agent-memory's AGENT_MEMORY_STORE are separate systems. Setup does not
read, write, copy, or synchronize Muse native memory. For attributable experiments, use a clean
workspace with no .agents/memory content and isolated HOME, XDG_CONFIG_HOME, and
XDG_DATA_HOME; the included tools/muse_sandbox_probe.py does this while reusing only the
explicit Muse auth file.
Muse's default sandbox can read outside the workspace but writes only to the workspace and temp
directories. Muse documents user hooks as outside the agent shell sandbox and MCP servers as
external processes; use the included live preflight to verify both write paths on your Muse
build before relying on the default external ~/agent-memory-store. A Muse shell command such
as mem --json recall ... can read it, but mem record from the shell cannot write it.
The experiment adapter keeps the sandbox enabled and roots its temporary Store and workdir under
one explicit experiment workspace. It never adds --yolo or --disable-sandbox.
AGENT_MEMORY_LIVE_MUSE=1 uv run python tools/muse_sandbox_probe.pySDK applications do not need a manually managed backend or the persistent pproxy process used by the interactive setup above. From the SDK application directory, run the managed bootstrap once:
export OPENROUTER_API_KEY="..." # preferably injected by the application's secret manager
mem setup --host muse-code --provider openrouter --sdkThe command initializes the selected Store, checks Muse, Node 20+, the project-local
@muse-code/sdk, the managed launcher, credential availability, and version compatibility, then
performs one real request through the same invocation-scoped bridge used by the SDK. It prints the
absolute museBin and arguments to use. It does not modify persistent Muse settings or auth, start
a persistent pproxy, edit package.json, or install system software. When a dependency is absent,
the failed check includes the exact next command; rerun bootstrap after applying it.
Run it from the application workspace, not from a workspace nested under the system temporary
directory: Muse requires its process-lifetime tool-output directory to live outside the workspace.
OPENROUTER_API_KEY must be present in the environment of both bootstrap and the application
process. In production, inject it with the application's secret manager. For a local shell, avoid
putting the key on a command line or in shell history:
read -rsp 'OpenRouter API key: ' OPENROUTER_API_KEY
echo
export OPENROUTER_API_KEY
mem setup --host muse-code --provider openrouter --sdk
node app.mjs
unset OPENROUTER_API_KEYAn existing Muse meta or openrouter credential is also recognized, so the environment variable
can be omitted in that case. Use the bootstrap result as the SDK launch contract:
import { MuseClient } from "@muse-code/sdk";
const client = await MuseClient.spawn({
museBin: "/absolute/path/printed/by/bootstrap/mem-muse",
args: ["serve"],
env: process.env,
clientInfo: { name: "my_app", version: "1.0.0" },
});
const session = await client.startSession({ workspaceRoot: process.cwd() });
// Send turns through session, then release the owned backend.
await client.close();mem-muse initializes the selected Store, creates a private invocation configuration, installs
the lifecycle hooks and skill, enables mem-mcp when that optional executable is installed,
serves the Muse model catalog through an invocation-scoped OpenRouter bridge, and starts the real
muse serve. It leaves the user's Muse settings and auth file unchanged. SDK close and startup
failure both tear down Muse, the generated credential copy, and the bridge; detached distillation
starts a fresh hook-free mem-muse exec, so it neither depends on the parent backend nor recurses.
The SDK requires Node 20 or newer. Pin @muse-code/sdk to the Muse Code CLI version because they
ship in lockstep. Bootstrap rejects major or minor version skew. Patch skew is reported as an
advisory warning and must still pass the live request. OpenRouter can reject an otherwise valid
request when the selected model is unavailable for the account or region; the Store hook log
retains that redacted upstream error and the archived session remains eligible for retry.
Start with the three ordered portability mechanics before a full four-host matrix:
mem-exp interop --workspace /tmp/muse-memory-smoke \
--pairs muse-code:muse-code,muse-code:codex,codex:muse-codeThen omit --pairs and pass --hosts claude-code,codex,hermes,muse-code for the 4×4 matrix.
Current limitations: only the root Muse session log is captured; child/observer logs are ignored,
setup installs MCP only when --mcp is requested, and live hook/MCP/sandbox behavior must be
verified on a machine with Muse Code installed and authenticated. Muse 1.4.3 with the tested
OpenRouter model may shorten a fully qualified MCP tool ID to memory_recall; Muse rejects that
shorthand even though the configured server handshake succeeds. Lifecycle hooks and SessionStart
injection do not depend on that optional model-driven MCP call. A missing Muse login is reported
as BLOCKED_BY_MUSE_AUTH; echo or mocked providers do not count as live E2E evidence.
mem sleep # host-backed consolidation; T0 applies, T1 files a proposal
mem proposals # what is waiting on you
mem decide <id> --acceptManage borrows its reasoning from the host CLI you point it at, writes a dream report for the pass, and cannot delete anything unattended.
uv run pytest -q && uv run ruff check . && uv run mypyThe task lifecycle and the invariants a change must not break are in CLAUDE.md.
BM25 is the low-latency baseline retrieval path. Install agent-memory-core[vector] (or run
uv sync --extra vector from this workspace), then set vector_enabled = true
in the store's [index] configuration. vector_model defaults to
BAAI/bge-small-en-v1.5. The first enabled Store loads FastEmbed/ONNX and may
need network access to download the model; disabled stores never load FastEmbed.
Run mem --store /path/to/store rebuild to rebuild the SQLite cache from Markdown.
The existing indexing path catches changed and deleted files, enabling vectors
on an existing store, and changes to vector_model. Recall fuses BM25 and vector
chunk candidates with reciprocal-rank fusion, then applies the existing lifecycle,
scope, as-of, weight and recency rules. Raw sessions are retained for audit and
provenance-bound trace; they do not enter Recall ranking. Recall never modifies Markdown truth.
This implements the existing optional-index design (ADR-003), using SQLite and exact cosine search. On the fixed 120-query retrieval acceptance set, optional vector fusion improved Recall@5 from 79.0% to 86.6% (+7.6 percentage points), while median retrieval latency increased from 5.1ms to 139.2ms. This establishes a retrieval-coverage/latency trade-off. Fixed-context answer replays scored 17/24 versus 18/24 and, on the expanded set, 28/36 versus 27/36. The existing answer-level experiments do not establish an end-to-end accuracy improvement, so vector retrieval remains optional.
mem --json read <name> includes a memory's provenance. To inspect a cited raw
message range, call mem --json trace <name> --pointer 'sessions/<session>#<start>-<end>'.
The pointer can select a smaller range within one citation; omitting it reads all
sources cited by the memory. Trace reports source, original message indices, roles,
times, validity, and a warning that historical content is data. Missing or unbound
evidence fails explicitly. Ordinary context and recall search Memory only; an agent
can increase --limit or reformulate its query before tracing a selected memory.
The experiment runner selects the tested host and judge independently. Pass
--host codex --judge-host codex and explicit --model / --judge-model values
for a Codex-only run. Omitting --judge-host retains the Claude Code judge and
its historical default model. calibrate and regrade also accept --judge-host.
Use calibrate --cases <labelled-cases.json> --output <calibration.json> to retain
individual votes and distinguish transport failures from label disagreements.
run --observe-reads retains bounded exam host output and CLI/read evidence in
observations/, outside store truth. Observation is off by default; missing or
truncated evidence is not proof of no tool calls. run.json fixes both host/model
pairs, configuration, source stores, code revision and episode identity. Replay
with --reuse-stores and a separate workspace for each configuration. Small panels
check execution and exploratory behavior, not a statistically established improvement.
Before scaling a read-side comparison, check that each copied store has a populated
Memory index and that a known query returns hits. Then run a small observed
agentic pilot and count successful, nonempty retrievals for each arm's intended
path (for example, vector candidates or bound Trace messages).
An enabled setting, a prompt instruction, or a tool call with zero hits does not
show that the intervention was used. Stop when the pilot does not exercise both
paths; report the exposure rate alongside scores when it does. Codex can also
read store files directly through its shell, so check the host command transcript
for bypasses before attributing an answer to a mem retrieval path.
MIT.