New paper · arXiv 2609.36059

Agents forget.Mnemon remembers.

Local, persistent memory for AI agents. The LLM you already use decides what to keep; Mnemon stores, links and recalls it without calling another model.

npm install -g @mnemon-dev/mnemon && mnemon setup

Node 22+ · macOS / Linux / Windows

$mnemon recall "why did we choose Qdrant" --intent WHY
  • temporal
  • entity
  • semantic
  • causal

Core · mnemon

Stop throwing every state problem at another LLM.

Most memory systems call another model on every write. Mnemon leaves judgment to the LLM you already run and does the deterministic work itself: storage, indexing, search and decay.

# record a decision
$ mnemon remember "Chose Qdrant over Milvus for vector search" --cat decision --imp 5 --entities "Qdrant,Milvus"

Store what the LLM decides should outlive the session. Exact repeats are skipped.

Who decides what to remember

Extra inference
  • LLM-embeddedMem0, LettaEvery write
  • File injectionCLAUDE.md and similarNone, but context fills up
  • Memory serverMCP toolsNone
  • LLM-supervisedMnemonNone
  • No API keys
  • One binary, SQLite storage
  • Deterministic and testable
  • Shared across agents

16

Supported agents

Run mnemon setup to detect and connect the agents on your machine.

  • Claude Code
  • Codex
  • Cursor
  • OpenCode
  • OpenClaw
  • Trae
  • Qoder
  • QoderWork
  • CodeBuddy
  • WorkBuddy
  • Kimi Code
  • Hermes
  • Pi
  • MiniMax Code
  • Nanobot
  • ZCode
  • DeepSeek HarnessThrough dsh-mnemon
  • NanoClawThrough its own skill
Compare approaches
ApproachDecidesExtra inferenceRuns as
LLM-embeddedMem0, LettaA second modelEvery writeUsually a service
File injectionCLAUDE.md and similarNobody; all loadedNone, but context fills upFiles
Memory serverMCP toolsThe agent, via toolsNoneA server per host
LLM-supervisedMnemonYour host LLMNoneLocal binary + SQLite

Ecosystem · dsh-mnemon

Layered memory for DeepSeek Harness

Frequently used facts stay in context; project documents and long-term memory are searched when needed.

  • Runtime memory

    Always loaded

    USER.md and MEMORY.md

  • Project documents

    On demand

    Markdown with revisions

  • Memory spaces

    On demand

    Mnemon Native or a third-party provider

    • Mnemon Nativedefault
    • OpenViking
    • Honcho
dsh plugin --profile web add dsh-mnemon
localhost:3080

Documents and memory recalled in a reply

Research · arXiv 2609.36059

Raw Records, Fast Judgments, Slow Thoughts

Conversations are kept as raw records and judged at question time: a fast decision model screens records, an LLM plans and answers.

  • Raw records

    verbatim · BM25 + vectors

  • Fast judgments

    Jev · System 1 · ≤ 16 records

  • Slow thoughts

    LLM · System 2 · < 4k tokens

Accuracy vs. contextResearch system; it does not use the mnemon binary. Comparison data from OmniMemEval, which uses a different grader (about 1–2 points apart).

All systems answer with gpt-4.1-mini

  • Mnemon
  • Other systems
10%20%30%40%50%60%70%80%90%100%1,0002,0005,00010K20K50K100KContext per question (tokens, log scale)AccuracyMnemon 91.7%
  • 91.7%LoCoMohighest of 15 systems (gpt-4.1-mini answering)
  • 3.8kcontext per questiontokens
  • 94.4%LongMemEval-Swith a reasoning model answering
  • ×1.11cost per questionfrom 100K to 10M tokens of history (BEAM)

Adoption

Estimated from public npm, npmmirror and GitHub data

Active installsInstalls that took a recent update, pooled over the latest complete update waves. The range is an 80% interval.

    

Estimated usersAfter discounting CI, containers and multiple machines per person. The range is an 80% interval.

    

New installs / dayMedian daily first-time installs over the last 7 measurable days.

    

Update half-lifeTime for half of active installs to take a new release.

    

FAQ

What is Mnemon?

An open-source memory layer for AI agents. A single binary keeps memories in SQLite on your machine; the agent you already use, such as Claude Code, Codex or Cursor, decides what to remember and when to recall it.

How is it different from Mem0 or Letta?

Those systems call a second model on every write. Mnemon calls no model of its own: your agent's LLM makes the judgment calls, and Mnemon does the deterministic work of storing, linking, searching and decaying memories.

Which agents does it work with?

mnemon setup connects 16 agents, including Claude Code, Codex, Cursor, OpenCode, OpenClaw, Trae, Qoder, QoderWork, CodeBuddy, WorkBuddy. DeepSeek Harness uses it through the dsh-mnemon plugin.

Does it need an API key? Where is my data stored?

No API key or account is needed. Memories are stored in a SQLite file on your machine.

How do I install it?

Run npm install -g @mnemon-dev/mnemon && mnemon setup. Homebrew and Go installs work too, and setup connects the agents it finds on your machine.

What is dsh-mnemon?

The memory plugin for DeepSeek Harness. Frequently used facts stay in context; project documents and long-term memory are searched when needed, from Mnemon or a third-party provider.

What does the Mnemon paper show?

The Mnemon memory agent keeps conversations as raw records and judges them at question time. It reaches 91.7% on LoCoMo, the highest of 15 systems with gpt-4.1-mini answering, with under 4k tokens of context per question. It is a research system, separate from the mnemon binary.

Thanks

To everyone who has contributed code or opened an issue.

Contributors–

Issue authors–

Get started

One binary, one command, no API keys.

npm install -g @mnemon-dev/mnemon && mnemon setup

Node 22+ · macOS / Linux / Windows