Index Agentica

Give an agent long-term memory

How to make an agent remember users, facts and lessons across sessions, choosing between file-based memory, a memory layer such as Mem0, a temporal knowledge graph such as Zep or Graphiti, a stateful agent platform such as Letta, or your own store.

Type
Guide
Author
Agentica Author
Published
Last verified
Difficulty
intermediate
Time
25 min

Prerequisites

What "memory" means here

Every agent already has short-term memory: the messages in the current context window, plus whatever state your framework checkpoints for the current thread. This guide is about long-term memory, the information that survives the end of a conversation and comes back in a later one, possibly in a different thread, on a different machine or with a different model.

The LangGraph docs give a useful split borrowed from psychology (and the CoALA paper):

TypeWhat is storedAgent example
SemanticFacts"The user's company is on AWS and prefers Terraform."
EpisodicExperiences"Last time, the migration failed because the staging DB was read-only."
ProceduralInstructionsAn updated system prompt or rules the agent has learned.

Most products called "agent memory" focus on semantic memory about users. If what you actually need is the agent getting better at a task, you want episodic or procedural memory, and the design looks different (few-shot examples, an evolving instruction file, a skills folder).

Decision 1: who writes memories, and when

There are two patterns, and LangGraph's docs name them well:

Many systems do both. Letta agents edit their own memory during work and also run "dreaming": background subagents that review recent conversations and consolidate lessons, triggered after a number of steps or on context compaction.

Decision 2: how memories are stored and found

ApproachHow retrieval worksGood atWatch out for
Files the agent reads and writesThe agent lists and opens files itselfProcedural memory, project notes, transparency (you can read and edit the files)Grows without bound unless the agent prunes it; no ranking
Vector store of extracted factsEmbedding similarity, often plus keyword search"What do I know that's relevant to this message?" over many small factsContradictions pile up; similar is not the same as current
Temporal knowledge graphGraph traversal plus semantic and keyword search, with timeFacts that change ("works at X" until March), relationships between entitiesMore moving parts: a graph database and an LLM extraction step
Stateful agent platformThe platform decides what is in context and what is paged inLong-lived agents with identity and self-edited memoryYou adopt the platform's runtime, not just a library

Option A: file-based memory (Claude memory tool)

The simplest long-term memory is a directory of notes. Anthropic's memory tool formalizes this: you add {"type": "memory_20250818", "name": "memory"} to tools, and Claude issues file commands (view, create, str_replace, insert, delete, rename) against a /memories directory. The tool is client-side, so your application executes each command against storage you control and returns the result. The docs say it is available on all Claude 4 and later models.

The Python and TypeScript SDKs include a local filesystem implementation and a tool runner that handles the loop:

import anthropic
from anthropic.tools import BetaLocalFilesystemMemoryTool

client = anthropic.Anthropic()
memory = BetaLocalFilesystemMemoryTool(base_path="./memory")

runner = client.beta.messages.tool_runner(
    model="claude-opus-5-5",  # any Claude 4+ model
    max_tokens=1024,
    messages=[{"role": "user", "content": "Remember that Acme Corp prefers email follow-ups."}],
    tools=[memory],
)
print(runner.until_done().content)

If you write your own handler (for example, to store memories in a database per user), the docs are explicit that you must validate every path: a request for /memories/../../secrets.env must be rejected. Treat each user's memory directory as a separate namespace.

This pattern fits procedural and episodic memory especially well: the agent can keep a lessons.md or a project log and read it at the start of each task. Letta's MemFS takes the same idea further, keeping each agent's memory as Markdown files in a git repository: files under system/ are always in the prompt, and the rest are listed as a tree the agent reads on demand.

Option B: a memory layer (Mem0)

Mem0 (Apache-2.0) sits beside your agent: you pass it conversations, it uses an LLM to extract memories, and you search them before each response. It comes as a Python and npm library, a self-hosted server (docker compose up) and a managed platform. The open-source library defaults to OpenAI models for extraction and embeddings, and you can swap in other providers.

from mem0 import Memory

memory = Memory()  # defaults: OpenAI LLM and embeddings; configure others as needed

# after a turn: extract and store memories scoped to this user
memory.add(
    [{"role": "user", "content": "I'm vegetarian and allergic to nuts."},
     {"role": "assistant", "content": "Got it, I'll keep that in mind."}],
    user_id="alice",
)

# before the next response: retrieve what's relevant
hits = memory.search(query="What should I cook for Alice?", filters={"user_id": "alice"}, top_k=3)
context = "\n".join(f"- {h['memory']}" for h in hits["results"])

Recent change: Mem0 shipped a new memory algorithm in April 2026. Extraction is now a single ADD-only pass (memories accumulate; nothing is updated or deleted in that step), retrieval fuses semantic, BM25 keyword and entity matching, and there is time-aware ranking. Upgrading from OSS v2 has a migration guide. Mem0 publishes benchmark scores for the new algorithm (for example 92.5 on LoCoMo and 94.4 on LongMemEval), but notes they are for its managed platform, which includes optimizations not in the open-source library. Treat them as vendor numbers and evaluate on your own conversations.

Because extraction is append-only, plan for how stale facts lose out: rely on the time-aware ranking, store timestamps, and give users a way to see and delete what was stored.

Option C: a temporal knowledge graph (Zep and Graphiti)

When facts change over time, plain vector memory struggles: "lives in Berlin" and "moved to Stockholm" are both similar to "where does she live?". Graphiti (Apache-2.0) builds a temporal context graph: entities, relationships with validity windows, and the raw "episodes" every fact came from. When new information contradicts an old fact, the old one is invalidated rather than deleted, so you can ask what is true now or what was true at a given time. Retrieval combines embeddings, BM25 and graph traversal.

Graphiti needs Python 3.10+, an LLM (OpenAI by default; it works best with providers that support structured output) and a graph database: Neo4j 5.26, FalkorDB, or Amazon Neptune with OpenSearch Serverless. Kuzu support is deprecated because the upstream project is no longer maintained. The repo also contains an MCP server.

Zep is the managed service built on the same ideas. It now describes itself as a unified context layer that combines business data, documents and conversations into temporal context graphs, with SDKs for Python, TypeScript and Go. Pick Zep when you want the graph without running a graph database; pick Graphiti when you want to self-host.

Option D: a stateful agent platform (Letta)

Letta (formerly MemGPT) treats memory as part of the agent itself: the agent's memory persists across conversations and follows it between models and computers, and the agent edits it as it learns. You define starting memory as labeled blocks at creation, configure dreaming, and run agents through the Letta Harness, the App Server, the desktop app or the TypeScript Agent SDK, locally or on Letta Cloud.

Recent change: the original letta-ai/letta Python server is retired. Its README points to letta-ai/letta-code as the current source, and the V1 API server lives on an archive branch. Older tutorials that pip install letta and call the V1 REST API describe the retired server.

Option E: build it on your own store

If you already run Postgres, pgvector adds vector columns and HNSW or IVFFlat indexes with cosine, L2, inner product and L1 distance, so memories can live next to your users table with normal row-level permissions. A dedicated vector database such as Qdrant makes sense at higher scale. Framework stores work too: the LangGraph store saves memories as JSON documents under a namespace (for example (user_id, "preferences")) and key, with optional semantic search:

from langgraph.store.memory import InMemoryStore  # use a DB-backed store in production

store = InMemoryStore()
namespace = ("user-123", "preferences")
store.put(namespace, "style", {"rules": ["Prefers short answers", "Writes Python"]})
item = store.get(namespace, "style")

You then write the extraction prompt, the deduplication and the retrieval policy yourself. That's more work, but every decision is visible and testable.

Doing it responsibly

Long-term memory turns an agent into a system that stores personal data, and it adds a new attack surface.

Which one to start with

Directory entries in this guide

Related

Sources

Machine-readable