{
  "type": "guide",
  "id": "give-an-agent-long-term-memory",
  "title": "Give an agent long-term memory",
  "summary": "How to make an agent remember users, facts and lessons across sessions, choosing between file-based memory, a memory layer such as Mem0, a temporal knowledge graph such as Zep or Graphiti, a stateful agent platform such as Letta, or your own store.",
  "author": "Agentica Author",
  "tags": [
    "memory",
    "agents",
    "personalization",
    "knowledge-graphs",
    "vector-search"
  ],
  "published": "2026-10-02",
  "last_verified": "2026-10-02",
  "difficulty": "intermediate",
  "time_estimate": "25 min",
  "entries": [
    "mem0",
    "zep",
    "graphiti",
    "letta",
    "langgraph",
    "cognee",
    "supermemory",
    "pgvector",
    "qdrant"
  ],
  "links": {
    "html": "https://indexagentica.com/guides/give-an-agent-long-term-memory/",
    "markdown": "https://indexagentica.com/guides/give-an-agent-long-term-memory.md",
    "json": "https://indexagentica.com/api/longform/guides/give-an-agent-long-term-memory.json",
    "source": "https://github.com/Drudley/indexagentica/blob/main/content-long/guides/give-an-agent-long-term-memory.md"
  },
  "status": "published",
  "prerequisites": [
    "An agent loop you control (any SDK or framework)",
    "An LLM API key; Mem0 and Graphiti default to OpenAI models for extraction and embeddings"
  ],
  "entries_detail": [
    {
      "id": "mem0",
      "name": "Mem0",
      "summary": "Memory layer for AI agents and apps: drop-in persistent memory that lets agents learn from past interactions.",
      "url": "https://indexagentica.com/entries/mem0/",
      "json": "https://indexagentica.com/api/entries/mem0.json"
    },
    {
      "id": "zep",
      "name": "Zep",
      "summary": "Unified context layer that combines business data, documents and conversations into governed context and memory for agents.",
      "url": "https://indexagentica.com/entries/zep/",
      "json": "https://indexagentica.com/api/entries/zep.json"
    },
    {
      "id": "graphiti",
      "name": "Graphiti",
      "summary": "Zep's open-source framework for building real-time knowledge graphs for AI agents.",
      "url": "https://indexagentica.com/entries/graphiti/",
      "json": "https://indexagentica.com/api/entries/graphiti.json"
    },
    {
      "id": "letta",
      "name": "Letta",
      "summary": "Platform for stateful agents with memory, identity and the ability to learn over time (formerly MemGPT); terminal UI, App Server, desktop app and Agent SDK.",
      "url": "https://indexagentica.com/entries/letta/",
      "json": "https://indexagentica.com/api/entries/letta.json"
    },
    {
      "id": "langgraph",
      "name": "LangGraph",
      "summary": "Low-level orchestration framework from LangChain for building resilient, stateful agents as graphs.",
      "url": "https://indexagentica.com/entries/langgraph/",
      "json": "https://indexagentica.com/api/entries/langgraph.json"
    },
    {
      "id": "cognee",
      "name": "Cognee",
      "summary": "Open-source AI memory platform for agents that turns documents and conversations into graph, vector and relational memory.",
      "url": "https://indexagentica.com/entries/cognee/",
      "json": "https://indexagentica.com/api/entries/cognee.json"
    },
    {
      "id": "supermemory",
      "name": "Supermemory",
      "summary": "Memory and context engine for AI agents with a Memory API; can run fully locally or hosted.",
      "url": "https://indexagentica.com/entries/supermemory/",
      "json": "https://indexagentica.com/api/entries/supermemory.json"
    },
    {
      "id": "pgvector",
      "name": "pgvector",
      "summary": "Open-source vector similarity search extension for Postgres.",
      "url": "https://indexagentica.com/entries/pgvector/",
      "json": "https://indexagentica.com/api/entries/pgvector.json"
    },
    {
      "id": "qdrant",
      "name": "Qdrant",
      "summary": "Open-source vector search engine and database written in Rust, self-hosted or as Qdrant Cloud.",
      "url": "https://indexagentica.com/entries/qdrant/",
      "json": "https://indexagentica.com/api/entries/qdrant.json"
    }
  ],
  "related": [
    {
      "type": "stack",
      "id": "research-agent",
      "title": "Research agent stack",
      "url": "https://indexagentica.com/stacks/research-agent/",
      "json": "https://indexagentica.com/api/longform/stacks/research-agent.json"
    },
    {
      "type": "stack",
      "id": "coding-agent-starter",
      "title": "Coding agent starter stack",
      "url": "https://indexagentica.com/stacks/coding-agent-starter/",
      "json": "https://indexagentica.com/api/longform/stacks/coding-agent-starter.json"
    },
    {
      "type": "guide",
      "id": "connect-an-agent-to-a-remote-mcp-server",
      "title": "Connect an agent to a remote MCP server",
      "url": "https://indexagentica.com/guides/connect-an-agent-to-a-remote-mcp-server/",
      "json": "https://indexagentica.com/api/longform/guides/connect-an-agent-to-a-remote-mcp-server.json"
    }
  ],
  "sources": [
    {
      "title": "LangChain docs, LangGraph memory overview",
      "url": "https://docs.langchain.com/oss/python/concepts/memory",
      "accessed": "2026-10-02"
    },
    {
      "title": "Claude docs, Memory tool",
      "url": "https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool",
      "accessed": "2026-10-02"
    },
    {
      "title": "Mem0 README",
      "url": "https://github.com/mem0ai/mem0",
      "accessed": "2026-10-02"
    },
    {
      "title": "Graphiti README",
      "url": "https://github.com/getzep/graphiti",
      "accessed": "2026-10-02"
    },
    {
      "title": "Zep documentation",
      "url": "https://help.getzep.com/",
      "accessed": "2026-10-02"
    },
    {
      "title": "Zep: A Temporal Knowledge Graph Architecture for Agent Memory (arXiv 2501.13956)",
      "url": "https://arxiv.org/abs/2501.13956",
      "accessed": "2026-10-02"
    },
    {
      "title": "Letta docs, Memory",
      "url": "https://docs.letta.com/agent-sdk/memory/",
      "accessed": "2026-10-02"
    },
    {
      "title": "Letta README (current source moved to letta-ai/letta-code)",
      "url": "https://github.com/letta-ai/letta",
      "accessed": "2026-10-02"
    },
    {
      "title": "pgvector README",
      "url": "https://github.com/pgvector/pgvector",
      "accessed": "2026-10-02"
    }
  ],
  "front_matter": {
    "id": "give-an-agent-long-term-memory",
    "type": "guide",
    "title": "Give an agent long-term memory",
    "summary": "How to make an agent remember users, facts and lessons across sessions, choosing between file-based memory, a memory layer such as Mem0, a temporal knowledge graph such as Zep or Graphiti, a stateful agent platform such as Letta, or your own store.",
    "description": "Long-term memory is not one feature but a set of design decisions: what kind of memory you need (facts, experiences or instructions), who writes it (the agent on the hot path or a background job), how it is scoped and searched, and how it is kept correct and private over time. This guide walks through those decisions and maps them to current tools, with verified, minimal examples for the Claude memory tool, Mem0 and the LangGraph store.",
    "author": "Agentica Author",
    "difficulty": "intermediate",
    "time_estimate": "25 min",
    "prerequisites": [
      "An agent loop you control (any SDK or framework)",
      "An LLM API key; Mem0 and Graphiti default to OpenAI models for extraction and embeddings"
    ],
    "tags": [
      "memory",
      "agents",
      "personalization",
      "knowledge-graphs",
      "vector-search"
    ],
    "entries": [
      "mem0",
      "zep",
      "graphiti",
      "letta",
      "langgraph",
      "cognee",
      "supermemory",
      "pgvector",
      "qdrant"
    ],
    "sources": [
      {
        "title": "LangChain docs, LangGraph memory overview",
        "url": "https://docs.langchain.com/oss/python/concepts/memory",
        "accessed": "2026-10-02"
      },
      {
        "title": "Claude docs, Memory tool",
        "url": "https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool",
        "accessed": "2026-10-02"
      },
      {
        "title": "Mem0 README",
        "url": "https://github.com/mem0ai/mem0",
        "accessed": "2026-10-02"
      },
      {
        "title": "Graphiti README",
        "url": "https://github.com/getzep/graphiti",
        "accessed": "2026-10-02"
      },
      {
        "title": "Zep documentation",
        "url": "https://help.getzep.com/",
        "accessed": "2026-10-02"
      },
      {
        "title": "Zep: A Temporal Knowledge Graph Architecture for Agent Memory (arXiv 2501.13956)",
        "url": "https://arxiv.org/abs/2501.13956",
        "accessed": "2026-10-02"
      },
      {
        "title": "Letta docs, Memory",
        "url": "https://docs.letta.com/agent-sdk/memory/",
        "accessed": "2026-10-02"
      },
      {
        "title": "Letta README (current source moved to letta-ai/letta-code)",
        "url": "https://github.com/letta-ai/letta",
        "accessed": "2026-10-02"
      },
      {
        "title": "pgvector README",
        "url": "https://github.com/pgvector/pgvector",
        "accessed": "2026-10-02"
      }
    ],
    "related": [
      "research-agent",
      "coding-agent-starter",
      "connect-an-agent-to-a-remote-mcp-server"
    ],
    "last_verified": "2026-10-02",
    "published": "2026-10-02"
  },
  "markdown": "\n## What \"memory\" means here\n\nEvery agent already has **short-term memory**: the messages in the current context window, plus whatever state your framework checkpoints for the current thread. This guide is about **long-term memory**, the information that survives the end of a conversation and comes back in a later one, possibly in a different thread, on a different machine or with a different model.\n\nThe LangGraph docs give a useful split borrowed from psychology (and the CoALA paper):\n\n| Type | What is stored | Agent example |\n|---|---|---|\n| Semantic | Facts | \"The user's company is on AWS and prefers Terraform.\" |\n| Episodic | Experiences | \"Last time, the migration failed because the staging DB was read-only.\" |\n| Procedural | Instructions | An updated system prompt or rules the agent has learned. |\n\nMost products called \"agent memory\" focus on semantic memory about users. If what you actually need is the agent getting better at a task, you want episodic or procedural memory, and the design looks different (few-shot examples, an evolving instruction file, a skills folder).\n\n## Decision 1: who writes memories, and when\n\nThere are two patterns, and LangGraph's docs name them well:\n\n- **On the hot path.** The agent decides, mid-conversation, to save something, usually through a tool call (`remember(...)`, a file write). Memories are available immediately and the agent can tell the user what it saved. The cost is latency and tokens on every turn, and the agent has to judge what's worth keeping while it is also doing the task.\n- **In the background.** A separate process reads finished conversations (or batches of them) and extracts, merges and rewrites memories. The main agent stays fast and focused, and the extractor can use a different model and prompt. The cost is that new memories aren't available until the job runs.\n\nMany systems do both. [Letta](https://indexagentica.com/entries/letta/) agents edit their own memory during work and also run \"dreaming\": background subagents that review recent conversations and consolidate lessons, triggered after a number of steps or on context compaction.\n\n## Decision 2: how memories are stored and found\n\n| Approach | How retrieval works | Good at | Watch out for |\n|---|---|---|---|\n| Files the agent reads and writes | The agent lists and opens files itself | Procedural memory, project notes, transparency (you can read and edit the files) | Grows without bound unless the agent prunes it; no ranking |\n| Vector store of extracted facts | Embedding similarity, often plus keyword search | \"What do I know that's relevant to this message?\" over many small facts | Contradictions pile up; similar is not the same as current |\n| Temporal knowledge graph | Graph traversal plus semantic and keyword search, with time | Facts that change (\"works at X\" until March), relationships between entities | More moving parts: a graph database and an LLM extraction step |\n| Stateful agent platform | The platform decides what is in context and what is paged in | Long-lived agents with identity and self-edited memory | You adopt the platform's runtime, not just a library |\n\n## Option A: file-based memory (Claude memory tool)\n\nThe simplest long-term memory is a directory of notes. Anthropic's **memory tool** formalizes this: you add `{\"type\": \"memory_20250818\", \"name\": \"memory\"}` to `tools`, and Claude issues file commands (`view`, `create`, `str_replace`, `insert`, `delete`, `rename`) against a `/memories` directory. The tool is **client-side**, so your application executes each command against storage you control and returns the result. The docs say it is available on all Claude 4 and later models.\n\nThe Python and TypeScript SDKs include a local filesystem implementation and a tool runner that handles the loop:\n\n```python\nimport anthropic\nfrom anthropic.tools import BetaLocalFilesystemMemoryTool\n\nclient = anthropic.Anthropic()\nmemory = BetaLocalFilesystemMemoryTool(base_path=\"./memory\")\n\nrunner = client.beta.messages.tool_runner(\n    model=\"claude-opus-5-5\",  # any Claude 4+ model\n    max_tokens=1024,\n    messages=[{\"role\": \"user\", \"content\": \"Remember that Acme Corp prefers email follow-ups.\"}],\n    tools=[memory],\n)\nprint(runner.until_done().content)\n```\n\nIf you write your own handler (for example, to store memories in a database per user), the docs are explicit that **you must validate every path**: a request for `/memories/../../secrets.env` must be rejected. Treat each user's memory directory as a separate namespace.\n\nThis pattern fits procedural and episodic memory especially well: the agent can keep a `lessons.md` or a project log and read it at the start of each task. Letta's MemFS takes the same idea further, keeping each agent's memory as Markdown files in a git repository: files under `system/` are always in the prompt, and the rest are listed as a tree the agent reads on demand.\n\n## Option B: a memory layer (Mem0)\n\n[Mem0](https://indexagentica.com/entries/mem0/) (Apache-2.0) sits beside your agent: you pass it conversations, it uses an LLM to extract memories, and you search them before each response. It comes as a Python and npm library, a self-hosted server (`docker compose up`) and a managed platform. The open-source library defaults to OpenAI models for extraction and embeddings, and you can swap in other providers.\n\n```python\nfrom mem0 import Memory\n\nmemory = Memory()  # defaults: OpenAI LLM and embeddings; configure others as needed\n\n# after a turn: extract and store memories scoped to this user\nmemory.add(\n    [{\"role\": \"user\", \"content\": \"I'm vegetarian and allergic to nuts.\"},\n     {\"role\": \"assistant\", \"content\": \"Got it, I'll keep that in mind.\"}],\n    user_id=\"alice\",\n)\n\n# before the next response: retrieve what's relevant\nhits = memory.search(query=\"What should I cook for Alice?\", filters={\"user_id\": \"alice\"}, top_k=3)\ncontext = \"\\n\".join(f\"- {h['memory']}\" for h in hits[\"results\"])\n```\n\n**Recent change:** Mem0 shipped a new memory algorithm in April 2026. Extraction is now a single ADD-only pass (memories accumulate; nothing is updated or deleted in that step), retrieval fuses semantic, BM25 keyword and entity matching, and there is time-aware ranking. Upgrading from OSS v2 has a migration guide. Mem0 publishes benchmark scores for the new algorithm (for example 92.5 on LoCoMo and 94.4 on LongMemEval), but notes they are for its managed platform, which includes optimizations not in the open-source library. Treat them as vendor numbers and evaluate on your own conversations.\n\nBecause extraction is append-only, plan for how stale facts lose out: rely on the time-aware ranking, store timestamps, and give users a way to see and delete what was stored.\n\n## Option C: a temporal knowledge graph (Zep and Graphiti)\n\nWhen facts change over time, plain vector memory struggles: \"lives in Berlin\" and \"moved to Stockholm\" are both similar to \"where does she live?\". [Graphiti](https://indexagentica.com/entries/graphiti/) (Apache-2.0) builds a **temporal context graph**: entities, relationships with validity windows, and the raw \"episodes\" every fact came from. When new information contradicts an old fact, the old one is invalidated rather than deleted, so you can ask what is true now or what was true at a given time. Retrieval combines embeddings, BM25 and graph traversal.\n\nGraphiti needs Python 3.10+, an LLM (OpenAI by default; it works best with providers that support structured output) and a graph database: Neo4j 5.26, FalkorDB, or Amazon Neptune with OpenSearch Serverless. Kuzu support is deprecated because the upstream project is no longer maintained. The repo also contains an MCP server.\n\n[Zep](https://indexagentica.com/entries/zep/) is the managed service built on the same ideas. It now describes itself as a unified context layer that combines business data, documents and conversations into temporal context graphs, with SDKs for Python, TypeScript and Go. Pick Zep when you want the graph without running a graph database; pick Graphiti when you want to self-host.\n\n## Option D: a stateful agent platform (Letta)\n\n[Letta](https://indexagentica.com/entries/letta/) (formerly MemGPT) treats memory as part of the agent itself: the agent's memory persists across conversations and follows it between models and computers, and the agent edits it as it learns. You define starting memory as labeled blocks at creation, configure dreaming, and run agents through the Letta Harness, the App Server, the desktop app or the TypeScript Agent SDK, locally or on Letta Cloud.\n\n**Recent change:** the original `letta-ai/letta` Python server is retired. Its README points to `letta-ai/letta-code` as the current source, and the V1 API server lives on an archive branch. Older tutorials that `pip install letta` and call the V1 REST API describe the retired server.\n\n## Option E: build it on your own store\n\nIf you already run Postgres, [pgvector](https://indexagentica.com/entries/pgvector/) adds vector columns and HNSW or IVFFlat indexes with cosine, L2, inner product and L1 distance, so memories can live next to your users table with normal row-level permissions. A dedicated vector database such as [Qdrant](https://indexagentica.com/entries/qdrant/) makes sense at higher scale. Framework stores work too: the [LangGraph](https://indexagentica.com/entries/langgraph/) store saves memories as JSON documents under a namespace (for example `(user_id, \"preferences\")`) and key, with optional semantic search:\n\n```python\nfrom langgraph.store.memory import InMemoryStore  # use a DB-backed store in production\n\nstore = InMemoryStore()\nnamespace = (\"user-123\", \"preferences\")\nstore.put(namespace, \"style\", {\"rules\": [\"Prefers short answers\", \"Writes Python\"]})\nitem = store.get(namespace, \"style\")\n```\n\nYou then write the extraction prompt, the deduplication and the retrieval policy yourself. That's more work, but every decision is visible and testable.\n\n## Doing it responsibly\n\nLong-term memory turns an agent into a system that stores personal data, and it adds a new attack surface.\n\n- **Scope everything.** Key every memory by user (and by tenant). Never run a search without the user filter. Shared \"agent\" memory should hold only what is safe for every user to see.\n- **Memory poisoning is prompt injection with persistence.** If the agent saves text from web pages, emails or tool output, an attacker can plant an instruction that comes back in every future session. Save facts the user stated or confirmed, record where each memory came from (Graphiti's episodes do this), and don't promote retrieved memories to system-prompt authority.\n- **Let users see, correct and delete.** Expose what is stored, support deletion end to end (including backups and the extraction provider's retention), and expire memories you no longer need.\n- **Measure it.** Build a small set of multi-session conversations with known answers and check recall, wrong recalls and stale facts before and after every change. Published benchmark scores won't tell you how memory behaves on your users' conversations.\n\n## Which one to start with\n\n- Single-user coding or research agent that should learn your preferences and project quirks: file-based memory (the Claude memory tool, or a `memory/` folder your harness reads).\n- Consumer or support assistant that should remember facts about many users: Mem0 or a similar memory layer ([Supermemory](https://indexagentica.com/entries/supermemory/) and [Cognee](https://indexagentica.com/entries/cognee/) are other open-source options).\n- Facts that change and relationships that matter (accounts, org charts, CRM-like data): Zep or Graphiti.\n- A long-lived agent with an identity of its own: Letta.\n- Strict data-residency or audit needs: your own Postgres with pgvector, or a self-hosted Mem0 or Graphiti.\n",
  "raw": "---\nid: give-an-agent-long-term-memory\ntype: guide\ntitle: Give an agent long-term memory\nsummary: How to make an agent remember users, facts and lessons across sessions, choosing between file-based memory, a memory layer such as Mem0, a temporal knowledge graph such as Zep or Graphiti, a stateful agent platform such as Letta, or your own store.\ndescription: \"Long-term memory is not one feature but a set of design decisions: what kind of memory you need (facts, experiences or instructions), who writes it (the agent on the hot path or a background job), how it is scoped and searched, and how it is kept correct and private over time. This guide walks through those decisions and maps them to current tools, with verified, minimal examples for the Claude memory tool, Mem0 and the LangGraph store.\"\nauthor: Agentica Author\ndifficulty: intermediate\ntime_estimate: 25 min\nprerequisites:\n  - An agent loop you control (any SDK or framework)\n  - An LLM API key; Mem0 and Graphiti default to OpenAI models for extraction and embeddings\ntags: [memory, agents, personalization, knowledge-graphs, vector-search]\nentries: [mem0, zep, graphiti, letta, langgraph, cognee, supermemory, pgvector, qdrant]\nsources:\n  - title: LangChain docs, LangGraph memory overview\n    url: https://docs.langchain.com/oss/python/concepts/memory\n    accessed: 2026-10-02\n  - title: Claude docs, Memory tool\n    url: https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool\n    accessed: 2026-10-02\n  - title: Mem0 README\n    url: https://github.com/mem0ai/mem0\n    accessed: 2026-10-02\n  - title: Graphiti README\n    url: https://github.com/getzep/graphiti\n    accessed: 2026-10-02\n  - title: Zep documentation\n    url: https://help.getzep.com/\n    accessed: 2026-10-02\n  - title: \"Zep: A Temporal Knowledge Graph Architecture for Agent Memory (arXiv 2501.13956)\"\n    url: https://arxiv.org/abs/2501.13956\n    accessed: 2026-10-02\n  - title: Letta docs, Memory\n    url: https://docs.letta.com/agent-sdk/memory/\n    accessed: 2026-10-02\n  - title: Letta README (current source moved to letta-ai/letta-code)\n    url: https://github.com/letta-ai/letta\n    accessed: 2026-10-02\n  - title: pgvector README\n    url: https://github.com/pgvector/pgvector\n    accessed: 2026-10-02\nrelated: [research-agent, coding-agent-starter, connect-an-agent-to-a-remote-mcp-server]\nlast_verified: 2026-10-02\npublished: 2026-10-02\n---\n\n## What \"memory\" means here\n\nEvery agent already has **short-term memory**: the messages in the current context window, plus whatever state your framework checkpoints for the current thread. This guide is about **long-term memory**, the information that survives the end of a conversation and comes back in a later one, possibly in a different thread, on a different machine or with a different model.\n\nThe LangGraph docs give a useful split borrowed from psychology (and the CoALA paper):\n\n| Type | What is stored | Agent example |\n|---|---|---|\n| Semantic | Facts | \"The user's company is on AWS and prefers Terraform.\" |\n| Episodic | Experiences | \"Last time, the migration failed because the staging DB was read-only.\" |\n| Procedural | Instructions | An updated system prompt or rules the agent has learned. |\n\nMost products called \"agent memory\" focus on semantic memory about users. If what you actually need is the agent getting better at a task, you want episodic or procedural memory, and the design looks different (few-shot examples, an evolving instruction file, a skills folder).\n\n## Decision 1: who writes memories, and when\n\nThere are two patterns, and LangGraph's docs name them well:\n\n- **On the hot path.** The agent decides, mid-conversation, to save something, usually through a tool call (`remember(...)`, a file write). Memories are available immediately and the agent can tell the user what it saved. The cost is latency and tokens on every turn, and the agent has to judge what's worth keeping while it is also doing the task.\n- **In the background.** A separate process reads finished conversations (or batches of them) and extracts, merges and rewrites memories. The main agent stays fast and focused, and the extractor can use a different model and prompt. The cost is that new memories aren't available until the job runs.\n\nMany systems do both. [Letta](https://indexagentica.com/entries/letta/) agents edit their own memory during work and also run \"dreaming\": background subagents that review recent conversations and consolidate lessons, triggered after a number of steps or on context compaction.\n\n## Decision 2: how memories are stored and found\n\n| Approach | How retrieval works | Good at | Watch out for |\n|---|---|---|---|\n| Files the agent reads and writes | The agent lists and opens files itself | Procedural memory, project notes, transparency (you can read and edit the files) | Grows without bound unless the agent prunes it; no ranking |\n| Vector store of extracted facts | Embedding similarity, often plus keyword search | \"What do I know that's relevant to this message?\" over many small facts | Contradictions pile up; similar is not the same as current |\n| Temporal knowledge graph | Graph traversal plus semantic and keyword search, with time | Facts that change (\"works at X\" until March), relationships between entities | More moving parts: a graph database and an LLM extraction step |\n| Stateful agent platform | The platform decides what is in context and what is paged in | Long-lived agents with identity and self-edited memory | You adopt the platform's runtime, not just a library |\n\n## Option A: file-based memory (Claude memory tool)\n\nThe simplest long-term memory is a directory of notes. Anthropic's **memory tool** formalizes this: you add `{\"type\": \"memory_20250818\", \"name\": \"memory\"}` to `tools`, and Claude issues file commands (`view`, `create`, `str_replace`, `insert`, `delete`, `rename`) against a `/memories` directory. The tool is **client-side**, so your application executes each command against storage you control and returns the result. The docs say it is available on all Claude 4 and later models.\n\nThe Python and TypeScript SDKs include a local filesystem implementation and a tool runner that handles the loop:\n\n```python\nimport anthropic\nfrom anthropic.tools import BetaLocalFilesystemMemoryTool\n\nclient = anthropic.Anthropic()\nmemory = BetaLocalFilesystemMemoryTool(base_path=\"./memory\")\n\nrunner = client.beta.messages.tool_runner(\n    model=\"claude-opus-5-5\",  # any Claude 4+ model\n    max_tokens=1024,\n    messages=[{\"role\": \"user\", \"content\": \"Remember that Acme Corp prefers email follow-ups.\"}],\n    tools=[memory],\n)\nprint(runner.until_done().content)\n```\n\nIf you write your own handler (for example, to store memories in a database per user), the docs are explicit that **you must validate every path**: a request for `/memories/../../secrets.env` must be rejected. Treat each user's memory directory as a separate namespace.\n\nThis pattern fits procedural and episodic memory especially well: the agent can keep a `lessons.md` or a project log and read it at the start of each task. Letta's MemFS takes the same idea further, keeping each agent's memory as Markdown files in a git repository: files under `system/` are always in the prompt, and the rest are listed as a tree the agent reads on demand.\n\n## Option B: a memory layer (Mem0)\n\n[Mem0](https://indexagentica.com/entries/mem0/) (Apache-2.0) sits beside your agent: you pass it conversations, it uses an LLM to extract memories, and you search them before each response. It comes as a Python and npm library, a self-hosted server (`docker compose up`) and a managed platform. The open-source library defaults to OpenAI models for extraction and embeddings, and you can swap in other providers.\n\n```python\nfrom mem0 import Memory\n\nmemory = Memory()  # defaults: OpenAI LLM and embeddings; configure others as needed\n\n# after a turn: extract and store memories scoped to this user\nmemory.add(\n    [{\"role\": \"user\", \"content\": \"I'm vegetarian and allergic to nuts.\"},\n     {\"role\": \"assistant\", \"content\": \"Got it, I'll keep that in mind.\"}],\n    user_id=\"alice\",\n)\n\n# before the next response: retrieve what's relevant\nhits = memory.search(query=\"What should I cook for Alice?\", filters={\"user_id\": \"alice\"}, top_k=3)\ncontext = \"\\n\".join(f\"- {h['memory']}\" for h in hits[\"results\"])\n```\n\n**Recent change:** Mem0 shipped a new memory algorithm in April 2026. Extraction is now a single ADD-only pass (memories accumulate; nothing is updated or deleted in that step), retrieval fuses semantic, BM25 keyword and entity matching, and there is time-aware ranking. Upgrading from OSS v2 has a migration guide. Mem0 publishes benchmark scores for the new algorithm (for example 92.5 on LoCoMo and 94.4 on LongMemEval), but notes they are for its managed platform, which includes optimizations not in the open-source library. Treat them as vendor numbers and evaluate on your own conversations.\n\nBecause extraction is append-only, plan for how stale facts lose out: rely on the time-aware ranking, store timestamps, and give users a way to see and delete what was stored.\n\n## Option C: a temporal knowledge graph (Zep and Graphiti)\n\nWhen facts change over time, plain vector memory struggles: \"lives in Berlin\" and \"moved to Stockholm\" are both similar to \"where does she live?\". [Graphiti](https://indexagentica.com/entries/graphiti/) (Apache-2.0) builds a **temporal context graph**: entities, relationships with validity windows, and the raw \"episodes\" every fact came from. When new information contradicts an old fact, the old one is invalidated rather than deleted, so you can ask what is true now or what was true at a given time. Retrieval combines embeddings, BM25 and graph traversal.\n\nGraphiti needs Python 3.10+, an LLM (OpenAI by default; it works best with providers that support structured output) and a graph database: Neo4j 5.26, FalkorDB, or Amazon Neptune with OpenSearch Serverless. Kuzu support is deprecated because the upstream project is no longer maintained. The repo also contains an MCP server.\n\n[Zep](https://indexagentica.com/entries/zep/) is the managed service built on the same ideas. It now describes itself as a unified context layer that combines business data, documents and conversations into temporal context graphs, with SDKs for Python, TypeScript and Go. Pick Zep when you want the graph without running a graph database; pick Graphiti when you want to self-host.\n\n## Option D: a stateful agent platform (Letta)\n\n[Letta](https://indexagentica.com/entries/letta/) (formerly MemGPT) treats memory as part of the agent itself: the agent's memory persists across conversations and follows it between models and computers, and the agent edits it as it learns. You define starting memory as labeled blocks at creation, configure dreaming, and run agents through the Letta Harness, the App Server, the desktop app or the TypeScript Agent SDK, locally or on Letta Cloud.\n\n**Recent change:** the original `letta-ai/letta` Python server is retired. Its README points to `letta-ai/letta-code` as the current source, and the V1 API server lives on an archive branch. Older tutorials that `pip install letta` and call the V1 REST API describe the retired server.\n\n## Option E: build it on your own store\n\nIf you already run Postgres, [pgvector](https://indexagentica.com/entries/pgvector/) adds vector columns and HNSW or IVFFlat indexes with cosine, L2, inner product and L1 distance, so memories can live next to your users table with normal row-level permissions. A dedicated vector database such as [Qdrant](https://indexagentica.com/entries/qdrant/) makes sense at higher scale. Framework stores work too: the [LangGraph](https://indexagentica.com/entries/langgraph/) store saves memories as JSON documents under a namespace (for example `(user_id, \"preferences\")`) and key, with optional semantic search:\n\n```python\nfrom langgraph.store.memory import InMemoryStore  # use a DB-backed store in production\n\nstore = InMemoryStore()\nnamespace = (\"user-123\", \"preferences\")\nstore.put(namespace, \"style\", {\"rules\": [\"Prefers short answers\", \"Writes Python\"]})\nitem = store.get(namespace, \"style\")\n```\n\nYou then write the extraction prompt, the deduplication and the retrieval policy yourself. That's more work, but every decision is visible and testable.\n\n## Doing it responsibly\n\nLong-term memory turns an agent into a system that stores personal data, and it adds a new attack surface.\n\n- **Scope everything.** Key every memory by user (and by tenant). Never run a search without the user filter. Shared \"agent\" memory should hold only what is safe for every user to see.\n- **Memory poisoning is prompt injection with persistence.** If the agent saves text from web pages, emails or tool output, an attacker can plant an instruction that comes back in every future session. Save facts the user stated or confirmed, record where each memory came from (Graphiti's episodes do this), and don't promote retrieved memories to system-prompt authority.\n- **Let users see, correct and delete.** Expose what is stored, support deletion end to end (including backups and the extraction provider's retention), and expire memories you no longer need.\n- **Measure it.** Build a small set of multi-session conversations with known answers and check recall, wrong recalls and stale facts before and after every change. Published benchmark scores won't tell you how memory behaves on your users' conversations.\n\n## Which one to start with\n\n- Single-user coding or research agent that should learn your preferences and project quirks: file-based memory (the Claude memory tool, or a `memory/` folder your harness reads).\n- Consumer or support assistant that should remember facts about many users: Mem0 or a similar memory layer ([Supermemory](https://indexagentica.com/entries/supermemory/) and [Cognee](https://indexagentica.com/entries/cognee/) are other open-source options).\n- Facts that change and relationships that matter (accounts, org charts, CRM-like data): Zep or Graphiti.\n- A long-lived agent with an identity of its own: Letta.\n- Strict data-residency or audit needs: your own Postgres with pgvector, or a self-hosted Mem0 or Graphiti.\n",
  "html": "<h2>What &quot;memory&quot; means here</h2>\n<p>Every agent already has <strong>short-term memory</strong>: the messages in the current context window, plus whatever state your framework checkpoints for the current thread. This guide is about <strong>long-term memory</strong>, the information that survives the end of a conversation and comes back in a later one, possibly in a different thread, on a different machine or with a different model.</p>\n<p>The LangGraph docs give a useful split borrowed from psychology (and the CoALA paper):</p>\n<table><thead><tr><th scope=\"col\">Type</th><th scope=\"col\">What is stored</th><th scope=\"col\">Agent example</th></tr></thead><tbody><tr><td>Semantic</td><td>Facts</td><td>&quot;The user&#39;s company is on AWS and prefers Terraform.&quot;</td></tr><tr><td>Episodic</td><td>Experiences</td><td>&quot;Last time, the migration failed because the staging DB was read-only.&quot;</td></tr><tr><td>Procedural</td><td>Instructions</td><td>An updated system prompt or rules the agent has learned.</td></tr></tbody></table>\n<p>Most products called &quot;agent memory&quot; focus on semantic memory about users. If what you actually need is the agent getting better at a task, you want episodic or procedural memory, and the design looks different (few-shot examples, an evolving instruction file, a skills folder).</p>\n<h2>Decision 1: who writes memories, and when</h2>\n<p>There are two patterns, and LangGraph&#39;s docs name them well:</p>\n<ul><li><strong>On the hot path.</strong> The agent decides, mid-conversation, to save something, usually through a tool call (<code>remember(...)</code>, a file write). Memories are available immediately and the agent can tell the user what it saved. The cost is latency and tokens on every turn, and the agent has to judge what&#39;s worth keeping while it is also doing the task.</li><li><strong>In the background.</strong> A separate process reads finished conversations (or batches of them) and extracts, merges and rewrites memories. The main agent stays fast and focused, and the extractor can use a different model and prompt. The cost is that new memories aren&#39;t available until the job runs.</li></ul>\n<p>Many systems do both. <a href=\"/entries/letta/\">Letta</a> agents edit their own memory during work and also run &quot;dreaming&quot;: background subagents that review recent conversations and consolidate lessons, triggered after a number of steps or on context compaction.</p>\n<h2>Decision 2: how memories are stored and found</h2>\n<table><thead><tr><th scope=\"col\">Approach</th><th scope=\"col\">How retrieval works</th><th scope=\"col\">Good at</th><th scope=\"col\">Watch out for</th></tr></thead><tbody><tr><td>Files the agent reads and writes</td><td>The agent lists and opens files itself</td><td>Procedural memory, project notes, transparency (you can read and edit the files)</td><td>Grows without bound unless the agent prunes it; no ranking</td></tr><tr><td>Vector store of extracted facts</td><td>Embedding similarity, often plus keyword search</td><td>&quot;What do I know that&#39;s relevant to this message?&quot; over many small facts</td><td>Contradictions pile up; similar is not the same as current</td></tr><tr><td>Temporal knowledge graph</td><td>Graph traversal plus semantic and keyword search, with time</td><td>Facts that change (&quot;works at X&quot; until March), relationships between entities</td><td>More moving parts: a graph database and an LLM extraction step</td></tr><tr><td>Stateful agent platform</td><td>The platform decides what is in context and what is paged in</td><td>Long-lived agents with identity and self-edited memory</td><td>You adopt the platform&#39;s runtime, not just a library</td></tr></tbody></table>\n<h2>Option A: file-based memory (Claude memory tool)</h2>\n<p>The simplest long-term memory is a directory of notes. Anthropic&#39;s <strong>memory tool</strong> formalizes this: you add <code>{&quot;type&quot;: &quot;memory_20250818&quot;, &quot;name&quot;: &quot;memory&quot;}</code> to <code>tools</code>, and Claude issues file commands (<code>view</code>, <code>create</code>, <code>str_replace</code>, <code>insert</code>, <code>delete</code>, <code>rename</code>) against a <code>/memories</code> directory. The tool is <strong>client-side</strong>, so your application executes each command against storage you control and returns the result. The docs say it is available on all Claude 4 and later models.</p>\n<p>The Python and TypeScript SDKs include a local filesystem implementation and a tool runner that handles the loop:</p>\n<pre><code class=\"language-python\">import anthropic\nfrom anthropic.tools import BetaLocalFilesystemMemoryTool\n\nclient = anthropic.Anthropic()\nmemory = BetaLocalFilesystemMemoryTool(base_path=&quot;./memory&quot;)\n\nrunner = client.beta.messages.tool_runner(\n    model=&quot;claude-opus-5-5&quot;,  # any Claude 4+ model\n    max_tokens=1024,\n    messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;Remember that Acme Corp prefers email follow-ups.&quot;}],\n    tools=[memory],\n)\nprint(runner.until_done().content)</code></pre>\n<p>If you write your own handler (for example, to store memories in a database per user), the docs are explicit that <strong>you must validate every path</strong>: a request for <code>/memories/../../secrets.env</code> must be rejected. Treat each user&#39;s memory directory as a separate namespace.</p>\n<p>This pattern fits procedural and episodic memory especially well: the agent can keep a <code>lessons.md</code> or a project log and read it at the start of each task. Letta&#39;s MemFS takes the same idea further, keeping each agent&#39;s memory as Markdown files in a git repository: files under <code>system/</code> are always in the prompt, and the rest are listed as a tree the agent reads on demand.</p>\n<h2>Option B: a memory layer (Mem0)</h2>\n<p><a href=\"/entries/mem0/\">Mem0</a> (Apache-2.0) sits beside your agent: you pass it conversations, it uses an LLM to extract memories, and you search them before each response. It comes as a Python and npm library, a self-hosted server (<code>docker compose up</code>) and a managed platform. The open-source library defaults to OpenAI models for extraction and embeddings, and you can swap in other providers.</p>\n<pre><code class=\"language-python\">from mem0 import Memory\n\nmemory = Memory()  # defaults: OpenAI LLM and embeddings; configure others as needed\n\n# after a turn: extract and store memories scoped to this user\nmemory.add(\n    [{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;I&#39;m vegetarian and allergic to nuts.&quot;},\n     {&quot;role&quot;: &quot;assistant&quot;, &quot;content&quot;: &quot;Got it, I&#39;ll keep that in mind.&quot;}],\n    user_id=&quot;alice&quot;,\n)\n\n# before the next response: retrieve what&#39;s relevant\nhits = memory.search(query=&quot;What should I cook for Alice?&quot;, filters={&quot;user_id&quot;: &quot;alice&quot;}, top_k=3)\ncontext = &quot;\\n&quot;.join(f&quot;- {h[&#39;memory&#39;]}&quot; for h in hits[&quot;results&quot;])</code></pre>\n<p><strong>Recent change:</strong> Mem0 shipped a new memory algorithm in April 2026. Extraction is now a single ADD-only pass (memories accumulate; nothing is updated or deleted in that step), retrieval fuses semantic, BM25 keyword and entity matching, and there is time-aware ranking. Upgrading from OSS v2 has a migration guide. Mem0 publishes benchmark scores for the new algorithm (for example 92.5 on LoCoMo and 94.4 on LongMemEval), but notes they are for its managed platform, which includes optimizations not in the open-source library. Treat them as vendor numbers and evaluate on your own conversations.</p>\n<p>Because extraction is append-only, plan for how stale facts lose out: rely on the time-aware ranking, store timestamps, and give users a way to see and delete what was stored.</p>\n<h2>Option C: a temporal knowledge graph (Zep and Graphiti)</h2>\n<p>When facts change over time, plain vector memory struggles: &quot;lives in Berlin&quot; and &quot;moved to Stockholm&quot; are both similar to &quot;where does she live?&quot;. <a href=\"/entries/graphiti/\">Graphiti</a> (Apache-2.0) builds a <strong>temporal context graph</strong>: entities, relationships with validity windows, and the raw &quot;episodes&quot; every fact came from. When new information contradicts an old fact, the old one is invalidated rather than deleted, so you can ask what is true now or what was true at a given time. Retrieval combines embeddings, BM25 and graph traversal.</p>\n<p>Graphiti needs Python 3.10+, an LLM (OpenAI by default; it works best with providers that support structured output) and a graph database: Neo4j 5.26, FalkorDB, or Amazon Neptune with OpenSearch Serverless. Kuzu support is deprecated because the upstream project is no longer maintained. The repo also contains an MCP server.</p>\n<p><a href=\"/entries/zep/\">Zep</a> is the managed service built on the same ideas. It now describes itself as a unified context layer that combines business data, documents and conversations into temporal context graphs, with SDKs for Python, TypeScript and Go. Pick Zep when you want the graph without running a graph database; pick Graphiti when you want to self-host.</p>\n<h2>Option D: a stateful agent platform (Letta)</h2>\n<p><a href=\"/entries/letta/\">Letta</a> (formerly MemGPT) treats memory as part of the agent itself: the agent&#39;s memory persists across conversations and follows it between models and computers, and the agent edits it as it learns. You define starting memory as labeled blocks at creation, configure dreaming, and run agents through the Letta Harness, the App Server, the desktop app or the TypeScript Agent SDK, locally or on Letta Cloud.</p>\n<p><strong>Recent change:</strong> the original <code>letta-ai/letta</code> Python server is retired. Its README points to <code>letta-ai/letta-code</code> as the current source, and the V1 API server lives on an archive branch. Older tutorials that <code>pip install letta</code> and call the V1 REST API describe the retired server.</p>\n<h2>Option E: build it on your own store</h2>\n<p>If you already run Postgres, <a href=\"/entries/pgvector/\">pgvector</a> adds vector columns and HNSW or IVFFlat indexes with cosine, L2, inner product and L1 distance, so memories can live next to your users table with normal row-level permissions. A dedicated vector database such as <a href=\"/entries/qdrant/\">Qdrant</a> makes sense at higher scale. Framework stores work too: the <a href=\"/entries/langgraph/\">LangGraph</a> store saves memories as JSON documents under a namespace (for example <code>(user_id, &quot;preferences&quot;)</code>) and key, with optional semantic search:</p>\n<pre><code class=\"language-python\">from langgraph.store.memory import InMemoryStore  # use a DB-backed store in production\n\nstore = InMemoryStore()\nnamespace = (&quot;user-123&quot;, &quot;preferences&quot;)\nstore.put(namespace, &quot;style&quot;, {&quot;rules&quot;: [&quot;Prefers short answers&quot;, &quot;Writes Python&quot;]})\nitem = store.get(namespace, &quot;style&quot;)</code></pre>\n<p>You then write the extraction prompt, the deduplication and the retrieval policy yourself. That&#39;s more work, but every decision is visible and testable.</p>\n<h2>Doing it responsibly</h2>\n<p>Long-term memory turns an agent into a system that stores personal data, and it adds a new attack surface.</p>\n<ul><li><strong>Scope everything.</strong> Key every memory by user (and by tenant). Never run a search without the user filter. Shared &quot;agent&quot; memory should hold only what is safe for every user to see.</li><li><strong>Memory poisoning is prompt injection with persistence.</strong> If the agent saves text from web pages, emails or tool output, an attacker can plant an instruction that comes back in every future session. Save facts the user stated or confirmed, record where each memory came from (Graphiti&#39;s episodes do this), and don&#39;t promote retrieved memories to system-prompt authority.</li><li><strong>Let users see, correct and delete.</strong> Expose what is stored, support deletion end to end (including backups and the extraction provider&#39;s retention), and expire memories you no longer need.</li><li><strong>Measure it.</strong> Build a small set of multi-session conversations with known answers and check recall, wrong recalls and stale facts before and after every change. Published benchmark scores won&#39;t tell you how memory behaves on your users&#39; conversations.</li></ul>\n<h2>Which one to start with</h2>\n<ul><li>Single-user coding or research agent that should learn your preferences and project quirks: file-based memory (the Claude memory tool, or a <code>memory/</code> folder your harness reads).</li><li>Consumer or support assistant that should remember facts about many users: Mem0 or a similar memory layer (<a href=\"/entries/supermemory/\">Supermemory</a> and <a href=\"/entries/cognee/\">Cognee</a> are other open-source options).</li><li>Facts that change and relationships that matter (accounts, org charts, CRM-like data): Zep or Graphiti.</li><li>A long-lived agent with an identity of its own: Letta.</li><li>Strict data-residency or audit needs: your own Postgres with pgvector, or a self-hosted Mem0 or Graphiti.</li></ul>"
}
