MemPalace/mempalace

Local-first AI memory that stores conversations verbatim, retrieves via semantic search, and offers a CLI, MCP server and pluggable backends; 96.6% R@5 on LongMemEval.

  • 59.5k GitHub stars
  • Python
  • ⚖️ MIT
  • 🎯 Intermediate
npx skills add MemPalace/mempalace
MemPalace/mempalace preview image

What it is

MemPalace stores conversation history as verbatim text and retrieves it with semantic search, without summarizing or paraphrasing. The index is organized into wings (people and projects), rooms (topics) and drawers (original content) so searches can be scoped. The retrieval backend is pluggable, with ChromaDB as the default, and nothing leaves your machine unless you opt in.

Who it's for

  • Developers using coding agents such as Claude Code, Codex CLI, Cursor IDE or Gemini CLI who want persistent conversation memory
  • Users who want local-first memory that works without API calls
  • Teams or users who want to swap in a different vector store backend (e.g. Qdrant, pgvector, Milvus)
  • Developers who want to reproduce retrieval benchmarks from the repository

Requirements

Requirements

  • Python 3.9+
  • A vector-store backend (ChromaDB by default)
  • About 300 MB disk for the embedding model (embeddinggemma-300m is recommended; all-MiniLM-L6-v2 is English-only, about 30 MB)
  • Docker is optional for the container image
  • No API key is required for the core benchmark path

Setup

  1. Agent-guided setup

    Install the MemPalace skills, then ask your coding agent to set up MemPalace. Installing a skill does not by itself install the CLI or MCP server; the setup skill guides the agent through those steps.

    bash
    npx skills add MemPalace/mempalace
  2. Direct CLI setup with uv

    Install the CLI in an isolated environment using uv (recommended), then initialize a project. pipx is also supported.

    bash
    uv tool install mempalace
    mempalace init ~/projects/myapp
  3. Plain pip inside a virtualenv

    Use pip only inside an activated virtualenv where you want import mempalace available.

    bash
    python -m venv .venv && source .venv/bin/activate
    pip install mempalace
  4. Docker image

    A multi-arch (amd64 + arm64) container image is available. Everything persists under /data, so mount a volume there.

    bash
    docker pull ghcr.io/mempalace/mempalace:latest

Examples

Mine content into the palace

bash
bash
mempalace mine ~/projects/myapp                    # project files
mempalace mine ~/.claude/projects/ --mode convos   # Claude Code sessions (scope with --wing per project)

What it does: Mines project files or Claude Code session transcripts into the palace.

Search and load context

bash
bash
mempalace search "why did we switch to GraphQL"

# Load context for a new session
mempalace wake-up

What it does: Runs a semantic search over stored content and loads context for a new session.

Run the MCP server in Docker

bash
bash
docker run -i --rm -v mempalace-data:/data ghcr.io/mempalace/mempalace

What it does: Starts the MCP server over stdio; the -i flag is needed because JSON-RPC uses stdin, and the volume persists the palace, config and embedding model.

Mine a directory via Docker

bash
bash
docker run --rm -v mempalace-data:/data -v /path/to/project:/work:ro \
  ghcr.io/mempalace/mempalace mine /work

What it does: Mounts a project read-only and mines it into the persistent data volume.

Per-message recall with sweep

bash
bash
mempalace sweep <transcript-dir>

What it does: Stores one verbatim drawer per user/assistant message; the source describes it as idempotent and resume-safe.

Pros & cons

Pros

  • Pro:Local-first: verbatim storage with no API calls required for the core retrieval path (96.6% R@5 raw on LongMemEval)
  • Pro:Pluggable backend contract with multiple options (chroma, sqlite_exact, rust_exact, milvus, qdrant, pgvector)
  • Pro:Includes a temporal knowledge graph on local SQLite and 45 MCP tools for agent integration
  • Pro:Benchmark results are reproducible from the repository, with methodology documented

Cons

  • Con:Native Termux/Android installation is not supported; it requires a Debian PRoot container workaround
  • Con:First embedding-dependent command downloads the model (about 80 MB for minilm, about 300 MB for embeddinggemma), needing network and time
  • Con:Claude Code sessions expire in 30 days unless auto-save hooks are wired, and the GPU Docker image is not published and is x86_64-only

Images