--- title: "Supermemory local" sidebarTitle: "Overview" description: "State-of-the-art memory, running on your machine. One binary, zero config." icon: "/icons/hugeicons/server-stack-01.svg" --- Supermemory runs on your own hardware. It's the same memory engine behind the [hosted platform](https://console.supermemory.ai) — ingestion, memory extraction, hybrid semantic search, and the full API — as a single self-contained binary. ```bash curl curl -fsSL https://supermemory.ai/install | bash ``` ```bash npx npx supermemory local ``` No Docker. No database to provision. No config files. It boots in seconds with everything built in. The SDKs and other components in the [public repository](https://git.new/memory) are open source; the downloadable self-hosted server binary is built from a separate, non-public codebase. ## Zero config, actually Run the binary with nothing set and you get a complete memory system: - **The Supermemory graph engine, embedded** — created automatically on first boot. No database to stand up, no connection strings. - **Built-in local embeddings** — default `Xenova/bge-base-en-v1.5` (768d) on your machine, no API key. Same provider stack as cloud if you opt into OpenAI, Gemini, or Ollama — see [Embeddings](/self-hosting/embeddings). - **An API key, generated for you** — printed on first boot, ready to paste into any SDK. - **The full Memory API** — `/ns/{namespace}/document`, `/ns/{namespace}/search`, `/ns/{namespace}/profile`, namespaces, the works. The only thing you bring is a model. In production, Supermemory runs its own proprietary models, purpose-tuned for long-horizon data understanding and memory extraction. Self-hosted, the same pipeline runs on whatever model you point it at — OpenAI, Anthropic, Gemini, Groq, or any OpenAI-compatible endpoint. Bring a key and go. Or don't bring one at all: ## Runs fully offline Supermemory works with any OpenAI-compatible endpoint, which means it runs end-to-end on your machine with a local model — Ollama, LM Studio, vLLM, llama.cpp. `gpt-oss-20b` is a great fit: ```bash OPENAI_BASE_URL=http://localhost:11434/v1 \ OPENAI_API_KEY=ollama \ OPENAI_MODEL=gpt-oss:20b \ supermemory-server ``` Local graph engine, local embeddings, local LLM. Text-memory processing can stay on your machine after the embedding model's first download. Supermemory collects telemetry; you can disable it with `SUPERMEMORY_DISABLE_TELEMETRY=1`. URL ingestion uses a hosted reader service; use local text or file inputs for offline operation. ## Drop-in with your existing code The v5 SDKs (`supermemory` 5.x on npm and PyPI) need `supermemory-server` v0.0.9 or later. Older servers only speak v3/v4; run `supermemory-server upgrade` first. The self-hosted server speaks the same API as the hosted platform. Point any Supermemory SDK at it with a one-line change: ```typescript const client = new Supermemory({ apiKey: "sm_...", // printed on first boot baseUrl: "http://localhost:6767", }) ``` Everything in the [Memory API docs](/quickstart) works the same way. The coding plugins do too — [Claude Code](/integrations/claude-code), [Muse Code](/integrations/muse-code), [Codex](/integrations/codex), and [OpenCode](/integrations/opencode) all target your local server with `SUPERMEMORY_API_URL=http://localhost:6767` (Muse: set `baseUrl` in `.muse/supermemory.json`, because hook env is cleared). ## Self-hosted vs. the platform Self-hosted is free within its lite license limit and useful for local development and privacy-sensitive workloads. The server binary is not open source; the hosted platform is where the full product lives: | | Self-hosted | Platform | |---|---|---| | Full Memory API | ✅ | ✅ | | Hybrid semantic search | ✅ | ✅ | | Embeddings | Local default (or OpenAI / Gemini / Ollama) | Same provider stack, managed | | File ingestion (PDFs, images) | ✅ | ✅ | | [Connectors](/connectors/overview) (Google Drive, Notion, Gmail, OneDrive) | — | ✅ | | [Supermemory MCP](/supermemory-mcp/mcp) | — | ✅ | | Memory extraction | Your model, your key | Proprietary long-horizon models — higher quality, cheaper at scale | | Infrastructure | Your machine | Globally distributed, scales with you | If you outgrow a single machine — or want connectors, MCP, and the best-tuned extraction pipeline — [the platform](https://console.supermemory.ai) is one `serverURL` change away. Running this for a team or organization? See [Local vs. Enterprise](/self-hosting/local-vs-enterprise). ## Next steps Install, run, and store your first memory in under two minutes Every environment variable: LLM providers, storage, auth, tuning Local default, remote providers, multilingual, dimension lock