diff --git a/apps/docs/self-hosting/configuration.mdx b/apps/docs/self-hosting/configuration.mdx index 09478722..655d40d6 100644 --- a/apps/docs/self-hosting/configuration.mdx +++ b/apps/docs/self-hosting/configuration.mdx @@ -69,10 +69,11 @@ Full provider table, multilingual guidance, remote examples (OpenAI / Gemini / O | Variable | Purpose | Default | |---|---|---| -| `SUPERMEMORY_EMBEDDING_PROVIDER` | `local`, `openai`, `gemini`, or OpenAI-compatible remote | `local` | -| `SUPERMEMORY_EMBEDDING_MODEL` | Model id for the chosen provider | `Xenova/bge-base-en-v1.5` | +| `SUPERMEMORY_EMBEDDING_PROVIDER` | `local`, `openai`, `openai-compatible`, or `gemini` (use `openai-compatible` for Ollama) | `local` | +| `SUPERMEMORY_EMBEDDING_MODEL` | Model id for the chosen provider (e.g. `Xenova/bge-m3`, `text-embedding-3-small`) | `Xenova/bge-base-en-v1.5` | | `SUPERMEMORY_EMBEDDING_DIMENSIONS` | Vector size; must match model and stored data | `768` | -| `SUPERMEMORY_EMBEDDING_BASE_URL` | Base URL for OpenAI-compatible embedding APIs | unset | +| `SUPERMEMORY_EMBEDDING_BASE_URL` | Base URL for OpenAI-compatible embedding APIs (Ollama, vLLM) | unset | +| `SUPERMEMORY_EMBEDDING_API_KEY` | API key for embedding endpoint (falls back to `OPENAI_API_KEY` / `GEMINI_API_KEY`) | unset | ### Embedding performance diff --git a/apps/docs/self-hosting/embeddings.mdx b/apps/docs/self-hosting/embeddings.mdx index 7f0f6257..4d66f223 100644 --- a/apps/docs/self-hosting/embeddings.mdx +++ b/apps/docs/self-hosting/embeddings.mdx @@ -1,143 +1,170 @@ ---- -title: "Embeddings (self-hosted)" -sidebarTitle: "Embeddings" -description: "Local and remote embedding providers for Supermemory local — defaults, env vars, multilingual options, and dimension lock." -icon: "waypoints" ---- - -Self-hosted Supermemory uses the **same embedding provider stack** as the hosted platform: local ONNX models, OpenAI, Gemini, or any OpenAI-compatible embeddings endpoint (including Ollama). LLM keys power extraction and summarization; embeddings are configured separately. - -## Defaults - -| | | -|---|---| -| Provider | `local` | -| Model | `Xenova/bge-base-en-v1.5` | -| Dimensions | `768` | -| API key | None — runs on your machine | - -Press Enter at the optional first-boot picker to keep this default. Nothing is sent off-box to embed. - - -The default local model is **English-only**. Non-English content can ingest successfully while dense semantic recall stays weak. See [Multilingual](#multilingual). - - -## First-time setup (interactive) - -On first boot with a TTY, Supermemory asks for an LLM API key (required), then optionally which embedding model to use. - -1. Choose or paste an LLM provider key (OpenAI, Anthropic, Gemini, Groq, or OpenAI-compatible). -2. Optionally pick an embedding provider/model. **Press Enter to keep the local English model.** -3. Choices are saved encrypted under your data directory (`$SUPERMEMORY_DATA_DIR`, typically `./.supermemory` / `~/.supermemory`). - -Boot order is intentional: LLM keys load first so remote embedding options can reuse them (for example OpenAI or Gemini embeddings with the same key). - - -**First boot (terminal):** Supermemory asks for an LLM API key (required), then optionally which embedding model to use. Press Enter to keep the local English model. Choices are saved encrypted under your data directory. - - -## Configuration (env) - -For Docker, CI, or any non-interactive deploy, set env vars — there is **no interactive prompt without a TTY**. - -| Variable | Purpose | Default | -|---|---|---| -| `SUPERMEMORY_EMBEDDING_PROVIDER` | Embedding backend: `local`, `openai`, `gemini`, or an OpenAI-compatible remote (`ollama` / custom base URL) | `local` | -| `SUPERMEMORY_EMBEDDING_MODEL` | Model id for the chosen provider | `Xenova/bge-base-en-v1.5` (local) | -| `SUPERMEMORY_EMBEDDING_DIMENSIONS` | Vector size; must match the model and any already-stored data | `768` (local default) | -| `SUPERMEMORY_EMBEDDING_BASE_URL` | Base URL for OpenAI-compatible embedding APIs (Ollama, vLLM, etc.) | unset | -| `OPENAI_API_KEY` | Used when provider is `openai` (or compatible) if not otherwise supplied | unset | -| `GEMINI_API_KEY` | Used when provider is `gemini` | unset | - -Local worker tuning (throughput only — does not change model or dimensions): - -| Variable | Purpose | Default | -|---|---|---| -| `SUPERMEMORY_LOCAL_EMBEDDING_POOL_SIZE` | Number of embedding workers | `1` | -| `SUPERMEMORY_LOCAL_EMBEDDING_WASM_THREADS` | Compute threads per worker | `1` | -| `SUPERMEMORY_LOCAL_EMBEDDING_BATCH_SIZE` | Texts per worker dispatch | `8` | -| `SUPERMEMORY_LOCAL_EMBEDDING_IDLE_TIMEOUT_MS` | Idle time before workers shut down | `120000` | -| `SUPERMEMORY_SKIP_EMBEDDING_PREWARM` | Skip startup prewarm, load on first use | unset | - -Ingestion memory headroom is controlled by `SUPERMEMORY_EMBEDDING_RAM_LIMIT` — see [Memory limits & ingestion queue](/self-hosting/configuration#memory-limits-&-ingestion-queue). - - -**Docker / production:** Set at least one LLM key and, if you don’t want local embeddings, set `SUPERMEMORY_EMBEDDING_PROVIDER` / `SUPERMEMORY_EMBEDDING_MODEL` / `SUPERMEMORY_EMBEDDING_DIMENSIONS` (and base URL or API key as needed). There is no interactive prompt without a TTY. - - -## Multilingual - -The default `Xenova/bge-base-en-v1.5` model is trained for English. For German, Dutch, and other non-English corpora, dense recall can fail even when hybrid keyword search still finds rare tokens. - -For multilingual or non-English deployments, switch **before** large backfills: - -```bash -# Example: local multilingual (set dimensions to match the model) -SUPERMEMORY_EMBEDDING_PROVIDER=local -SUPERMEMORY_EMBEDDING_MODEL=Xenova/bge-m3 -SUPERMEMORY_EMBEDDING_DIMENSIONS=1024 -``` - -Or use a remote multilingual embedding API (OpenAI, Gemini, or Ollama with a multilingual embed model). Set provider, model, and dimensions together. Changing them later requires a fresh data directory or full re-ingestion — see below. - -## Remote providers - -### Local (default) - -```bash -# Explicit local default — no embedding API key -SUPERMEMORY_EMBEDDING_PROVIDER=local -SUPERMEMORY_EMBEDDING_MODEL=Xenova/bge-base-en-v1.5 -SUPERMEMORY_EMBEDDING_DIMENSIONS=768 -``` - -### OpenAI - -```bash -OPENAI_API_KEY=sk-... -SUPERMEMORY_EMBEDDING_PROVIDER=openai -SUPERMEMORY_EMBEDDING_MODEL=text-embedding-3-small -SUPERMEMORY_EMBEDDING_DIMENSIONS=1536 -``` - -### Gemini - -```bash -GEMINI_API_KEY=... -SUPERMEMORY_EMBEDDING_PROVIDER=gemini -SUPERMEMORY_EMBEDDING_MODEL=text-embedding-004 -SUPERMEMORY_EMBEDDING_DIMENSIONS=768 -``` - -### Ollama (OpenAI-compatible) - -```bash -SUPERMEMORY_EMBEDDING_PROVIDER=openai -SUPERMEMORY_EMBEDDING_BASE_URL=http://localhost:11434/v1 -OPENAI_API_KEY=ollama -SUPERMEMORY_EMBEDDING_MODEL=nomic-embed-text -SUPERMEMORY_EMBEDDING_DIMENSIONS=768 -``` - -Use the dimension published for your chosen model. A mismatch with vectors already in the store fails boot. - -## Changing models later - - -**Not supported in place.** Embeddings from different models (or different dimensions) are not comparable. Start from a fresh data directory or re-ingest all content so vectors stay in one space. If configured dimensions disagree with stored data, the server **refuses to boot**. - - -**Changing embeddings later:** Not supported in place. Start from a fresh data directory or re-ingest all content so vectors stay comparable. - -> [!IMPORTANT] -> **Model Mixing Bug in v0.0.5 (Exact match returns nothing)** -> -> In version `v0.0.5`, there was a bug where the server could mix different embedding models between write and read paths (e.g., document ingestion using OpenAI but memory queries using local default embeddings). In multilingual contexts like Japanese (which lacks space tokenization for fallback lexical FTS matching), this caused exact-text memory searches through `/v4/search` and `/v4/profile` to silently return `{"results":[],"total":0}`. -> -> **Resolution:** -> This was fully resolved in `v0.0.7` by locking the embedding plan uniformly across all document and query embedding paths (enforced via a locked plan in the database store). If you are running `v0.0.5` and experiencing this issue, you should upgrade to `v0.0.7` or later. - -## Related - -- [Configuration](/self-hosting/configuration) — LLM providers, storage, ingestion limits -- [Quickstart](/self-hosting/quickstart) — install and first memory +--- +title: "Embeddings (self-hosted)" +sidebarTitle: "Embeddings" +description: "Local and remote embedding providers for Supermemory local — defaults, env vars, multilingual options, and dimension lock." +icon: "waypoints" +--- + +Self-hosted Supermemory uses the **same embedding provider stack** as the hosted platform: local ONNX models, OpenAI, Gemini, or any OpenAI-compatible embeddings endpoint (including Ollama). LLM keys power extraction and summarization; embeddings are configured separately. + +## Defaults + +| | | +|---|---| +| Provider | `local` | +| Model | `Xenova/bge-base-en-v1.5` | +| Dimensions | `768` | +| API key | None — runs on your machine | + +Press Enter at the optional first-boot picker to keep this default. Nothing is sent off-box to embed. + + +The default local model is **English-only**. Non-English content can ingest successfully while dense semantic recall stays weak. See [Multilingual](#multilingual). + + +## First-time setup (interactive) + +On first boot with a TTY, Supermemory asks for an LLM API key (required), then optionally which embedding model to use. + +1. Choose or paste an LLM provider key (OpenAI, Anthropic, Gemini, Groq, or OpenAI-compatible). +2. Optionally pick an embedding provider/model. **Press Enter to keep the local English model.** +3. Choices are saved encrypted under your data directory (`$SUPERMEMORY_DATA_DIR`, typically `./.supermemory` / `~/.supermemory`). + +Boot order is intentional: LLM keys load first so remote embedding options can reuse them (for example OpenAI or Gemini embeddings with the same key). + + +**First boot (terminal):** Supermemory asks for an LLM API key (required), then optionally which embedding model to use. Press Enter to keep the local English model. Choices are saved encrypted under your data directory. + + +## Configuration (env) + +For Docker, CI, or any non-interactive deploy, set env vars — there is **no interactive prompt without a TTY**. + +| Variable | Purpose | Default | +|---|---|---| +| `SUPERMEMORY_EMBEDDING_PROVIDER` | Embedding backend: `local`, `openai`, `openai-compatible`, or `gemini` (for Ollama, use `openai-compatible` with `SUPERMEMORY_EMBEDDING_BASE_URL`) | `local` | +| `SUPERMEMORY_EMBEDDING_MODEL` | Model id for the chosen provider (e.g., `Xenova/bge-m3`, `text-embedding-3-small`) | `Xenova/bge-base-en-v1.5` (local) | +| `SUPERMEMORY_EMBEDDING_DIMENSIONS` | Vector size; must match the model's native dimensions and stored data | `768` (local default) | +| `SUPERMEMORY_EMBEDDING_BASE_URL` | Base URL for OpenAI-compatible embedding APIs (Ollama, vLLM, etc.) | unset | +| `SUPERMEMORY_EMBEDDING_API_KEY` | API key for embedding endpoint (falls back to `OPENAI_API_KEY` / `GEMINI_API_KEY`; required for `openai-compatible`, use any non-empty string if unauthenticated) | unset | +| `OPENAI_API_KEY` | Used when provider is `openai` (or compatible) if not otherwise supplied | unset | +| `GEMINI_API_KEY` | Used when provider is `gemini` | unset | + +Local worker tuning (throughput only — does not change model or dimensions): + +| Variable | Purpose | Default | +|---|---|---| +| `SUPERMEMORY_LOCAL_EMBEDDING_POOL_SIZE` | Number of embedding workers | `1` | +| `SUPERMEMORY_LOCAL_EMBEDDING_WASM_THREADS` | Compute threads per worker | `1` | +| `SUPERMEMORY_LOCAL_EMBEDDING_BATCH_SIZE` | Texts per worker dispatch | `8` | +| `SUPERMEMORY_LOCAL_EMBEDDING_IDLE_TIMEOUT_MS` | Idle time before workers shut down | `120000` | +| `SUPERMEMORY_SKIP_EMBEDDING_PREWARM` | Skip startup prewarm, load on first use | unset | + +Ingestion memory headroom is controlled by `SUPERMEMORY_EMBEDDING_RAM_LIMIT` — see [Memory limits & ingestion queue](/self-hosting/configuration#memory-limits-&-ingestion-queue). + + +**Docker / production:** Set at least one LLM key and, if you don’t want local embeddings, set `SUPERMEMORY_EMBEDDING_PROVIDER` / `SUPERMEMORY_EMBEDDING_MODEL` / `SUPERMEMORY_EMBEDDING_DIMENSIONS` (and base URL or API key as needed). There is no interactive prompt without a TTY. + + +## Multilingual + +The default `Xenova/bge-base-en-v1.5` model is trained for English. For German, Dutch, and other non-English corpora, dense recall can fail even when hybrid keyword search still finds rare tokens. + +For multilingual or non-English deployments, switch **before** large backfills: + +```bash +# Example: local multilingual (set dimensions to match the model) +SUPERMEMORY_EMBEDDING_PROVIDER=local +SUPERMEMORY_EMBEDDING_MODEL=Xenova/bge-m3 +SUPERMEMORY_EMBEDDING_DIMENSIONS=1024 +``` + +Or use a remote multilingual embedding API (OpenAI, Gemini, or Ollama with a multilingual embed model). Set provider, model, and dimensions together. Changing them later requires a fresh data directory or full re-ingestion — see below. + +## Remote providers + +### Local (default) + +```bash +# Explicit local default — no embedding API key +SUPERMEMORY_EMBEDDING_PROVIDER=local +SUPERMEMORY_EMBEDDING_MODEL=Xenova/bge-base-en-v1.5 +SUPERMEMORY_EMBEDDING_DIMENSIONS=768 +``` + +#### Supported local models & dimensions + +When using `SUPERMEMORY_EMBEDDING_PROVIDER=local`, local ONNX models do not support dimension reduction. `SUPERMEMORY_EMBEDDING_DIMENSIONS` must match the model's native dimensions: + +| Model ID | Native Dimensions | Description | +|---|---|---| +| `Xenova/bge-base-en-v1.5` | `768` | Default local model (English) | +| `Xenova/bge-m3` | `1024` | Multilingual (uses `cls` pooling) | +| `Xenova/bge-small-en-v1.5` | `384` | English (lightweight) | +| `Xenova/bge-large-en-v1.5` | `1024` | English (high capacity) | +| `Xenova/multilingual-e5-small` | `384` | Multilingual (lightweight) | +| `Xenova/multilingual-e5-base` | `768` | Multilingual | +| `Xenova/multilingual-e5-large` | `1024` | Multilingual (high capacity) | +| `Xenova/all-MiniLM-L6-v2` | `384` | English general purpose | +| `Xenova/paraphrase-multilingual-MiniLM-L12-v2` | `384` | Multilingual sentence similarity | + +### OpenAI + +```bash +OPENAI_API_KEY=sk-... +SUPERMEMORY_EMBEDDING_PROVIDER=openai +SUPERMEMORY_EMBEDDING_MODEL=text-embedding-3-small +SUPERMEMORY_EMBEDDING_DIMENSIONS=1536 +``` + +### Gemini + +```bash +GEMINI_API_KEY=... +SUPERMEMORY_EMBEDDING_PROVIDER=gemini +SUPERMEMORY_EMBEDDING_MODEL=text-embedding-004 +SUPERMEMORY_EMBEDDING_DIMENSIONS=768 +``` + +### Ollama (OpenAI-compatible) + +```bash +SUPERMEMORY_EMBEDDING_PROVIDER=openai-compatible +SUPERMEMORY_EMBEDDING_BASE_URL=http://localhost:11434/v1 +SUPERMEMORY_EMBEDDING_API_KEY=ollama +SUPERMEMORY_EMBEDDING_MODEL=nomic-embed-text +SUPERMEMORY_EMBEDDING_DIMENSIONS=768 +``` + +Use the dimension published for your chosen model. A mismatch with vectors already in the store fails boot. + +## Changing models later + + +**Not supported in place.** Embeddings from different models (or different dimensions) are not comparable. Start from a fresh data directory or re-ingest all content so vectors stay in one space. If configured dimensions disagree with stored data, the server **refuses to boot**. + + +**Changing embeddings later:** Not supported in place. Start from a fresh data directory or re-ingest all content so vectors stay comparable. + +> [!IMPORTANT] +> **Release Binary Env Var Support (v0.0.6 / v0.0.7-rc.2 vs v0.0.7+)** +> +> In release `v0.0.6` and `v0.0.7-rc.2`, the compiled standalone release binaries were built without pluggable embedding configuration hooks (environment variables like `SUPERMEMORY_EMBEDDING_PROVIDER`, `SUPERMEMORY_EMBEDDING_MODEL`, and `SUPERMEMORY_EMBEDDING_DIMENSIONS` were omitted from the binary and defaulted to local English `Xenova/bge-base-en-v1.5`). +> +> **Resolution & Upgrade Steps:** +> - Full pluggable embedding configuration via environment variables is active in `v0.0.7` and `v0.0.8+`. +> - **Existing Data Directory Notice:** If your data directory (`$SUPERMEMORY_DATA_DIR`) was created on `v0.0.6`, it contains pre-existing rows embedded with the 768d local model without an `embedding-plan.json` lock file. When upgrading to `v0.0.8+`, the server automatically assumes and locks to legacy `local · Xenova/bge-base-en-v1.5 · 768d` to preserve vector compatibility. +> - To switch to a custom provider or multilingual model (`bge-m3`, `openai`, etc.), you **must wipe the data directory** (e.g. `rm -rf "$SUPERMEMORY_DATA_DIR"`) or point `SUPERMEMORY_DATA_DIR` to a fresh directory before starting the server with your new embedding variables. + +> [!IMPORTANT] +> **Model Mixing Bug in v0.0.5 (Exact match returns nothing)** +> +> In version `v0.0.5`, there was a bug where the server could mix different embedding models between write and read paths (e.g., document ingestion using OpenAI but memory queries using local default embeddings). In multilingual contexts like Japanese (which lacks space tokenization for fallback lexical FTS matching), this caused exact-text memory searches through `/v4/search` and `/v4/profile` to silently return `{"results":[],"total":0}`. +> +> **Resolution:** +> This was fully resolved in `v0.0.7` by locking the embedding plan uniformly across all document and query embedding paths (enforced via a locked plan in the database store). If you are running `v0.0.5` and experiencing this issue, you should upgrade to `v0.0.7` or later. + +## Related + +- [Configuration](/self-hosting/configuration) — LLM providers, storage, ingestion limits +- [Quickstart](/self-hosting/quickstart) — install and first memory