`supermemoryProfileSearch` in `shared/memory-client.ts` is the only
Supermemory HTTP call in this package with neither a request timeout nor
redirect handling. The identical `/v4/profile` call in
`openai/middleware.ts` sets both, and `/v4/conversations`
(`conversations-client.ts`) and `/v4/memories` (`shared/forget-memory.ts`)
each set a 30s budget.
Two consequences:
- **Unbounded request.** A timeout only applied when the caller supplied a
signal. `withSupermemory` passes one (5s), but `buildMemoriesText` is
called with no signal by the Mastra processor and the VoltAgent
middleware, and by the exported `buildMemoriesText` / `addSystemPrompt`
helpers. `fetch` has no default deadline, so a stalled connection blocks
the agent turn indefinitely — the failure both integrations' surrounding
try/catch is written to absorb, but which never surfaces as an error.
- **Redirects followed.** The request carries `Authorization: Bearer
<apiKey>`; a 3xx from a misconfigured or attacker-influenced `baseUrl`
was followed silently rather than refused.
Apply a 30s `PROFILE_REQUEST_TIMEOUT_MS` unconditionally and set
`redirect: "error"`. A caller signal is composed with the timeout via
`AbortSignal.any` rather than replacing it, so a caller-side budget can
only shorten the request, never leave it unbounded — the wrapper is kept
separate so the composition is stated once rather than re-derived at the
call site.
`src/shared/memory-client.test.ts` existed but was absent from the
`test:unit` file list CI runs, so its assertions never ran on a pull
request; add it alongside the new coverage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7GkmUn6skD6dCzKtcHDbe
Same pattern as forgetMemoryRequest: `options?.signal ?? AbortSignal.timeout(...)`
dropped the 30s bound whenever a caller supplied its own signal. Compose the
two with AbortSignal.any so the caller signal adds cancellation instead of
replacing the timeout.
`forgetMemoryRequest` combined the caller's signal and the 30s abort with
`??`, making them mutually exclusive. Passing a cancellation signal removed
the timeout, so a hung `DELETE /v4/memories` could wedge the tool call again
— the exact condition #1451 set out to remove. There was also no way for a
caller to ask for both cancellation and a timeout.
Compose the two with `AbortSignal.any` instead of choosing between them.
`AbortSignal.any` is available in Node 20.3+, Bun and workerd.
No production call site passes `options` today (`ai-sdk.ts` and
`openai/tools.ts` both omit it), so this was latent rather than live.
The existing test asserted the buggy behaviour (`init.signal` being the
caller's own signal), so it is replaced by two tests that pin the composed
semantics: aborting the caller aborts the request, and the timeout leg still
aborts the request on its own. Both fail against the previous implementation.
Fixes#1549
Three places where ClaudeMemoryTool diverges from the documented
memory_20250818 wire format:
- rename sends old_path/new_path, not path. handleCommand validated
command.path, so every rename coming from a real model died with
"Cannot read properties of undefined (reading 'startsWith')".
path is still accepted as the source for existing callers.
- insert_line means "insert after this line" (0 = top of file), but we
spliced at insertLine - 1 and rejected 0, so every insert landed one
line above where Claude asked and inserting at the top was impossible.
- str_replace with new_str omitted is a deletion per the spec; we
rejected it.
The new tests mock the supermemory client so they run without an API
key. Also fixed the rename example in the docs, which showed the same
path shape the code expected.
Anthropic's memory_20250818 spec defines insert as: insert_text is inserted
AFTER line insert_line, 0 inserts at the beginning of the file, and the valid
range is [0, n_lines]. The implementation treated insert_line as a 1-based
insert-BEFORE index with range [1, n_lines + 1].
Since the caller of this tool is Claude itself, which is trained on the spec
semantics, every model-driven insert landed one line earlier than intended,
insert_line: 0 (insert at top of file) was rejected as invalid, and
insert_line: n_lines (append) inserted before the last line instead of after
it.
Fix the validation range to [0, n_lines], splice at insert_line directly
(0-based insert-after), and update the error and success messages to match.
One existing tool-operations test encoded the old insert-before behavior; its
insert_line is adjusted so its expected output is unchanged under spec
semantics. Adds four regression tests covering top-of-file, middle,
append, and both out-of-range directions.
- 402 now tells the agent the org is out of credits and links to console.supermemory.ai/billing (the old message said "memory limit" and pointed at the retired app).
- PostHog MCP analytics tags these calls `$mcp_error_type: out_of_credits` with `$mcp_is_error: false`, so an empty balance no longer inflates the MCP error rate.
The server is stateless HTTP, so every tool call landed in its own PostHog session. Enable conversation ids and let $mcp_conversation_id and $session_id through the metadata filter so echoed handles group calls.
## Summary
- Instrument the MCP v2 server with the pinned PostHog MCP Analytics SDK and replace the custom `mcp_tool_executed` wrapper with standard `$mcp_*` events.
- Send only allowlisted metadata, disable schema injection and exception autocapture, and use personless user IDs with person-profile processing disabled.
- Deliver events through immediate capture guarded by Cloudflare `waitUntil`.
## Verification
- The initial implementation passed the MCP typecheck, existing unit suite, Biome, and Wrangler dry-run bundle.
- A local `who_am_i` MCP call returned successfully and emitted a metadata-only `$mcp_tool_call` on the initial implementation.
- All five checks passed on the final `54a2384` head, including the Cloudflare MCP build.
Actual PostHog ingestion is not verified yet; this workspace still needs a PostHog project token and authenticated Supermemory MCP credential.
Document-to-memory links and actual `derives` relations were both emitted as `derives`, so they shared the same color and legend entry.
This separates structural document links into a `document` edge type, adds a dedicated theme color with `--graph-edge-document` support, and updates force-layout and level-of-detail handling to preserve existing structural behavior. The package and MCP widget legends/themes now distinguish document links from derived-memory relations.
Adds regression coverage for edge classification and validates the package plus its MCP consumer.
<!-- capy-badge:start -->
<a href="https://capy.ai/thread/jam_01M36F4MPFXZCA8J14T53YY029"><picture><source media="(prefers-color-scheme: dark)" srcset="https://capy.ai/badge/accent-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://capy.ai/badge/accent-light.svg"><img alt="Open in Capy" src="https://capy.ai/badge/accent-light.svg"></picture></a>
<!-- capy-badge:end -->
## Summary
The billing guide omits Max and contradicts the pricing page about Gmail access. Add Max ($100/month with $130 in monthly credits), correct the plan feature matrix, list all six current coding plugins as available on Free with credit-consuming usage, and update both remaining Scale-only requirements on the Gmail connector page to Max or above.
Clarify that paid-plan automatic top-up limits cap automatic credit purchases, not the total invoice, and correct the documented auto-topup update endpoint from PATCH to POST.
Companion website changes: https://github.com/supermemoryai/landing-2/pull/32
## Validation
- Compared prices, credit inclusions, connector gates, and top-up behavior with the console and API source on main. Verified that all six coding plugins are on FREE_TIER_PLUGIN_IDS and the authentication route bypasses the generic Pro gate for them. Usage still consumes plan credits.
- Checked Markdown table column counts and required Max entries; `git diff --check` passed.
- Confirmed the Gmail introduction, prerequisite, and troubleshooting requirement consistently say Max or above.
- Documentation-only change. A full Mintlify build was not run.
Crawler-access verification, AI-search benchmarking, and analytics/measurement work are excluded from this PR.
Before: npm serves 0.2.3 without the merged fixes.
After: merging publishes 0.2.4 with the theme, initial fit, and node settling fixes.
Validation: package typecheck and build passed.
**Before:** Console lost its theme and dots. Incoming nodes needed a drag to reorganize, fit missed later pages, and clicks or drag release could leave the layout moving.
**After:** Restore themed rendering with configurable dots. Automatically settle and fit incoming nodes, keep clicks from reheating forces, and cool the layout after drag release.
**Checked:** Package types/build and Console build linked to this package.
Console companion: [mono#3295](https://github.com/supermemoryai/mono/pull/3295).
## Summary
The `workflow_run` auto-fix job has write permissions and checks out the triggering PR branch. It now runs only when all three conditions hold:
- the tracked CI workflow failed on a pull request;
- the pull request branch belongs to this repository, not a fork;
- the triggering actor is not a bot account.
The same-repository check closes the privileged fork boundary. The generic `[bot]` suffix check covers Polylane, Graphite, Dependabot, and other GitHub App bot users without maintaining a name list.
## Validation
- `go run github.com/rhysd/actionlint/cmd/actionlint@v1.7.12 .github/workflows/claude-auto-fix-ci.yml`
- YAML parse with the repository's installed parser
- Final diff audit: one workflow, +3/-1, no added comments
Human-triggered same-repository pull requests keep the existing auto-fix behavior.
chore(web): reduce the app to a redirect shell, drop the browser extension
app.supermemory.ai now forwards everything to the console: plugin, OAuth and invite paths get an immediate 308 with the query intact, everything else shows a short notice first. Removes the browser extension workspace.
chore(web): give the moved notice a proper design
Hostnames become the headline, one primary action, a draining line for the countdown, DM Sans and the dot-grid backdrop from the brand.
chore(docs): point docs and README at the console, drop stale sections
Docs and both READMEs now link to console.supermemory.ai for API keys. Removed the company-brain docs tab with a redirect, and trimmed the README app section.
chore(web): give the redirect five seconds
getDocument filtered on the caller's active space, so an ID from listDocuments in any other space returned "Document not found". With activeSpace unset the fallback is sm_project_default, which broke most cross-space reads.
The API already scopes document reads to the caller's org, so the extra filter added no protection. Verified locally against the mono API: own-space and cross-space IDs now resolve, foreign-org IDs still 404.
Ship dedicated integration docs and point the plugin catalog at them. Grok Bot stays to install, auth, and skills — no Cursor-only config or repo tags.
## Stack Context
Part 3 (top) of a 3-PR stack moving memory deduplication into the SDKs. See `sdk-dedup/tools-ts` for full context.
## What?
Update the SDK playground so its debug view reflects the SDK-owned memory block.
- Displays the current deduplicated `<supermemory>` replacement block produced by the SDK middleware, instead of the old browser-side "seen facts" delta.
- Adds a `memory-dedupe` helper and ignores local `*.tsbuildinfo`.
## Why?
The previous debug cards were misleading — they showed an incremental browser-filtered delta while the middleware actually re-injected the full profile. Now the visualization matches what the SDK really sends.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- CURSOR_SUMMARY -->
---
> [!NOTE]
> **Low Risk**
> Playground-only visualization and chat gating changes; no production SDK or API behavior.
>
> **Overview**
> The playground **debug trace** now shows the **deduplicated memory block** the SDK middleware would inject (static → dynamic → search, mode-aware), instead of a misleading browser-side “new facts” delta. A new **`memory-dedupe`** helper mirrors `@supermemory/tools` middleware behavior and is applied when fetching container context and building middleware memory debug entries; the context preview card is relabeled to reflect that each turn **replaces** the prior `<supermemory>` block.
>
> **Chat UX:** messaging is enabled when API keys are configured on the **server** (`hasSupermemoryKey` / `hasOpenAiKey` from `/api/chat`), not only when keys are typed in the panel. The message input stays editable while waiting for text; Send still requires non-empty input.
>
> Also ignores `*.tsbuildinfo` in `.gitignore`.
>
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit ed15364eb3. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
## Stack Context
Part 2 of a 3-PR stack moving memory deduplication into the SDKs. See `sdk-dedup/tools-ts` (parent) for the full context and the TypeScript implementation this mirrors.
## What?
Port the normalized, priority-ordered (`static > dynamic > search`) profile deduplication into the Python SDKs.
- Each request injects one **owned memory block that replaces** the prior block rather than accumulating.
- Dedup is **request-local** (no shared state), so it stays correct under concurrency.
Covers OpenAI, Agent Framework (middleware + context provider), Cartesia, and Pipecat.
## Why?
Keeps the Python SDKs at behavioral parity with the TypeScript SDK so all integrations deduplicate memory the same way.
## Testing
- OpenAI: 31 passed, 11 skipped (live)
- Agent Framework: 59 passed
- Cartesia: 8 passed
- Pipecat: 8 passed
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- CURSOR_SUMMARY -->
---
> [!NOTE]
> **Medium Risk**
> Changes memory formatting and system-prompt injection across multiple SDK integrations; incorrect dedup or replacement could alter LLM context, but there is no auth or data-store risk.
>
> **Overview**
> Ports **normalized cross-source memory deduplication** and **replace-not-append injection** into the Python OpenAI, Agent Framework, Cartesia, and Pipecat packages so they match the TypeScript SDK behavior.
>
> **Deduplication** uses request-local keys: strip optional `[YYYY-MM-DD]` prefixes, normalize whitespace, and compare with `casefold`, with priority **static → dynamic → search**. In **`query` mode**, profile static/dynamic are excluded from dedup input so facts that only appear in search (or overlap profile) are not dropped before formatting.
>
> **Injection** no longer appends memory text every turn. OpenAI and Agent Framework middleware **strip prior owned `<supermemory context="user-memories" readonly>` blocks** and **replace** them once per request while keeping the caller’s system instructions; extra system messages lose stale blocks only. New helpers (`strip`/`replace`/`wrap`) live in each package’s utils.
>
> Tests cover normalized fact variants, query-mode search retention, and stale block replacement.
>
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 42f308b224. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
## Stack Context
This stack moves memory deduplication **out of the playground UI and into the SDKs themselves**, so every integration injects a single, deduplicated, self-replacing memory block. Three PRs:
1. **`sdk-dedup/tools-ts`** (this PR) — TypeScript SDK core + integrations
2. `sdk-dedup/python` — Python SDKs
3. `sdk-dedup/playground` — playground debug view reflects the SDK-owned block
## What?
Move profile deduplication into the SDK middleware for the TypeScript tools package.
- Facts are normalized (strip leading `[YYYY-MM-DD]`, trim, collapse whitespace, casefold) and deduplicated in **`static > dynamic > search`** priority within a single request.
- The result is injected as one **owned `<supermemory>` block** that *replaces* the previous block instead of accumulating a new one each turn.
- Dedup is **mode-aware**: in query mode, search results are not dropped against a profile that isn't being injected.
- Deduplication is **request-local** — no global/browser `Set`. Safe for multiple users, concurrent requests, and Cloudflare Worker isolates.
Covers AI SDK, OpenAI (Chat + Responses), Mastra, and VoltAgent. New `shared/memory-context.ts` owns the block-replacement logic.
## Why?
The earlier "conversation-scoped deduplication" was only a playground browser `Set` — a UI debug affordance that did not change what the SDK sent to the model, and would have been unsafe as server-side global state. Real cross-source dedup belongs in the SDK, applied fresh per stateless model request.
## Testing
- `bun run test` in `packages/tools`: 145 passed (the one failing suite, `claude-memory.test.ts`, is a pre-existing broken import unrelated to this change).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- CURSOR_SUMMARY -->
---
> [!NOTE]
> **Medium Risk**
> Changes how system prompts and instructions are built across all TypeScript integrations; behavior is well-covered by unit tests but incorrect strip/replace logic could drop or duplicate context in production prompts.
>
> **Overview**
> Moves **cross-source memory deduplication** and **owned prompt injection** into `@supermemory/tools` so every integration sends one deduplicated memory block per request instead of growing context each turn.
>
> **Deduplication:** Facts are normalized via `normalizeMemoryFact` (strip `[YYYY-MM-DD]`, trim, collapse whitespace, lowercase) and deduplicated with **static → dynamic → search** priority. `deduplicateMemoriesForMode` keeps search hits in **query** mode when the profile is not injected.
>
> **Owned `<supermemory>` block:** New `shared/memory-context.ts` wraps memories in `<supermemory context="user-memories" readonly>`, strips stale blocks, and **replaces** prior SDK context while preserving caller system instructions. Applied in AI SDK (`injectMemoriesIntoParams`), OpenAI Chat/Responses middleware, Mastra input processor (`wrapMemoryContext`), and VoltAgent hooks.
>
> **Tests:** Unit coverage for block replacement (with-supermemory, OpenAI, VoltAgent), Mastra wrapper tag assertion, normalized dedup variants, and concurrent `containerTag` isolation.
>
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 2fa2e0d85c. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
## Summary
- Add `apps/sdk-playground` — chat UI to test TS/Python SDK integrations
- Context panel with document memories, API keys in dashboard, tools reference tab
- Python FastAPI server on port 8792; portless entry in `portless.json`
Stacked on #1436
## Test plan
- [ ] `cd apps/sdk-playground && bun run check-types`
- [ ] `bun run dev` with Supermemory + OpenAI keys in UI
- [ ] Switch SDKs and verify chat + context panel
Made with [Cursor](https://cursor.com)