supermemory/apps/docs/integrations/livekit.mdx
Ishaan Gupta 2d281f8cbc fix(livekit): recall once per turn in llm_node and remember instantly
Found in live LiveKit Cloud voice calls:

- Recalling in on_user_turn_completed changed the turn context, so
  LiveKit discarded its preemptive generation on every turn with memory.
  SupermemoryAgent now recalls only in llm_node, and uses the turn hook
  only for realtime models, which skip llm_node.
- A turn with no memory recalled twice (hook, then llm_node), and tool
  follow-ups recalled again. Each user message is now recalled once;
  retries and follow-ups share the result.
- A preloaded profile made llm_node skip recall for the rest of the
  call. Fresh recall now replaces the earlier injection.
- remember used the default dynamic dreaming, so an explicit fact took
  about 18 minutes to become recallable. It now uses instant (~40s).
  The call transcript mode is configurable via capture_dreaming.
- recall_timeout defaults to 2s; profile calls measured 0.5-1.4s.
- Require livekit-agents>=1.3.6: AgentServer (used in the quick start)
  arrived in 1.3.1, and 1.3.1-1.3.5 no longer import with current
  opentelemetry-sdk.
2026-09-28 20:56:25 +05:30

181 lines
7.7 KiB
Text

---
title: "LiveKit"
sidebarTitle: "LiveKit (Voice)"
description: "Add persistent memory to LiveKit Agents voice sessions"
icon: "/images/livekit.svg"
---
Supermemory integrates with [LiveKit Agents](https://docs.livekit.io/agents/), so a voice agent can remember a caller across sessions and use that context on the next turn. Recall runs before the model replies. Completed turns are stored automatically. The model can also search, remember, and forget explicitly.
A Supermemory timeout or outage is skipped. It does not end the call.
## Installation
```bash
pip install supermemory-livekit
```
```bash
export SUPERMEMORY_API_KEY=your_supermemory_api_key
```
Create a key at [console.supermemory.ai](https://console.supermemory.ai). For a self-hosted API, pass `base_url` to `SupermemoryLiveKit`.
## Scope
Memory is isolated by container tag. Use a stable caller id, the same one you use in the rest of your app, so a LiveKit call and a chat session share one profile.
Container tags may only contain letters, numbers, `_`, `-`, and `:`, and must be 100 characters or fewer. Two ways to set it:
- Pass `container_tag` yourself, usually from [job metadata](https://docs.livekit.io/agents/build/external-data/).
- Let the plugin read the participant attribute `supermemory_container_tag`, or fall back to the participant identity.
Identities that are not valid container tags are sanitized to a stable tag. Set the attribute when the identity is an email, phone number, or SIP address and you already have memories under a different id.
Set the attribute in the access token your backend issues. Do not grant the caller `canUpdateOwnMetadata`, or they could change the attribute and read another caller's memory.
One call is stored as a single document. Pass the LiveKit room name as `session_id`. The document id is `lk-<session_id>`, so a later update with the same id appends to that call instead of creating another.
## Quick start
Create one `SupermemoryLiveKit` inside the job. A worker handles many calls, and each call needs its own scope.
```python
import json
import os
from livekit import agents
from livekit.agents import AgentServer, AgentSession, ChatContext, JobContext
from supermemory_livekit import SupermemoryAgent, SupermemoryLiveKit
server = AgentServer()
@server.rtc_session(agent_name="memory-agent")
async def entrypoint(ctx: JobContext):
memory = SupermemoryLiveKit(api_key=os.environ["SUPERMEMORY_API_KEY"])
metadata = json.loads(ctx.job.metadata or "{}")
container_tag = metadata.get("container_tag")
if container_tag:
memory.bind(container_tag=container_tag, session_id=ctx.room.name)
chat_ctx = ChatContext()
if container_tag:
await memory.preload(chat_ctx)
await ctx.connect()
if not container_tag:
participant = await ctx.wait_for_participant()
memory.bind(participant=participant, session_id=ctx.room.name)
session = AgentSession(
stt="deepgram/nova-3:en",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3",
)
memory.attach(session)
await session.start(
room=ctx.room,
agent=SupermemoryAgent(
memory,
chat_ctx=chat_ctx,
instructions=(
"You are a helpful voice assistant. You remember this caller across calls. "
"Use that naturally, and do not mention the memory system."
),
),
)
if __name__ == "__main__":
agents.cli.run_app(server)
```
`preload` puts the caller's profile into the first turn, so the greeting can use it. `attach` stores the conversation after each agent reply, and stores anything left when the session closes, including a caller who hangs up mid-turn. `SupermemoryAgent` recalls inside `llm_node` before each reply and adds the memory tools. That covers voice turns, text turns from `generate_reply` or `session.run`, and LiveKit's preemptive generation. Each user message is recalled once, and tool follow-ups reuse it.
`agent_name` turns on explicit dispatch, which is how job metadata reaches the agent. Remove it to join every new room automatically, for example when testing in the [Agents Playground](https://agents-playground.livekit.io).
## Your own agent
If you already subclass `Agent`, keep that class. Pass the tools in, and recall from `llm_node`.
```python
from livekit.agents import Agent
from supermemory_livekit import SupermemoryLiveKit
class Assistant(Agent):
def __init__(self, memory: SupermemoryLiveKit) -> None:
super().__init__(
instructions="You are a helpful voice assistant.",
tools=memory.tools(),
)
self.memory = memory
async def llm_node(self, chat_ctx, tools, model_settings):
await self.memory.enrich(chat_ctx)
return Agent.default.llm_node(self, chat_ctx, tools, model_settings)
```
Call `memory.attach(session)` before `session.start`.
Recall in `llm_node` rather than `on_user_turn_completed`. Changing the turn context in that hook makes LiveKit discard its preemptive generation, which adds latency to every turn.
Realtime models skip `llm_node`. With a realtime model, call `await self.memory.on_user_turn_completed(turn_ctx, new_message)` from `on_user_turn_completed` instead. That hook only runs when turn detection runs in your agent, not inside the model. See [LiveKit's external data guide](https://docs.livekit.io/agents/build/external-data/). `SupermemoryAgent` picks the right hook for you. Turn capture still listens for `conversation_item_added`.
## What gets recalled
| Mode | Static profile | Dynamic profile | Search | Use when |
| --- | --- | --- | --- | --- |
| `profile` | Yes | Yes | No | Durable facts are enough |
| `query` | No | No | Yes | Only this turn's related memories |
| `full` | Yes | Yes | Yes | Default |
```python
from supermemory_livekit import InputParams, SupermemoryLiveKit
memory = SupermemoryLiveKit(
container_tag="user_123",
params=InputParams(
mode="full",
search_limit=10,
search_threshold=0.1,
recall_timeout=2.0,
capture="always",
capture_dreaming="dynamic",
),
)
```
Recall waits at most `recall_timeout` seconds (default 2). If the profile call is slower than that, the turn proceeds without memory. Retrieved text is inserted immediately before the user message and is not stored back as something the agent said.
| Parameter | Default | Description |
| --- | --- | --- |
| `search_limit` | `10` | Maximum search results merged into the turn |
| `search_threshold` | `0.1` | Minimum similarity, from 0 to 1 |
| `mode` | `"full"` | `profile`, `query`, or `full` |
| `recall_timeout` | `2.0` | Seconds to wait before skipping recall |
| `capture` | `"always"` | `never` disables storing the call |
| `capture_dreaming` | `"dynamic"` | How the call transcript becomes memories. `dynamic` batches related documents and can take several minutes. `instant` is usually ready within a minute and bills one extra operation per document. Until then, `search_memories` still finds the raw call text. |
## Tools
`SupermemoryAgent` adds these tools. They are scoped to the bound container tag, so the model cannot read or write another caller.
| Tool | Used for |
| --- | --- |
| `search_memories` | Look up facts and past calls when the automatic recall is not enough |
| `remember` | Save one explicit fact, preference, or correction |
| `forget` | Forget one fact by id from `search_memories`, or by exact text |
`remember` stores a standalone fact and processes it right away, so the next call can recall it. It is not appended to the call transcript. Automatic capture is what records the conversation.
## Self-hosting
```python
memory = SupermemoryLiveKit(
api_key=os.environ["SUPERMEMORY_API_KEY"],
base_url="http://localhost:6767",
container_tag="user_123",
)
```