From review of 5b2a435e:
- Captured turns now keep the container tag and document id they were
spoken under. Rebinding the instance used to flush earlier turns into
the new caller's scope. Turns captured before any bind still go to the
first caller bound.
- Recall strips earlier injected memory before it runs, so a failed or
empty recall can no longer leave another caller's memory in context.
For the same caller, a slow or failed recall falls back to the profile
loaded by preload.
- The recall cache is keyed on the user message, not its text, so a
later turn with the same words recalls again.
- Capture writes time out after 10s and keep their turns for retry,
including on cancellation. Large calls are written in chunks of at
most 100k characters, and the buffer is capped.
- remember falls back to the default processing schedule when the
organization has no balance for instant processing (HTTP 402).
- Docs state the measured delays: about a minute for remember, 10 to 20
minutes for captured calls on the default dynamic schedule.
Found in live LiveKit Cloud voice calls:
- Recalling in on_user_turn_completed changed the turn context, so
LiveKit discarded its preemptive generation on every turn with memory.
SupermemoryAgent now recalls only in llm_node, and uses the turn hook
only for realtime models, which skip llm_node.
- A turn with no memory recalled twice (hook, then llm_node), and tool
follow-ups recalled again. Each user message is now recalled once;
retries and follow-ups share the result.
- A preloaded profile made llm_node skip recall for the rest of the
call. Fresh recall now replaces the earlier injection.
- remember used the default dynamic dreaming, so an explicit fact took
about 18 minutes to become recallable. It now uses instant (~40s).
The call transcript mode is configurable via capture_dreaming.
- recall_timeout defaults to 2s; profile calls measured 0.5-1.4s.
- Require livekit-agents>=1.3.6: AgentServer (used in the quick start)
arrived in 1.3.1, and 1.3.1-1.3.5 no longer import with current
opentelemetry-sdk.
session.run and generate_reply never call on_user_turn_completed, so a
text turn reached the model with no memory. SupermemoryAgent now recalls
inside llm_node and skips the fetch when the voice hook already injected.