Real Prisma translates {agent: {connect: {id: X}}} into the underlying
agent_id FK column at write time. The in-memory FakePrismaClient we
share with the proxy tests stores the dict verbatim, so the agent_id
column was None on every row our managed_agents code wrote. Tests
asserting session.agent_id == 'agent_1' (and the same for session_id /
run_id) all failed at the same boundary.
Fix: wrap each table's create() in a small adapter that unwraps the
relation-syntax keys into bare FK columns before delegating. Local to
this test conftest so the change does not leak into the proxy tests
that re-use the same fake.
The package __init__ now re-exports the public surface so SDK callers
can do a single 'from litellm.managed_agents import Agent, Session, ...'
instead of digging into submodules.
Also pins what's public — anything not in __all__ is implementation
detail and may move.
Session is the unit of work in the managed-agents API. send(prompt)
mirrors what the HTTP create_run + event_stream pair does on the proxy:
1. INSERT LiteLLM_AgentRun (status=queued, prompt={text}).
2. Flip session status ready -> busy and run status queued -> running.
3. async for event in self.runtime.run(...) — persist each event to
LiteLLM_AgentRunEvent and yield it back to the consumer.
4. On exit (clean or error), set the run terminal status and restore
session to ready in a finally block so a single failed send never
locks the session forever.
Also adds get_run/list_runs/conversation read helpers — same wire shape
the corresponding HTTP endpoints serve.
Agent is the SDK-side handle for one LiteLLM_Agent row. It wraps:
* the static agent definition (name, model, system_prompt, tools_config)
* the AgentRuntime + Sandbox that spawned sessions will use
* session lifecycle helpers (create_session, get_session, list_sessions,
delete)
create_session inserts the LiteLLM_AgentSession row directly in 'ready'
status, mints a daemon JWT (same shape the proxy mints in /v2/sessions),
and returns a Session ready to send() prompts.
The HTTP /v2/sessions path uses 'provisioning' because it kicks off a
remote VM provider call; the managed_agents Python layer owns the sandbox
in-process and is ready immediately.
Run is the read-only Python accessor for one LiteLLM_AgentRun row. It is
the SDK-side counterpart to GET /v2/sessions/{sid}/runs/{rid} and to the
SSE event stream.
Two entry points:
* Run.from_db_row(row, db) — build from a Prisma row.
* await run.stream(starting_seq=N) — async iterator over persisted
events ordered by seq, paginated through the DB to bound memory on
chatty runs.
Stream is snapshot-based; live tailing remains the SSE endpoint's job.
LiteLLMAgentRuntime drives a manual tool loop on top of litellm.acompletion
so the same managed agent can run against any provider that supports tool
calling (Anthropic, OpenAI, Gemini, Bedrock, etc.).
Unlike ClaudeSDKAgentRuntime, this runtime routes every tool call through
sandbox.execute_tool(), which means EC2SandboxViaSSM and any future
remote-execution sandbox actually get used.
Tools come from AgentConfig.tools_config (accepts both {tools: [...]} and
bare list shapes), with a sensible default that mirrors the LocalSandbox
surface (Bash/Read/Write/Edit/ls).
Backend SSE frames are {seq, event_type, payload} but the SDK was
reading obj.type / obj.data without any snake-to-camel transform —
all events were being dropped. Apply the same snakeToCamel walk used
for JSON responses, then map event_type/payload to type/data so the
public RunEvent shape stays stable. Also extend the terminal-event
check so the stream returns on the actual backend terminal types
(run_finished | run_cancelled | run_error) instead of hanging until
the keep-alive timeout.
Backend RunCreate expects {prompt: {...}} but the SDK was posting
{text, images} directly. Wrap the normalized payload accordingly so
session.send returns a real Run instead of 422.
Backend exposes POST /v2/sessions (with agent_id in body) and
GET /v2/sessions; there is no /v2/agents/{id}/sessions route. Also
align getSession/listSessions to the same routes so SDK -> proxy
calls actually resolve. listSessions filters by agent_id via the
existing query param.
Daemon events:append was failing with MissingRequiredValueError because
the bare run_id FK and dict payload don't pass Prisma's input validation
on the LiteLLM_AgentRunEvent.create call.
The Prisma generated client wants {"connect": {"id": ...}} for the
session relation and prisma.Json(...) wrapping for the prompt + event
payload Json columns.
The generated Prisma client requires {"connect": {"id": ...}} for
relation fields and prisma.Json(...) for Json columns. Bare agent_id
strings and bare dict/list values both raise MissingRequiredValueError
on session create.
Prisma rejects bare None / dict / list for optional Json columns with
MissingRequiredValueError. Wrap dict/list values in prisma.Json(...) and
drop the key entirely when the source is None so the column resolves to
SQL NULL.
Reconciliation:
- adopt B's richer base.py types (ProvisionContext, VMHandle, AwsCreds,
Ec2Config, ProvisionError) as canonical; keep A's NoopVMProvider alias
and the registry helpers (register_vm_provider, reset_vm_provider_registry)
for tests.
- rename B's factory entry point from get_vm_provider to build_vm_provider so
it doesn't collide with the runtime registry's get_vm_provider(name).
- update A's session_endpoints._provision_in_background to construct a
ProvisionContext and call provider.provision(ctx); pass team_id from
the caller's API key.
- update _terminate_session_internal to construct a VMHandle when a vm_id
is recorded on the session row.
- extend B's NoopProvider with provision_calls/terminate_calls recording for
backward compat with A's tests; accept either VMHandle-style or legacy
keyword-style terminate args.
- de-duplicate LiteLLM_AgentVMConfig from all 3 schema.prisma files —
G's (LIT-2891) version wins; B's stub at the top is collapsed to a
comment pointing to G's section.
- delete B's 20260506220000_add_agent_vm_config migration (collides with
G's 20260506220000_add_cloud_agent_settings_tables which already creates
the table).
- rewrite team_config.py to read G's per-field encrypted columns
(aws_access_key_id_enc, aws_secret_access_key_enc, aws_region) instead
of B's single-blob aws_creds_enc; rewrite test_team_config.py to match.
Resolved schema conflicts in schema.prisma, litellm/proxy/schema.prisma,
and litellm-proxy-extras/litellm_proxy_extras/schema.prisma by unioning
A's agent/session/run/event tables with G's VM config / secrets /
worker / pairing-token tables. Both sets of tables now coexist.
Mock proxy now serializes responses using snake_case keys (agent_id,
created_at, system_prompt, run_id, etc.) and reads request bodies as
snake_case so it matches the real backend that the SDK now talks to via
the new transform layer. Also update the status string literals to the
new SessionStatus and RunStatus values, and read the followup body as
{prompt: {text}}.
There is no per-run conversation endpoint on the backend; conversation
history is session-scoped. Callers should use SessionHandle.conversation()
instead. Also update TERMINAL_STATES to the new RunStatus values
(finished/cancelled/error).
The backend's followup endpoint expects a FollowupCreate payload with a
nested prompt object ({prompt: {text: ...}}), not a flat {message: ...}.
Public method signature followup(message: string) is unchanged - only
the wire body changes.
Backend speaks snake_case (Python idiom) while the SDK's public TS API
is camelCase. Add recursive snakeToCamel and camelToSnake helpers and
wire them into the HTTP layer so request bodies are camel->snake before
JSON.stringify and response JSON is snake->camel before being returned
to callers. Single-word keys like id/type/data/seq/status pass through
unchanged in both directions, and non-object values are not touched.
Backend (LIT-2890) emits provisioning/ready/busy/error/terminated for
session status and queued/running/finished/cancelled/error for run status.
Update the SDK type aliases so callers compare against the actual values
the proxy returns over the wire.
Greptile P3 (regression coverage): the cascade filter in
delete_agent was changed from terminated_at is None to
status not in SESSION_TERMINAL_STATUSES in commit a0015e8564.
Add an explicit test that exercises the bug surface — a session
flipped to error (terminal) but with terminated_at deliberately
left None. The legacy filter would have re-terminated it; the
status-based filter must skip it.