Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
The package __init__ now re-exports AgentRuntime/AgentConfig/SessionState
plus both runtime implementations (ClaudeSDKAgentRuntime and the new
LiteLLMAgentRuntime) so callers can do
'from litellm.managed_agents.agent_runtime import LiteLLMAgentRuntime'
without poking into submodules.
Required for the package-level __init__.py re-exports to resolve.
Real Prisma translates {agent: {connect: {id: X}}} into the underlying
agent_id FK column at write time. The in-memory FakePrismaClient we
share with the proxy tests stores the dict verbatim, so the agent_id
column was None on every row our managed_agents code wrote. Tests
asserting session.agent_id == 'agent_1' (and the same for session_id /
run_id) all failed at the same boundary.
Fix: wrap each table's create() in a small adapter that unwraps the
relation-syntax keys into bare FK columns before delegating. Local to
this test conftest so the change does not leak into the proxy tests
that re-use the same fake.
The package __init__ now re-exports the public surface so SDK callers
can do a single 'from litellm.managed_agents import Agent, Session, ...'
instead of digging into submodules.
Also pins what's public — anything not in __all__ is implementation
detail and may move.
Session is the unit of work in the managed-agents API. send(prompt)
mirrors what the HTTP create_run + event_stream pair does on the proxy:
1. INSERT LiteLLM_AgentRun (status=queued, prompt={text}).
2. Flip session status ready -> busy and run status queued -> running.
3. async for event in self.runtime.run(...) — persist each event to
LiteLLM_AgentRunEvent and yield it back to the consumer.
4. On exit (clean or error), set the run terminal status and restore
session to ready in a finally block so a single failed send never
locks the session forever.
Also adds get_run/list_runs/conversation read helpers — same wire shape
the corresponding HTTP endpoints serve.
Agent is the SDK-side handle for one LiteLLM_Agent row. It wraps:
* the static agent definition (name, model, system_prompt, tools_config)
* the AgentRuntime + Sandbox that spawned sessions will use
* session lifecycle helpers (create_session, get_session, list_sessions,
delete)
create_session inserts the LiteLLM_AgentSession row directly in 'ready'
status, mints a daemon JWT (same shape the proxy mints in /v2/sessions),
and returns a Session ready to send() prompts.
The HTTP /v2/sessions path uses 'provisioning' because it kicks off a
remote VM provider call; the managed_agents Python layer owns the sandbox
in-process and is ready immediately.
Run is the read-only Python accessor for one LiteLLM_AgentRun row. It is
the SDK-side counterpart to GET /v2/sessions/{sid}/runs/{rid} and to the
SSE event stream.
Two entry points:
* Run.from_db_row(row, db) — build from a Prisma row.
* await run.stream(starting_seq=N) — async iterator over persisted
events ordered by seq, paginated through the DB to bound memory on
chatty runs.
Stream is snapshot-based; live tailing remains the SSE endpoint's job.
LiteLLMAgentRuntime drives a manual tool loop on top of litellm.acompletion
so the same managed agent can run against any provider that supports tool
calling (Anthropic, OpenAI, Gemini, Bedrock, etc.).
Unlike ClaudeSDKAgentRuntime, this runtime routes every tool call through
sandbox.execute_tool(), which means EC2SandboxViaSSM and any future
remote-execution sandbox actually get used.
Tools come from AgentConfig.tools_config (accepts both {tools: [...]} and
bare list shapes), with a sensible default that mirrors the LocalSandbox
surface (Bash/Read/Write/Edit/ls).
Backend SSE frames are {seq, event_type, payload} but the SDK was
reading obj.type / obj.data without any snake-to-camel transform —
all events were being dropped. Apply the same snakeToCamel walk used
for JSON responses, then map event_type/payload to type/data so the
public RunEvent shape stays stable. Also extend the terminal-event
check so the stream returns on the actual backend terminal types
(run_finished | run_cancelled | run_error) instead of hanging until
the keep-alive timeout.