Phase 2 architecture (LIT-2879): the agent runtime runs on the
per-Session VM, not in the proxy. Add a pytest that scans
litellm/proxy/ and fails if any source file imports claude_agent_sdk.
Currently passes (zero offenders). The test enforces that the boundary
stays intact going forward.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
Phase 2 architecture rule: VM management code lives in
litellm/managed_agents/vms/, not litellm/proxy/. The proxy is control
plane only.
Tracks LIT-2879.
The package __init__ now re-exports AgentRuntime/AgentConfig/SessionState
plus both runtime implementations (ClaudeSDKAgentRuntime and the new
LiteLLMAgentRuntime) so callers can do
'from litellm.managed_agents.agent_runtime import LiteLLMAgentRuntime'
without poking into submodules.
Required for the package-level __init__.py re-exports to resolve.
Real Prisma translates {agent: {connect: {id: X}}} into the underlying
agent_id FK column at write time. The in-memory FakePrismaClient we
share with the proxy tests stores the dict verbatim, so the agent_id
column was None on every row our managed_agents code wrote. Tests
asserting session.agent_id == 'agent_1' (and the same for session_id /
run_id) all failed at the same boundary.
Fix: wrap each table's create() in a small adapter that unwraps the
relation-syntax keys into bare FK columns before delegating. Local to
this test conftest so the change does not leak into the proxy tests
that re-use the same fake.
The package __init__ now re-exports the public surface so SDK callers
can do a single 'from litellm.managed_agents import Agent, Session, ...'
instead of digging into submodules.
Also pins what's public — anything not in __all__ is implementation
detail and may move.
Session is the unit of work in the managed-agents API. send(prompt)
mirrors what the HTTP create_run + event_stream pair does on the proxy:
1. INSERT LiteLLM_AgentRun (status=queued, prompt={text}).
2. Flip session status ready -> busy and run status queued -> running.
3. async for event in self.runtime.run(...) — persist each event to
LiteLLM_AgentRunEvent and yield it back to the consumer.
4. On exit (clean or error), set the run terminal status and restore
session to ready in a finally block so a single failed send never
locks the session forever.
Also adds get_run/list_runs/conversation read helpers — same wire shape
the corresponding HTTP endpoints serve.