Backend RunCreate expects {prompt: {...}} but the SDK was posting
{text, images} directly. Wrap the normalized payload accordingly so
session.send returns a real Run instead of 422.
Backend exposes POST /v2/sessions (with agent_id in body) and
GET /v2/sessions; there is no /v2/agents/{id}/sessions route. Also
align getSession/listSessions to the same routes so SDK -> proxy
calls actually resolve. listSessions filters by agent_id via the
existing query param.
Daemon events:append was failing with MissingRequiredValueError because
the bare run_id FK and dict payload don't pass Prisma's input validation
on the LiteLLM_AgentRunEvent.create call.
The Prisma generated client wants {"connect": {"id": ...}} for the
session relation and prisma.Json(...) wrapping for the prompt + event
payload Json columns.
The generated Prisma client requires {"connect": {"id": ...}} for
relation fields and prisma.Json(...) for Json columns. Bare agent_id
strings and bare dict/list values both raise MissingRequiredValueError
on session create.
Prisma rejects bare None / dict / list for optional Json columns with
MissingRequiredValueError. Wrap dict/list values in prisma.Json(...) and
drop the key entirely when the source is None so the column resolves to
SQL NULL.
Reconciliation:
- adopt B's richer base.py types (ProvisionContext, VMHandle, AwsCreds,
Ec2Config, ProvisionError) as canonical; keep A's NoopVMProvider alias
and the registry helpers (register_vm_provider, reset_vm_provider_registry)
for tests.
- rename B's factory entry point from get_vm_provider to build_vm_provider so
it doesn't collide with the runtime registry's get_vm_provider(name).
- update A's session_endpoints._provision_in_background to construct a
ProvisionContext and call provider.provision(ctx); pass team_id from
the caller's API key.
- update _terminate_session_internal to construct a VMHandle when a vm_id
is recorded on the session row.
- extend B's NoopProvider with provision_calls/terminate_calls recording for
backward compat with A's tests; accept either VMHandle-style or legacy
keyword-style terminate args.
- de-duplicate LiteLLM_AgentVMConfig from all 3 schema.prisma files —
G's (LIT-2891) version wins; B's stub at the top is collapsed to a
comment pointing to G's section.
- delete B's 20260506220000_add_agent_vm_config migration (collides with
G's 20260506220000_add_cloud_agent_settings_tables which already creates
the table).
- rewrite team_config.py to read G's per-field encrypted columns
(aws_access_key_id_enc, aws_secret_access_key_enc, aws_region) instead
of B's single-blob aws_creds_enc; rewrite test_team_config.py to match.
Resolved schema conflicts in schema.prisma, litellm/proxy/schema.prisma,
and litellm-proxy-extras/litellm_proxy_extras/schema.prisma by unioning
A's agent/session/run/event tables with G's VM config / secrets /
worker / pairing-token tables. Both sets of tables now coexist.
Mock proxy now serializes responses using snake_case keys (agent_id,
created_at, system_prompt, run_id, etc.) and reads request bodies as
snake_case so it matches the real backend that the SDK now talks to via
the new transform layer. Also update the status string literals to the
new SessionStatus and RunStatus values, and read the followup body as
{prompt: {text}}.
There is no per-run conversation endpoint on the backend; conversation
history is session-scoped. Callers should use SessionHandle.conversation()
instead. Also update TERMINAL_STATES to the new RunStatus values
(finished/cancelled/error).
The backend's followup endpoint expects a FollowupCreate payload with a
nested prompt object ({prompt: {text: ...}}), not a flat {message: ...}.
Public method signature followup(message: string) is unchanged - only
the wire body changes.
Backend speaks snake_case (Python idiom) while the SDK's public TS API
is camelCase. Add recursive snakeToCamel and camelToSnake helpers and
wire them into the HTTP layer so request bodies are camel->snake before
JSON.stringify and response JSON is snake->camel before being returned
to callers. Single-word keys like id/type/data/seq/status pass through
unchanged in both directions, and non-object values are not touched.
Backend (LIT-2890) emits provisioning/ready/busy/error/terminated for
session status and queued/running/finished/cancelled/error for run status.
Update the SDK type aliases so callers compare against the actual values
the proxy returns over the wire.
Greptile P3 (regression coverage): the cascade filter in
delete_agent was changed from terminated_at is None to
status not in SESSION_TERMINAL_STATUSES in commit a0015e8564.
Add an explicit test that exercises the bug surface — a session
flipped to error (terminal) but with terminated_at deliberately
left None. The legacy filter would have re-terminated it; the
status-based filter must skip it.
Asserts:
* _is_seq_collision matches RuntimeError('event_seq_collision') and
prisma.errors.UniqueViolationError, but rejects unrelated errors.
* Real collision: insert at seq=N when seq=N already exists, retry
at N+1 succeeds (no 409 surfaced).
* Unrelated DB outage: a non-collision RuntimeError propagates as
itself, NOT misclassified as 409 event_seq_collision.
Greptile follow-up regression coverage.
Greptile P1 regression coverage. The sweeper test only checked
run.status == RUN_STATUS_ERROR, leaving the session-status gap
undetected. Add a force-busy step before the sweep + assert the
session transitions to ready after.
The seq-collision retry in daemon_append_event previously caught all
Exception types and re-raised them as 409 event_seq_collision. A
transient DB error during retry would surface to the daemon as a
misleading 409, hiding the real outage and the daemon would respond
incorrectly (they treat 409 as 'data conflict, drop the event' rather
than 'retry').
Add a narrow _is_seq_collision predicate that matches:
* prisma.errors.UniqueViolationError (production)
* RuntimeError('event_seq_collision') marker (test stand-in)
Other errors bubble up unchanged so callers can distinguish a real
seq conflict from a real outage.
Greptile (review #PRR_kwDOKALCgc78u_NS — overly broad exception
catch in seq-collision retry).
_sweep_stuck_runs marks idle-timeout runs as error but never called
refresh_session_status_from_runs, leaving the parent session
permanently busy. Every other run-terminal path (cancel_run,
daemon_append_event, /followup) calls the helper to flip
busy -> ready; the sweeper was the only path that skipped it.
After flipping each run to error, call refresh_session_status_from_runs
inside the loop so a session whose only active run was reaped here
transitions back to ready.
Greptile P1 (review #PRR_kwDOKALCgc78u_NS, inline comment line 172).
Walks the SDK-visible session.status across:
* POST /v2/sessions/{sid}/runs — ready -> busy
* POST /v2/sessions/{sid}/runs/{rid}/cancel — busy -> ready
* POST /v2/sessions/{sid}/followup — ready -> busy via /followup
* Cancellation via /followup-created run — busy -> ready
Plus two helper tests:
* idempotent (no-op when no transition is needed)
* quiet on missing session (race with cascade delete)
Greptile P1 regression coverage.
Greptile P1 regression coverage: dead-daemon sweep must route through
_terminate_session_internal so provider.terminate gets called. Without
this assertion the regression silently returned (NoopVMProvider would
still mark rows correctly via update_many).
delete_agent's cascade filter previously used 'terminated_at is None'
to find non-terminal sessions. The fix is safe in practice because
_terminate_session_internal has its own SESSION_TERMINAL_STATUSES guard,
but it's inconsistent with the rest of the module which uses
'status in/notin SESSION_TERMINAL_STATUSES' everywhere else.
Switch to the status-based check to match.
Greptile P3 (review #PRR_kwDOKALCgc78u9En).
_sweep_dead_daemons previously ran update_many directly on the session
row, skipping provider.terminate. With NoopVMProvider this was harmless,
but once Epic B swaps in a real VM provider it would orphan EC2
instances every time a daemon stopped heartbeating.
Mirror the pattern from _sweep_expired_sessions: call
_terminate_session_internal per-row so the provider is notified, then
explicitly downgrade status from 'terminated' to 'error' (both are
terminal — no further state transitions). Also drops the per-run
update loop since _terminate_session_internal already cancels active
runs and emits run_cancelled events.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).