Reconciliation:
- adopt B's richer base.py types (ProvisionContext, VMHandle, AwsCreds,
Ec2Config, ProvisionError) as canonical; keep A's NoopVMProvider alias
and the registry helpers (register_vm_provider, reset_vm_provider_registry)
for tests.
- rename B's factory entry point from get_vm_provider to build_vm_provider so
it doesn't collide with the runtime registry's get_vm_provider(name).
- update A's session_endpoints._provision_in_background to construct a
ProvisionContext and call provider.provision(ctx); pass team_id from
the caller's API key.
- update _terminate_session_internal to construct a VMHandle when a vm_id
is recorded on the session row.
- extend B's NoopProvider with provision_calls/terminate_calls recording for
backward compat with A's tests; accept either VMHandle-style or legacy
keyword-style terminate args.
- de-duplicate LiteLLM_AgentVMConfig from all 3 schema.prisma files —
G's (LIT-2891) version wins; B's stub at the top is collapsed to a
comment pointing to G's section.
- delete B's 20260506220000_add_agent_vm_config migration (collides with
G's 20260506220000_add_cloud_agent_settings_tables which already creates
the table).
- rewrite team_config.py to read G's per-field encrypted columns
(aws_access_key_id_enc, aws_secret_access_key_enc, aws_region) instead
of B's single-blob aws_creds_enc; rewrite test_team_config.py to match.
Resolved schema conflicts in schema.prisma, litellm/proxy/schema.prisma,
and litellm-proxy-extras/litellm_proxy_extras/schema.prisma by unioning
A's agent/session/run/event tables with G's VM config / secrets /
worker / pairing-token tables. Both sets of tables now coexist.
Greptile P3 (regression coverage): the cascade filter in
delete_agent was changed from terminated_at is None to
status not in SESSION_TERMINAL_STATUSES in commit a0015e8564.
Add an explicit test that exercises the bug surface — a session
flipped to error (terminal) but with terminated_at deliberately
left None. The legacy filter would have re-terminated it; the
status-based filter must skip it.
Asserts:
* _is_seq_collision matches RuntimeError('event_seq_collision') and
prisma.errors.UniqueViolationError, but rejects unrelated errors.
* Real collision: insert at seq=N when seq=N already exists, retry
at N+1 succeeds (no 409 surfaced).
* Unrelated DB outage: a non-collision RuntimeError propagates as
itself, NOT misclassified as 409 event_seq_collision.
Greptile follow-up regression coverage.
Greptile P1 regression coverage. The sweeper test only checked
run.status == RUN_STATUS_ERROR, leaving the session-status gap
undetected. Add a force-busy step before the sweep + assert the
session transitions to ready after.
The seq-collision retry in daemon_append_event previously caught all
Exception types and re-raised them as 409 event_seq_collision. A
transient DB error during retry would surface to the daemon as a
misleading 409, hiding the real outage and the daemon would respond
incorrectly (they treat 409 as 'data conflict, drop the event' rather
than 'retry').
Add a narrow _is_seq_collision predicate that matches:
* prisma.errors.UniqueViolationError (production)
* RuntimeError('event_seq_collision') marker (test stand-in)
Other errors bubble up unchanged so callers can distinguish a real
seq conflict from a real outage.
Greptile (review #PRR_kwDOKALCgc78u_NS — overly broad exception
catch in seq-collision retry).
_sweep_stuck_runs marks idle-timeout runs as error but never called
refresh_session_status_from_runs, leaving the parent session
permanently busy. Every other run-terminal path (cancel_run,
daemon_append_event, /followup) calls the helper to flip
busy -> ready; the sweeper was the only path that skipped it.
After flipping each run to error, call refresh_session_status_from_runs
inside the loop so a session whose only active run was reaped here
transitions back to ready.
Greptile P1 (review #PRR_kwDOKALCgc78u_NS, inline comment line 172).
Walks the SDK-visible session.status across:
* POST /v2/sessions/{sid}/runs — ready -> busy
* POST /v2/sessions/{sid}/runs/{rid}/cancel — busy -> ready
* POST /v2/sessions/{sid}/followup — ready -> busy via /followup
* Cancellation via /followup-created run — busy -> ready
Plus two helper tests:
* idempotent (no-op when no transition is needed)
* quiet on missing session (race with cascade delete)
Greptile P1 regression coverage.
Greptile P1 regression coverage: dead-daemon sweep must route through
_terminate_session_internal so provider.terminate gets called. Without
this assertion the regression silently returned (NoopVMProvider would
still mark rows correctly via update_many).
delete_agent's cascade filter previously used 'terminated_at is None'
to find non-terminal sessions. The fix is safe in practice because
_terminate_session_internal has its own SESSION_TERMINAL_STATUSES guard,
but it's inconsistent with the rest of the module which uses
'status in/notin SESSION_TERMINAL_STATUSES' everywhere else.
Switch to the status-based check to match.
Greptile P3 (review #PRR_kwDOKALCgc78u9En).
_sweep_dead_daemons previously ran update_many directly on the session
row, skipping provider.terminate. With NoopVMProvider this was harmless,
but once Epic B swaps in a real VM provider it would orphan EC2
instances every time a daemon stopped heartbeating.
Mirror the pattern from _sweep_expired_sessions: call
_terminate_session_internal per-row so the provider is notified, then
explicitly downgrade status from 'terminated' to 'error' (both are
terminal — no further state transitions). Also drops the per-run
update loop since _terminate_session_internal already cancels active
runs and emits run_cancelled events.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
When /followup creates a fresh queued run (terminal-or-absent latest_run
branch), call refresh_session_status_from_runs so the session moves
'ready' -> 'busy'. Match the same hook added to POST /runs.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Two call sites added:
* After POST /v2/sessions/{sid}/runs creates a queued run, call
refresh_session_status_from_runs so the parent session moves
'ready' -> 'busy'.
* After POST /v2/sessions/{sid}/runs/{rid}/cancel marks a run
cancelled, call the same helper so the session moves
'busy' -> 'ready' if no other active runs exist.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Two call sites added:
* After _claim_next_queued_run flips a queued run to running, call
refresh_session_status_from_runs so a session that was 'ready'
transitions to 'busy'. Idempotent for already-busy sessions.
* After daemon_append_event finalizes a terminal event (run_finished /
run_cancelled / run_error), call refresh_session_status_from_runs so
a session with no remaining active runs transitions back to 'ready'.
Without these hooks, session.status stayed permanently 'busy' after the
first run started — clients polling GET /v2/sessions/{id} always saw
'busy' regardless of run state.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Pure-function logic for the session status state machine already lives
in state_machine.derive_session_status_from_runs. This module is the
I/O wrapper that reads the session row, counts active runs, and
persists the new status only when the helper says it should change.
Used by every code path that flips a run's status:
* POST /v2/sessions/{sid}/runs (queued -> session busy)
* POST /v2/sessions/{sid}/followup (new run -> same)
* GET /v2/sessions/{sid}/runs/next/internal/poll (queued -> running)
* POST /v2/sessions/{sid}/runs/{rid}/cancel (terminal -> ready)
* POST /v2/sessions/{sid}/runs/{rid}/events:append (terminal -> ready)
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Reproduces the original race deterministically: a session has both
an OLDER active run AND a NEWER terminal run. The buggy code path
used latest_run (newest by created_at) to decide whether to fall
through to the create-new-run branch — and since the newest is
terminal, it skipped the busy check and would have inserted a
duplicate run.
Three tests:
* 409 run_busy when an older active run exists.
* Happy path: empty session creates the first run.
* Happy path: latest run is active so /followup appends a
user_message event (unchanged by the busy guard).
Greptile P1 regression coverage (validation #16).
Exercise every state-mutating endpoint as PROXY_ADMIN_VIEW_ONLY and
confirm they all return 403:
* POST /v2/agents
* PATCH /v2/agents/{id}
* DELETE /v2/agents/{id}
* POST /v2/sessions
* DELETE /v2/sessions/{id}
* POST /v2/sessions/{id}/runs
* POST /v2/sessions/{id}/runs/{rid}/cancel
* POST /v2/sessions/{id}/followup
Plus a happy-path test confirming view-only admins still get
cross-tenant READ access (the whole point of the role).
Greptile P1 SECURITY regression coverage (validation #14).
Asserts:
* _get_signing_secret raises AgentJWTSecretNotConfiguredError when
LITELLM_AGENT_JWT_SECRET is unset, even with LITELLM_MASTER_KEY set
(no silent fallback).
* is_agent_jwt_secret_configured returns False for unset / empty,
True for any non-empty value.
* mint_daemon_token and decode_daemon_token both surface the same
error when the env var is missing.
Greptile P1 SECURITY regression coverage (validation #15).
When LITELLM_AGENT_JWT_SECRET is not set, the daemon JWT auth layer
cannot operate safely (no master-key fallback). Refuse to mount the
four /v2/agents+/v2/sessions routers in that case and log a clear
error pointing operators at the env var. Mounting them anyway would
expose an auth surface that can never validate a token and crash on
every request.
Greptile P1 SECURITY follow-up.
asyncio.get_event_loop() is deprecated inside a running coroutine
(Python 3.10+). Replace the two call sites in
daemon_get_next_queued_run with asyncio.get_running_loop().
Greptile P2.
Two fixes in this file:
1. asyncio.get_event_loop() is deprecated inside a running coroutine
(Python 3.10+). Replace both call sites in _stream_run_events with
asyncio.get_running_loop().
2. Call assert_caller_can_mutate on create_run and cancel_run so
view-only admins get 403 instead of bypassing ownership.
Greptile P2 (deprecation) + P1 SECURITY (view-only admin write bypass).
Two fixes in this file:
1. /followup new-run branch was bypassing the run-busy concurrency
guard. When latest_run was terminal-or-absent, the endpoint went
straight to litellm_agentrun.create without calling
_has_active_run. Two concurrent /followup calls on an idle session
both passed the latest_run.status check and both inserted runs,
breaking the 'one active run per session' invariant that POST /runs
enforces via 409 run_busy. Add the same _has_active_run check
before the fallthrough create, plus a defensive insert-time retry
that re-queries for active runs after IntegrityError and surfaces
409 run_busy if it lost the race.
2. Call assert_caller_can_mutate on create_session, delete_session,
and followup so view-only admins get 403 instead of bypassing
ownership.
Greptile P1 (concurrency) + P1 SECURITY (view-only admin write bypass).