Resolved schema conflicts in schema.prisma, litellm/proxy/schema.prisma,
and litellm-proxy-extras/litellm_proxy_extras/schema.prisma by unioning
A's agent/session/run/event tables with G's VM config / secrets /
worker / pairing-token tables. Both sets of tables now coexist.
Greptile P3 (regression coverage): the cascade filter in
delete_agent was changed from terminated_at is None to
status not in SESSION_TERMINAL_STATUSES in commit a0015e8564.
Add an explicit test that exercises the bug surface — a session
flipped to error (terminal) but with terminated_at deliberately
left None. The legacy filter would have re-terminated it; the
status-based filter must skip it.
Asserts:
* _is_seq_collision matches RuntimeError('event_seq_collision') and
prisma.errors.UniqueViolationError, but rejects unrelated errors.
* Real collision: insert at seq=N when seq=N already exists, retry
at N+1 succeeds (no 409 surfaced).
* Unrelated DB outage: a non-collision RuntimeError propagates as
itself, NOT misclassified as 409 event_seq_collision.
Greptile follow-up regression coverage.
Greptile P1 regression coverage. The sweeper test only checked
run.status == RUN_STATUS_ERROR, leaving the session-status gap
undetected. Add a force-busy step before the sweep + assert the
session transitions to ready after.
The seq-collision retry in daemon_append_event previously caught all
Exception types and re-raised them as 409 event_seq_collision. A
transient DB error during retry would surface to the daemon as a
misleading 409, hiding the real outage and the daemon would respond
incorrectly (they treat 409 as 'data conflict, drop the event' rather
than 'retry').
Add a narrow _is_seq_collision predicate that matches:
* prisma.errors.UniqueViolationError (production)
* RuntimeError('event_seq_collision') marker (test stand-in)
Other errors bubble up unchanged so callers can distinguish a real
seq conflict from a real outage.
Greptile (review #PRR_kwDOKALCgc78u_NS — overly broad exception
catch in seq-collision retry).
_sweep_stuck_runs marks idle-timeout runs as error but never called
refresh_session_status_from_runs, leaving the parent session
permanently busy. Every other run-terminal path (cancel_run,
daemon_append_event, /followup) calls the helper to flip
busy -> ready; the sweeper was the only path that skipped it.
After flipping each run to error, call refresh_session_status_from_runs
inside the loop so a session whose only active run was reaped here
transitions back to ready.
Greptile P1 (review #PRR_kwDOKALCgc78u_NS, inline comment line 172).
Walks the SDK-visible session.status across:
* POST /v2/sessions/{sid}/runs — ready -> busy
* POST /v2/sessions/{sid}/runs/{rid}/cancel — busy -> ready
* POST /v2/sessions/{sid}/followup — ready -> busy via /followup
* Cancellation via /followup-created run — busy -> ready
Plus two helper tests:
* idempotent (no-op when no transition is needed)
* quiet on missing session (race with cascade delete)
Greptile P1 regression coverage.
Greptile P1 regression coverage: dead-daemon sweep must route through
_terminate_session_internal so provider.terminate gets called. Without
this assertion the regression silently returned (NoopVMProvider would
still mark rows correctly via update_many).
delete_agent's cascade filter previously used 'terminated_at is None'
to find non-terminal sessions. The fix is safe in practice because
_terminate_session_internal has its own SESSION_TERMINAL_STATUSES guard,
but it's inconsistent with the rest of the module which uses
'status in/notin SESSION_TERMINAL_STATUSES' everywhere else.
Switch to the status-based check to match.
Greptile P3 (review #PRR_kwDOKALCgc78u9En).
_sweep_dead_daemons previously ran update_many directly on the session
row, skipping provider.terminate. With NoopVMProvider this was harmless,
but once Epic B swaps in a real VM provider it would orphan EC2
instances every time a daemon stopped heartbeating.
Mirror the pattern from _sweep_expired_sessions: call
_terminate_session_internal per-row so the provider is notified, then
explicitly downgrade status from 'terminated' to 'error' (both are
terminal — no further state transitions). Also drops the per-run
update loop since _terminate_session_internal already cancels active
runs and emits run_cancelled events.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
When /followup creates a fresh queued run (terminal-or-absent latest_run
branch), call refresh_session_status_from_runs so the session moves
'ready' -> 'busy'. Match the same hook added to POST /runs.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Two call sites added:
* After POST /v2/sessions/{sid}/runs creates a queued run, call
refresh_session_status_from_runs so the parent session moves
'ready' -> 'busy'.
* After POST /v2/sessions/{sid}/runs/{rid}/cancel marks a run
cancelled, call the same helper so the session moves
'busy' -> 'ready' if no other active runs exist.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Two call sites added:
* After _claim_next_queued_run flips a queued run to running, call
refresh_session_status_from_runs so a session that was 'ready'
transitions to 'busy'. Idempotent for already-busy sessions.
* After daemon_append_event finalizes a terminal event (run_finished /
run_cancelled / run_error), call refresh_session_status_from_runs so
a session with no remaining active runs transitions back to 'ready'.
Without these hooks, session.status stayed permanently 'busy' after the
first run started — clients polling GET /v2/sessions/{id} always saw
'busy' regardless of run state.
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Pure-function logic for the session status state machine already lives
in state_machine.derive_session_status_from_runs. This module is the
I/O wrapper that reads the session row, counts active runs, and
persists the new status only when the helper says it should change.
Used by every code path that flips a run's status:
* POST /v2/sessions/{sid}/runs (queued -> session busy)
* POST /v2/sessions/{sid}/followup (new run -> same)
* GET /v2/sessions/{sid}/runs/next/internal/poll (queued -> running)
* POST /v2/sessions/{sid}/runs/{rid}/cancel (terminal -> ready)
* POST /v2/sessions/{sid}/runs/{rid}/events:append (terminal -> ready)
Greptile P1 (review #PRR_kwDOKALCgc78u9En).
Reproduces the original race deterministically: a session has both
an OLDER active run AND a NEWER terminal run. The buggy code path
used latest_run (newest by created_at) to decide whether to fall
through to the create-new-run branch — and since the newest is
terminal, it skipped the busy check and would have inserted a
duplicate run.
Three tests:
* 409 run_busy when an older active run exists.
* Happy path: empty session creates the first run.
* Happy path: latest run is active so /followup appends a
user_message event (unchanged by the busy guard).
Greptile P1 regression coverage (validation #16).
Exercise every state-mutating endpoint as PROXY_ADMIN_VIEW_ONLY and
confirm they all return 403:
* POST /v2/agents
* PATCH /v2/agents/{id}
* DELETE /v2/agents/{id}
* POST /v2/sessions
* DELETE /v2/sessions/{id}
* POST /v2/sessions/{id}/runs
* POST /v2/sessions/{id}/runs/{rid}/cancel
* POST /v2/sessions/{id}/followup
Plus a happy-path test confirming view-only admins still get
cross-tenant READ access (the whole point of the role).
Greptile P1 SECURITY regression coverage (validation #14).
Asserts:
* _get_signing_secret raises AgentJWTSecretNotConfiguredError when
LITELLM_AGENT_JWT_SECRET is unset, even with LITELLM_MASTER_KEY set
(no silent fallback).
* is_agent_jwt_secret_configured returns False for unset / empty,
True for any non-empty value.
* mint_daemon_token and decode_daemon_token both surface the same
error when the env var is missing.
Greptile P1 SECURITY regression coverage (validation #15).
When LITELLM_AGENT_JWT_SECRET is not set, the daemon JWT auth layer
cannot operate safely (no master-key fallback). Refuse to mount the
four /v2/agents+/v2/sessions routers in that case and log a clear
error pointing operators at the env var. Mounting them anyway would
expose an auth surface that can never validate a token and crash on
every request.
Greptile P1 SECURITY follow-up.
asyncio.get_event_loop() is deprecated inside a running coroutine
(Python 3.10+). Replace the two call sites in
daemon_get_next_queued_run with asyncio.get_running_loop().
Greptile P2.
Two fixes in this file:
1. asyncio.get_event_loop() is deprecated inside a running coroutine
(Python 3.10+). Replace both call sites in _stream_run_events with
asyncio.get_running_loop().
2. Call assert_caller_can_mutate on create_run and cancel_run so
view-only admins get 403 instead of bypassing ownership.
Greptile P2 (deprecation) + P1 SECURITY (view-only admin write bypass).
Two fixes in this file:
1. /followup new-run branch was bypassing the run-busy concurrency
guard. When latest_run was terminal-or-absent, the endpoint went
straight to litellm_agentrun.create without calling
_has_active_run. Two concurrent /followup calls on an idle session
both passed the latest_run.status check and both inserted runs,
breaking the 'one active run per session' invariant that POST /runs
enforces via 409 run_busy. Add the same _has_active_run check
before the fallthrough create, plus a defensive insert-time retry
that re-queries for active runs after IntegrityError and surfaces
409 run_busy if it lost the race.
2. Call assert_caller_can_mutate on create_session, delete_session,
and followup so view-only admins get 403 instead of bypassing
ownership.
Greptile P1 (concurrency) + P1 SECURITY (view-only admin write bypass).
Block PROXY_ADMIN_VIEW_ONLY from create_agent / update_agent /
delete_agent. Reads (GET) still pass through is_proxy_admin_read
so view-only admins keep cross-tenant visibility for the support UI.
Greptile P1 SECURITY follow-up (view-only admin write bypass).
is_proxy_admin previously returned True for both PROXY_ADMIN and
PROXY_ADMIN_VIEW_ONLY, letting view-only admins skip
assert_caller_owns_agent / assert_caller_owns_session on every write
endpoint and create / update / delete other tenants' agents,
sessions, and runs.
Split the helpers:
* is_proxy_admin: now full-admin only (used for write paths via the
fall-through to per-tenant ownership; view-only fails and gets 404).
* is_proxy_admin_read: full + view-only, used on read paths so the
support UI can still render any tenant's resources.
* assert_caller_can_mutate: explicit 403 guard for view-only on
every state-mutating endpoint.
The mutating endpoints in agent/session/run files call
assert_caller_can_mutate before any DB write — see follow-up commits.
Greptile P1 SECURITY (review #PRR_kwDOKALCgc78uM7F).
The daemon JWT secret must be a SEPARATE credential from the proxy
master key. The previous fallback to LITELLM_MASTER_KEY conflated
two distinct auth surfaces — a captured daemon JWT could be used
to mint regular API keys with master-key authority.
Replace _get_signing_secret with a strict check that raises
AgentJWTSecretNotConfiguredError if the dedicated env var is unset.
Add is_agent_jwt_secret_configured() so proxy_server.py can refuse
to mount the routers when the secret is missing.
Greptile P1 SECURITY (review #PRR_kwDOKALCgc78uM7F).
LIT-2890 / B2's long-poll heartbeat calls find_worker_by_jwt on every
request, which filters by worker_jwt_hash. Without an index that's a
full table scan per heartbeat once the worker pool grows past a handful
of rows. Adding the index now so it lands with the table create rather
than as a follow-up migration during B2 ramp-up.
Covers all four resolution paths and the security-critical default:
* LITELLM_CLOUD_AGENT_PROXY_BASE_URL wins regardless of forwarded
headers (and trailing slashes get stripped).
* X-Forwarded-Host is IGNORED unless LITELLM_TRUST_PROXY_HEADERS=1.
A forged header pointing at attacker.example must not show up.
* Trust-flag opt-in honors X-Forwarded-Host + X-Forwarded-Proto,
but x-forwarded-proto alone (no host) doesn't poison the response.
* Last-resort fallback to localhost:4000 when no Host header.
Resolution order for the URL embedded in the install one-liner is now:
1. LITELLM_CLOUD_AGENT_PROXY_BASE_URL env var (operator-configured,
fully trusted) — recommended for production.
2. X-Forwarded-Host / X-Forwarded-Proto, ONLY when the operator
opts in via LITELLM_TRUST_PROXY_HEADERS=1.
3. The request's direct Host header — safe by default because it
reflects the actual TCP destination, not an attacker-supplied hop.
Previously any authenticated caller could forge X-Forwarded-Host to
embed an attacker-controlled URL in the install command. If a second
operator ran that command, the worker would send its raw pair token to
the attacker's host, who could then call POST /v2/agent-workers/register
and gain a long-lived worker JWT.
Also adds structured logging on /register failures (invalid / replayed
/ expired tokens) so operators running the proxy behind a WAF / fail2ban
can detect abuse at the network layer (the proxy itself doesn't ship a
built-in per-IP limiter).
Locks the production-safe default for LITELLM_CLOUD_AGENT_MOCK_AWS so a
future revert to the unsafe "1" default trips CI. Also covers the
strict-string parsing — typos like "true" / "yes" must NOT silently
flip on the mock.
Previously the test-connection endpoint defaulted to mock-on, so a fresh
production proxy would silently return a synthetic success for any
non-empty AWS access key. Operators saving incorrect credentials would
only discover the failure later when VMs failed to launch.
Default is now "0" — operators must set LITELLM_CLOUD_AGENT_MOCK_AWS=1
explicitly to opt into the mock path during local development.
Also adds an inline comment on _build_update_payload's `is not None`
guard so a future contributor doesn't silently drop `False` / `0`
updates by switching to truthy comparison.
Adds three regression cases that fail if a future change drops the
shlex.quote() pass on proxy_url, raw_token, or install_script_url. The
existing simple-input cases still pass unchanged because shlex.quote
returns alnum/colon/slash/dot strings verbatim.
shlex.quote() the install_script_url, proxy_url, and raw_token before
interpolating into the curl-pipe-sh one-liner. Without quoting, a
proxy_url containing spaces or shell metacharacters (e.g. via a misconfig
or the X-Forwarded-Host issue Greptile also flagged) could produce a
malformed or exploitable command on the worker box.