Two fixes in this file:
1. /followup new-run branch was bypassing the run-busy concurrency
guard. When latest_run was terminal-or-absent, the endpoint went
straight to litellm_agentrun.create without calling
_has_active_run. Two concurrent /followup calls on an idle session
both passed the latest_run.status check and both inserted runs,
breaking the 'one active run per session' invariant that POST /runs
enforces via 409 run_busy. Add the same _has_active_run check
before the fallthrough create, plus a defensive insert-time retry
that re-queries for active runs after IntegrityError and surfaces
409 run_busy if it lost the race.
2. Call assert_caller_can_mutate on create_session, delete_session,
and followup so view-only admins get 403 instead of bypassing
ownership.
Greptile P1 (concurrency) + P1 SECURITY (view-only admin write bypass).
Block PROXY_ADMIN_VIEW_ONLY from create_agent / update_agent /
delete_agent. Reads (GET) still pass through is_proxy_admin_read
so view-only admins keep cross-tenant visibility for the support UI.
Greptile P1 SECURITY follow-up (view-only admin write bypass).
is_proxy_admin previously returned True for both PROXY_ADMIN and
PROXY_ADMIN_VIEW_ONLY, letting view-only admins skip
assert_caller_owns_agent / assert_caller_owns_session on every write
endpoint and create / update / delete other tenants' agents,
sessions, and runs.
Split the helpers:
* is_proxy_admin: now full-admin only (used for write paths via the
fall-through to per-tenant ownership; view-only fails and gets 404).
* is_proxy_admin_read: full + view-only, used on read paths so the
support UI can still render any tenant's resources.
* assert_caller_can_mutate: explicit 403 guard for view-only on
every state-mutating endpoint.
The mutating endpoints in agent/session/run files call
assert_caller_can_mutate before any DB write — see follow-up commits.
Greptile P1 SECURITY (review #PRR_kwDOKALCgc78uM7F).
The daemon JWT secret must be a SEPARATE credential from the proxy
master key. The previous fallback to LITELLM_MASTER_KEY conflated
two distinct auth surfaces — a captured daemon JWT could be used
to mint regular API keys with master-key authority.
Replace _get_signing_secret with a strict check that raises
AgentJWTSecretNotConfiguredError if the dedicated env var is unset.
Add is_agent_jwt_secret_configured() so proxy_server.py can refuse
to mount the routers when the secret is missing.
Greptile P1 SECURITY (review #PRR_kwDOKALCgc78uM7F).
Add 4 new tables for the agent_session_endpoints module (Cursor SDK on LiteLLM):
- LiteLLM_Agent: agent definition (model, system prompt, default repos)
- LiteLLM_AgentSession: VM-backed conversation owned by an agent
- LiteLLM_AgentRun: single turn within a session
- LiteLLM_AgentRunEvent: append-only event log for resumable SSE
Includes cascade deletes, idempotency unique constraints, and indexes
for the cleanup sweeper and ownership lookups.
* fix(auth): pass team_id in member-level model access check
_check_team_member_model_access calls _can_object_call_model without
team_id, so access groups defined via model_info.access_groups cannot
resolve for team-scoped DB models (their internal router name is
model_name_<team>_<uuid>, not the public name). The team-level check
already passes team_id; this mirrors that.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(auth): add tests for member-level access group resolution with team_id
Eight tests covering _can_object_call_model and
_check_team_member_model_access with team-scoped DB models:
- access group resolves when team_id is passed
- access group fails without team_id (pre-fix behavior)
- literal model name still works with team_id (no regression)
- denied model still denied with team_id
- second model in group also reachable
- end-to-end member access via access group (mocked membership)
- end-to-end member denied for model not in allowed list
- no-override member inherits team-level check
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* proxy: hot-reload config YAML when --reload is set
Uvicorn's --reload only watches *.py by default, so editing the
--config YAML did not restart the proxy. _get_reload_options() now
extends reload_dirs/reload_includes with the config file's directory
and basename when --config is provided.
* proxy: qualify reload_includes with absolute config path
Address Greptile review on PR #27274. When the --config file lives
outside cwd, reload_includes previously stored only the basename, which
meant uvicorn/watchfiles would also reload on edits to any same-named
file inside cwd. Use the absolute config path as the include pattern in
that case so only the actual proxy config triggers a restart.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(proxy): use basename for reload_includes config pattern
Uvicorn's resolve_reload_patterns() calls pathlib.Path.glob(), which
raises NotImplementedError on absolute patterns (uvicorn discussion
2156). Passing config_abs (an absolute path) when the config file lived
outside cwd crashed startup under --reload. The config_dir is already
added to reload_dirs, so using just the basename as the include pattern
is sufficient to match the specific config file.
* fix: make it reload app when yaml changes
* style: remove unneeded comments
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* Include model name + configured TPM/RPM in priority rate-limit 429 errors (#27215)
* Include model name + configured TPM/RPM in priority rate-limit 429 errors
The current 429 message ('Priority-based rate limit exceeded. Priority: prod,
Rate limit type: tokens, Remaining: -664145, Model saturation: 86.3%') doesn't
tell the operator which model was hit or what the configured limit is, so they
can't tell whether the priority allocation needs tuning or the model TPM is
just too small.
Add Model, Model TPM, and Model RPM to both the priority-based 429 and the
sibling Model-capacity 429 in dynamic_rate_limiter_v3._check_rate_limits.
Pure error-message change — no behavior or schema impact.
* test: assert priority 429 includes model name + configured TPM/RPM
Adds a regression test for the new fields in the priority-based 429 detail
('Model:', 'Model TPM:', 'Model RPM:'). Verified locally that the test
fails against the unpatched dynamic_rate_limiter_v3.py and passes after
the patch.
---------
Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>
* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>