Hiding the tab behind all_admin_roles took it away from org admins, who are
entitled to it: /v2/team/list?status=deleted returns 200 for them, scoped to
their own organizations. An org admin is an organization membership rather
than a global role, so their session carries user_role "internal_user" and no
role-based gate can ever see them.
Lift the membership lookup the left nav already did into a shared
useIsOrgAdmin hook, and let a capability opt into allowing org admins.
viewDeletedTeams is the only one that opts in; the backend still refuses org
admins on /v1/tool/list, /policies/list, /prompts/list and /audit, so those
gates stay as they are. The hook also accepts a session role of org_admin, in
case a deployment maps one through SSO.
* fix(bedrock): reject Anthropic server-side web_search tool with actionable error
Bedrock's Anthropic Messages endpoints cannot execute Anthropic's
server-side web_search tool, so forwarding it returns an opaque
"The provided request is not valid" 400 from Bedrock. Fail fast in the
invoke transform with an error that names the unsupported tool, the
model, and links the web search interception docs as the fix.
* refactor(bedrock): address review nits on web_search guard typing
Drops Any-typed router plumbing, immutable bucketing, generated dashboard API types, and adds endpoint plus resolution-fallback tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The proxy documentation lives in BerriAI/litellm-docs and any new env
key flagged by os.getenv() must be added there before the
test_env_keys.py CI check passes. Rather than fork the docs repo for
two niche tunables, hard-code the defaults:
- DEFAULT_DEPRECATION_WARN_DAYS = 30 (already overridable per-request
via ?warn_within_days=N on /model/deprecations).
- DEFAULT_DEPRECATION_CHECK_INTERVAL_SECONDS = 24h.
Both can still be raised as env-var follow-ups together with their docs
update if operators ask for it.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Surfaces deprecation_date metadata that is already shipped in
model_prices_and_context_window.json so operators get lead time to
migrate before a provider sunsets a model.
- New helper litellm.proxy.common_utils.model_deprecation classifies the
router's configured models into deprecated / imminent / upcoming
buckets. Resolution order: explicit model_info.deprecation_date >
model_info.base_model > litellm_params.model.
- New GET /model/deprecations (and /v1/model/deprecations) endpoint
returns a ModelDeprecationResponse, gated by user_api_key_auth.
- New AlertType.model_deprecation_warnings (in DEFAULT_ALERT_TYPES) plus
SlackAlerting.send_model_deprecation_alert dispatches a Slack message
for deprecated/imminent models. Severity is High when any model is
already past its date, Medium when only imminent.
- ProxyLogging.startup_event schedules a daily background task
(_run_scheduled_deprecation_check) when the alert type is enabled. The
interval is configurable via LITELLM_MODEL_DEPRECATION_CHECK_INTERVAL
and the warn window via LITELLM_MODEL_DEPRECATION_WARN_DAYS.
- Tests: 16 unit tests for the helper plus 4 for the Slack hook in
tests/test_litellm/.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
knip derives its entry points from vitest's `test.include`, which does not
cover `test.typecheck.include`, so the new `*.test-d.tsx` file read as an
unused file and failed the lint job. Declare the glob as an entry point.
Also drops the doc comment on `DataTableResolvedProps`; the rationale for
the resolved/public split belongs in the commit that introduced it.
* fix(proxy): treat SAML as configured in UI SSO detection
_has_user_setup_sso only checked OAuth client IDs, so SAML-only setups
left /.well-known/litellm-ui-config sso_configured=false and the login
button gray even when SAML IdP metadata was set. Include
SAML_IDP_METADATA_URL / SAML_IDP_METADATA_XML so UI discovery matches
the login redirect path.
* chore: adhere to comment policy
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix: apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
The removed comments restated the test names and the assertions directly
below them. The reasoning they carried is already recorded in the commit
that introduced the fix and in the pull request body.
Workflow Runs, Memory and Guardrails Monitor were visible to every role
while their page-load routes are proxy-admin-only, so a non-admin got a
page shell and a 401. Cost Optimization was half-broken the same way: its
Overall charts run on /user/daily/activity, which every role may call, but
tool spend, prompt caching, prompt compression and auto-router benchmarks
are all proxy-admin-only.
Add viewWorkflowRuns, viewMemory, viewGuardrailUsage and
viewProxyWideCostData, each gating the nav entry, the page and the request
together. The first three hide their page, including the direct-URL path,
since nothing on them works for a non-admin. Cost Optimization keeps its
page and drops only the parts a non-admin cannot read.
Gating both Agentic children left roles with no visible child rendering the
parent as a leaf link to ?page=agentic, which is not a route, so a parent
whose children are all filtered out is now dropped.
Role lists follow what the proxy actually grants: proxy_admin and
proxy_admin_viewer are served, and org admins are not, because
_user_is_org_admin needs an organization_id that a page-load GET never
carries.
* feat(ui): show vector store indexes on the Vector Stores page
Adds a proxy-admin-only Indexes tab listing rows from GET /v1/indexes:
index name, backing vector store, provider index, creator, and created
date. The tab is hidden for non proxy-admin roles to match the
endpoint's gate, and data loads lazily on first visit.
* feat(ui): link index rows to their vector store and creator
Vector Store cells open the store's info view when the name resolves to
a registered store, and Created By cells deep link to the users page via
a new userDetailHref, with the users page reading the user query param
through nuqs so the link is shareable.
* feat(ui): link docs and note supported providers on Indexes tab
* fix(ui): show not-found state instead of infinite loading for missing vector store
DataTable accepted any mix of its 40-odd props and rejected the incoherent
combinations at runtime, from a validator that threw during the first render.
A caller only found out it had wired server sorting without a `sorting` prop
when the page blew up in front of them.
Split the public prop type into mode-keyed unions instead, so the compiler
rejects those combinations at the call site. `validateDataTableConfig` and
`DataTableConfigError` go away; the component body reads an unchanged flat
`DataTableResolvedProps`, which every union member is assignable to, so there
is no narrowing inside it.
All 44 existing call sites typecheck against the new union unchanged, which
`next build` covers. That build only typechecks the app module graph, so the
prop type itself needed a gate of its own: `npm run test:types` runs vitest's
typecheck mode over `*.test-d.tsx`, and the unit workflow now runs it. The
four guards deleted from `DataTable.test.tsx` come back there as compile-time
assertions, and loosening the union back to the flat shape fails all five.
The Old Usage nav entry carried no role restriction, so every role saw it
and the page immediately fired eight /global/spend/* requests that the
proxy withholds from non-admins, producing a wall of 401s.
Gate the nav entry, the page, and both of its mount effects behind a
single viewGlobalSpend capability scoped to proxy_admin and
proxy_admin_viewer, matching what the backend actually serves.
Also drop the session JWT that adminspendByProvider put in the
/global/spend/provider query string; the handler never read it.
* fix(scripts): stop type-discipline checker reading Literal strings as forward refs
The checker re-parsed every string constant inside an annotation as a
forward reference, so Literal["list"] was counted as the mutable list
type. Skip Literal subtrees and ratchet the LIT001 ceiling down to the
corrected count.
* fix(proxy): keep lazy openapi snapshot fragments for transitively imported features
generate_snapshot skipped register_fn for any feature module already in
sys.modules, so a module pulled in transitively by an earlier feature
never mounted its routes and its fragment silently vanished on regen
(vector_store_management). Route collection also matched path_prefixes
only, dropping suffix-matched routes from fragments. Register every
feature and collect routes with feat.matches, mirroring the runtime
loader.
* feat(proxy): add GET /v1/indexes to list vector store indexes
/v1/indexes was POST-only, so indexes created through it could never be
viewed again. Add an admin-only list endpoint returning the stored index
rows newest first, fix the stale index_create docstring curl, and
regenerate the lazy openapi snapshot and dashboard schema types.
* chore(proxy): defer lazy openapi snapshot catch-up regen to a follow-up
Reverts _lazy_openapi_snapshot.json and schema.d.ts to the staging
versions. The snapshot was months stale, so regenerating it here buried
the actual change under ten thousand generated lines. A follow-up will
land the regen together with CI enforcement that keeps the snapshot
current. Until then GET /v1/indexes is served but absent from the
dashboard's generated types, which the UI step needs anyway.
* fix(proxy): use Annotated dependency to avoid new B008 violation
The Virtual Keys table and the Logs page team filter both asked for every
team on the proxy, which /v2/team/list and /team/list reject with a 401 for
any role below proxy admin or org admin. Both endpoints answer the same
request with the caller's own teams when it carries a user_id, so send one.
Only the two unscoped call sites change. The remaining callers either
already role-branch or render on surfaces gated to roles the endpoints
answer broadly, and scoping those would shrink the list they see: a proxy
admin scoped to their own id gets nothing back, and an org admin scoped on
/team/list loses the org teams they administer but do not belong to.
The shared helper reads the display-form session role rather than
all_admin_roles, which mixes display labels with raw role names and so does
not match the "Org Admin" value the dashboard actually holds.
The runbook still read as a fully manual flow: dispatch the publish workflow by
hand, then approve a second gate in the mirror repo. Neither is true now.
project-releaser checks the provider changelog on every release except adhoc,
nightly included, and dispatches the publish itself when the topmost released
heading has moved ahead of the mirror's tags, so cutting the version heading is
what ships the provider. The mirror's own release workflow no longer gates,
leaving one approval in project-releaser.
Ports the daily cron VM publisher from the unmerged tests/claude_code
checkout so the automation runs the e2e suite from litellm_internal_staging.
Adds find_regressions to matrix_builder for the green to red auto-merge
gate, pins the cron venv to Python 3.12, and ships the systemd units, env
template, and runbook alongside
* fix(reset_budget_job): advance budget_reset_at atomically with the spend cascade
A postgres timeout mid-cascade previously left LiteLLM_BudgetTable rows
stamped for the next window while team member, enduser, org and tag spend
stayed at cap, so every later tick skipped them until the window rolled
over. All cascade writes and the budget_reset_at advance now share one
prisma batch transaction; a failed run persists nothing and the rows stay
due for the next ~10 minute tick. Cache and counter invalidation runs only
after commit, and the catch-all enduser log line now names the cascade.
* fix(reset_budget_job): elect one runner per tick and chunk the reset scans
Every pod and worker previously ran the reset job every ~10 minutes,
each fetching every expired row with no limit and writing one giant
transaction at the same calendar-aligned boundary; that concurrency is
what piled up postgres lock contention and timeouts. The job now takes
the shared PodLockManager redis lock (no redis keeps the old behavior),
and each phase walks its due rows in 500-row chunks, one transaction per
chunk, stopping when a chunk is short, makes no forward progress, or
hits the per-run cap; leftovers wait for the next tick.
* chore(lint): ratchet budget ceilings down for fixed violations
* fix(reset_budget_job): harden chunk loop, fail open on redis errors, heartbeat the lock
Review fixes on the two prior commits. Reset scans now skip rows with no
budget_duration, so permanently due rows can neither starve a phase nor
have a lifetime cap zeroed every tick. Chunk progress counts rows whose
new budget_reset_at actually cleared the cutoff, so a zero-length
duration cannot burn the per-run chunk cap. A failed lock acquire only
skips the run when another pod verifiably holds the lock; a broken redis
runs unguarded instead of silently disabling resets fleet-wide. Partial
row failures report real progress and fire the failure hook without
killing the phase. The leader re-asserts the lock between phases and
stops if another pod took over, and the budget window advance uses
update_many so a tier deleted mid-chunk cannot abort the transaction.
Lint budget ceilings re-ratcheted for the net-fixed violations.
* fix(reset_budget_job): renew the leader lease and reject non-positive budget durations
Bot review follow-ups. PodLockManager now extends the lock TTL when the
holding pod re-acquires, via an atomic compare-and-expire script with a
plain SET fallback, so a run longer than the TTL keeps its lease instead
of silently sharing the job with another pod. The positive-duration
validation that team member endpoints already had is hoisted to
management common_utils and applied to key, internal user, budget,
customer and team intake, so a tenant can no longer create zero-duration
budgets whose permanently due rows starve other tenants' resets. Such
durations now return 400 at intake; existing rows are untouched.
* refactor(reset_budget_job): defer leader election to a follow-up PR
* fix(reset_budget_job): satisfy strict lint gates
String defaults for the two getenv calls (PLW1508) and the chunk
outcome returns moved to try/else (TRY300).
* fix(proxy): isolate guardrail load failures per row
One DB guardrail row that fails to initialize aborted the whole
_init_guardrails_in_db loop, so a single typo'd guardrail type or a
missing required param left the proxy running with zero DB guardrails
registered and requests that should have been blocked reaching the
provider.
Catch per row around sync_guardrail_from_db, log the guardrail name, id
and error, and continue with the remaining rows. The failing row's id is
still added to db_guardrail_ids before the attempt so reconcile_db_guardrails
cannot mistake a live row for a deleted one.
* test(proxy): drop inline note and record reconcile via a handler double
Replaces the patched bound method with an InMemoryGuardrailHandler subclass
that records what reconcile_db_guardrails received, so the test injects a
double instead of swapping a method on a live object.
Merging staging's flat-cost summary work with the capability gating pushed
EntityUsage.tsx to 815 counted lines, over the 800-line eslint cap. Move the
four pure top-N/rollup helpers to entityUsageAggregations.ts and pass their
inputs explicitly.
TopKeyView and TopModelView were mocked to render static text, so nothing
asserted which breakdown fed which table. The mocks now surface their rows and
a new case pins each table to its own data source.
Pass static AWS credentials through optional_params so the real
credential-resolution path runs locally instead of patching
BedrockRerankHandler._get_boto_credentials_from_optional_params.
BedrockRerankHandler._prepare_request duplicated ad-hoc SigV4 signing
instead of using BaseAWSLLM.get_request_headers, the helper every other
Bedrock handler (embeddings, converse, invoke, image) already uses.
The duplicate skipped header filtering before signing, so any forwarded
header (e.g. x-forwarded-for) got included in the signed set and could
invalidate the signature if rewritten downstream between signing and
delivery, the same class of bug fixed for the invoke path in #19111.
LITELLM_ENABLE_PTU_COST_ATTRIBUTION, read through get_secret_bool and defaulting to
false, makes the whole PTU flat-cost feature inert unless an operator opts in. The
daily rollup cron is not registered at all, so no sentinel row is ever written;
/model/new and /model/{id}/update reject a request that carries any PTU model_info
field with a 400 naming the fields and the env var rather than dropping them; the
daily activity read path reports zero flat cost; and the model add and edit forms
hide the four PTU inputs.
The read gate lives where flat cost enters SpendMetrics rather than in the aggregated
SQL select. /team/daily/activity, the endpoint the Usage page reads, is served by the
paginated find_many path and never runs that query, so forcing the select to a
constant zero would have left the reporting surface that matters still showing flat
cost.
Sentinel row filtering is deliberately not gated. An operator can enable the flag,
accrue rows under the __ptu_flat_cost__ api_key, then disable it, and those rows stay
in LiteLLM_DailyTeamSpend; gating the filter too would surface the sentinel as a bogus
api_key and mint a provider bucket for its empty provider. Response fields keep their
shape and report 0.0, so typed clients are unaffected, and the migration and the
ModelInfo field declarations are untouched.
The write gate reads the incoming request rather than the merged deployment, so a
model configured during an earlier opt-in stays editable, and the edit form drops the
PTU keys from the payload instead of sending nulls that would clear stored config.
The dashboard reads the flag from a read-only enable_ptu_cost_attribution key on
/get/ui_settings, computed from the environment on every read. It is deliberately not
an allowlisted persisted setting, and PATCH /update/ui_settings rejects it with a 400,
so an admin cannot flip an env-gated feature from the UI.
Two review findings on the gate itself. The PTU clear loop now runs only when the
feature is enabled: the write gate rejects a value but lets an explicit null through,
and a client round-tripping a model_info blob sends the PTU keys as nulls, so a
disabled proxy would have quietly erased a billing configuration set up during an
earlier opt-in. Disabling pauses PTU rather than discarding its setup. And the
dashboard flag is re-read every thirty seconds instead of the hour the other UI settings
use, since those are persisted records while this one tracks the proxy process; a
restart that flips the variable would otherwise leave the model form offering inputs
the backend now rejects. The flag is polled rather than only marked stale, since a form
that stays mounted and focused never refetches on its own.
The read gate checks the row before the flag. It runs once per metric accumulation and a
record fans out across roughly a dozen breakdowns, while the flag reads through the secret
manager uncached, so consulting it for every accumulation put thousands of lookups on a
shared endpoint that made none before. Only a row actually carrying flat cost reaches it.
The timeout contract check skipped a job whenever either budget came from
a `with:` value it could not parse, or from a matrix column no `include`
row supplied as a number. Both paths produced no pairs and no errors, so
the guard printed "invariants hold" for a caller whose budgets were never
compared at all. A caller reading `${{ matrix.timeout }}` off a mistyped
column while capping the job at 1 minute passed clean.
Unresolvable budgets now come back as the reason they could not be read
and are reported as violations, which is the whole point of a guard built
to catch checks that silently do not run. `Column` tags a matrix
reference so it stays distinguishable from that reason string, and the
report names only the columns that resolve nowhere, since a column every
row supplies is not what left the pair unchecked.
calculate_image_response_cost_from_usage read input_tokens_details with
getattr(), but OpenAI images.edit responses carry it as a plain dict, so
both fields came back None and all input tokens were priced at
input_cost_per_token instead of input_cost_per_image_token (e.g. $5/M
instead of $8/M for gpt-image-2). Read it with the dict-tolerant
_get_token_detail_value helper, as the output side of the same function
already does.
Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Add PTU count, cost per PTU per hour, and effective-from/to date-time pickers to
the model add and edit forms; the create submit and the edit save map the picker
values to model_info as ISO strings
On the Team Usage Cost tab the money tile becomes Total Cost once a team has
accrued flat cost, and expands to a Request Cost and Flat Cost breakdown, so the
summary row stays at five tiles and the cards keep their width. Each of the three
carries a tooltip, including that flat cost is reported rather than charged
against budgets. The Daily Spend chart stacks Flat cost on Request cost, with a
tooltip that splits the two and shows the total
A team that has accrued no flat cost renders exactly as it did before, and other
entity views are unchanged. CSV export gains Flat Cost and Total Cost columns
when a team accrued non-zero flat cost; the existing Spend header is left alone
so downstream parsers keep working
Both forms validate the PTU pair through one shared module. The count rule rejects
a fractional, zero or negative value, and a rate rule rejects a negative one, each
mirroring a contract the backend enforces. Keeping the rate rule shared rather than
on a single form is deliberate: the edit form previously validated only the count,
so a negative rate typed past the input's min reached the backend and failed the
save with a 400 the operator had no way to anticipate
Both forms require PTU Effective From once PTU Count is set, matching the backend, which
rejects PTU config without a start because flat cost accrues from that instant and an
inferred one would bill days a deployment did not exist. The rule lives beside the count
and rate rules in the shared module, so the add and edit paths cannot drift.
The timeout contract check resolved the test and job budgets
independently and compared every value against every other, so two
matrix-sourced columns were paired across different include rows. A
row-wise-valid matrix could be rejected on a pairing no shard actually
runs with. Budgets now resolve per include row, so each shard's test
budget is checked only against that same shard's job budget.
* feat(router): add per-deployment allowed_fails_policy and cooldown_time override support
Three bugs fixed in the router cooldown system: (1) deployment-level allowed_fails and
allowed_fails_policy in model_info now take precedence over router-level settings in
_should_cooldown_deployment; (2) failed fallback deployments now get evaluated for
cooldown via _trigger_cooldown_for_failed_deployment, bypassing the Logging dedup gate;
(3) DualCache promotes Redis cooldown entries using default 600s TTL instead of true
remaining cooldown time -- _corrected_active_cooldown now evicts expired entries and
corrects stale in-memory TTLs on backfill. Adds ServiceUnavailableError, BadGatewayError,
and NotFoundError fields to AllowedFailsPolicy and cooldown_time to LiteLLMParamsTypedDict.
* fix(router): gate fallback cooldown trigger on has_logged_async_failure; use only litellm_metadata for deployment ID
* fix(router): use X | Y union syntax to fix UP007 strict lint gate
* test(router_utils): add coverage for _trigger_cooldown_for_failed_deployment and has_logged_async_failure gate
* test(router_utils): cover deployment cooldown override and exception swallow paths
* fix(router): add InternalServerError/ServiceUnavailableError/BadGatewayError/NotFoundError to router-level get_allowed_fails_from_policy
* fix(router): format router.py and add router-level policy tests
* test(router): add CI-visible coverage for per-deployment cooldown policy
Tests for `_get_deployment_cooldown_policy`, `_resolve_allowed_fails_from_policy`,
and `_should_cooldown_based_on_deployment_policy` (cooldown_handlers.py), the
`_corrected_active_cooldown` branches in CooldownCache, and the four new
exception-type branches in `Router.get_allowed_fails_from_policy` (router.py) --
all in `tests/test_litellm/` which the enterprise-routing CI job runs.
* fix(router): use is not None guard for cooldown_time_override in should_cooldown_based_on_allowed_fails_policy
A cooldown_time_override of 0 was previously treated as falsy and silently
fell through to the router-level cooldown_time value. Switched to an explicit
is not None check so that zero is honored as a valid override.
Added a regression test covering the zero case.
* fix(router): honor has_logged_async_failure and metadata for fallback cooldown; support both model_info and litellm_params locations
Manual verification against a live proxy surfaced that the fallback-cooldown-gap
trigger never actually fired: the has_logged_async_failure check read a plain
attribute that Logging never sets (the real flag lives in model_call_details),
and the deployment_id lookup only trusted litellm_metadata, which regular chat
completions never populate (only batch/thread/file endpoints do). Router
overwrites model_info on whichever key is present before every attempt, so
metadata is equally authoritative there, not caller-controlled as previously
assumed. Also let allowed_fails/allowed_fails_policy/cooldown_time be set under
either model_info or litellm_params, each preferring its own canonical location.
* fix(router): fix ContentPolicyViolationError policy shadowing and partial-policy zero-threshold
Two bugs from Greptile review on PR #34416:
- ContentPolicyViolationError subclasses BadRequestError, so listing
BadRequestError first in _EXCEPTION_POLICY_FIELDS made the isinstance
check always match BadRequestError for content-policy errors, using the
wrong allowed_fails threshold. Reordered so the subclass is checked first.
- A deployment with a partial allowed_fails_policy and no deployment-wide
allowed_fails forced allowed_fails_override=0 for any exception type its
policy didn't cover, cooling the deployment down on the first unrelated
failure. Now defers to router-level behavior for uncovered exception
types instead of forcing an immediate cooldown.
* fix(router): only trust a metadata/litellm_metadata bucket the router itself wrote deployment info into
veria-ai flagged that preferring litellm_metadata whenever present could pick up a
caller-supplied litellm_metadata.model_info.id (preserved via allow_client_pricing_override)
instead of the metadata bucket the router actually populated for a regular completion's
fallback attempt, naming an arbitrary "victim" deployment for cooldown.
Router._update_kwargs_with_deployment() always writes model_info and
deployment_model_name into the same bucket together. Only trust a bucket that
carries deployment_model_name alongside model_info, since that marker is only
ever set by the router itself, not by request-body metadata.
* test(router): add regression coverage for ContentPolicyViolationError policy shadowing
The subclass-ordering fix in commit 38fe4e4490 had no regression test.
Verified the new test fails on the pre-fix ordering (asserts 2, got 10)
before restoring the fix, and confirmed the same behavior through the full
_should_cooldown_deployment call path against a real Router instance.
* fix(router): let explicit allowed_fails_policy entries override the generic 4XX cooldown exclusion
_is_cooldown_required skips cooldown evaluation for any 4XX status outside
{429, 401, 408, 404} by default, since a generic client error is usually not
the deployment's fault. BadRequestError and ContentPolicyViolationError both
carry status 400, so their AllowedFailsPolicy fields (BadRequestErrorAllowedFails,
ContentPolicyViolationErrorAllowedFails, both router-level pre-existing and the
new deployment-level ones) were silently unreachable: an operator could set
them to any value with no effect, since _is_cooldown_required blocked cooldown
evaluation before that policy was ever consulted.
_should_run_cooldown_logic now also checks whether an explicit allowed_fails_policy
entry (deployment-level or router-level) covers the exception's type, and if so,
proceeds with cooldown evaluation regardless of the generic status-code exclusion.
The exclusion remains the default for exception types with no explicit policy.
Verified live against a mock-triggered ContentPolicyViolationError (config-level
mock_response, azure/gpt-4.1-mini deployment) with BadRequestErrorAllowedFails=100
and ContentPolicyViolationErrorAllowedFails=0 on the same deployment: it now cools
down after exactly one ContentPolicyViolationError instead of never cooling down.
* fix(router): use the router-stamped failed_deployment_id for fallback cooldown targeting
Greptile flagged a real gap in the metadata-bucket-based deployment lookup:
for a generic-API-call fallback, the router writes the current attempt into
litellm_metadata, but a stale "metadata" bucket carrying the same
deployment_model_name marker (from an earlier point) would be picked first,
cooling the wrong deployment.
Router already has a more robust, pre-existing mechanism for this exact
problem: _set_failed_deployment_id_on_exception stamps the failing
deployment's id directly onto the exception at the point of failure,
immune to metadata-bucket ambiguity since a caller can't influence it and
it doesn't depend on which bucket the current call type happens to use.
It just wasn't called from _ageneric_api_call_with_fallbacks_helper's
except block, unlike _completion/_acompletion.
Added the missing call there (matching the existing pattern exactly), and
changed _trigger_cooldown_for_failed_deployment to prefer
exception.failed_deployment_id when present, falling back to metadata-bucket
inspection only for call paths that don't stamp it yet.
Verified live: the standard fallback-cooldown-gap scenario (two bad-key
deployments in a fallback chain) still correctly cools down both the
originally-called and fallback deployment.
* fix(router): address human review on per-deployment cooldown overrides
Scope allowed_fails_policy override to deployment-level only (a router-level
policy predates this feature and must keep its existing behavior), exempt
advisor-orchestration failures from the fallback cooldown trigger, keep the
single-deployment model group protection intact against a generic
deployment-level allowed_fails, make cooldown_time precedence consistent
across resolution paths, fix a falsy-zero swallowing bug in the router-level
allowed_fails fallback, and make allowed_fails_policy resolution fall through
to the next matching exception type instead of stopping at the first unset
field.
Also restrict allowed_fails/allowed_fails_policy/cooldown_time to model_info:
litellm_params gets copied into the actual provider request, so a router-only
setting placed there would leak into that request.
* test(router): update test_cooldown_handlers.py for the deployment-policy signature change
Surfaced by the rebase: this mirrored test file (tests/test_litellm/ mirrors
litellm/) predates the router_unit_tests/ coverage added earlier in this PR and
was still calling _should_cooldown_based_on_deployment_policy with its old
4-argument signature and asserting the now-removed litellm_params cooldown_time
location.
* test(router): update test_fallback_event_handlers.py for model_info-only cooldown_time
Another mirrored test file surfaced by the rebase that still asserted the
now-removed litellm_params.cooldown_time location.
* fix(router): match cooldown-duration precedence in the fallback path to the primary path
_trigger_cooldown_for_failed_deployment only checked deployment config before
falling back to the router default, skipping the response Retry-After header
step that Router.deployment_callback_on_failure applies on the primary path.
* fix(router): restore litellm_params.cooldown_time as a pre-existing fallback
cooldown_time already had litellm_params support on Router.deployment_callback_on_failure
before this PR; the earlier model_info-only restriction (aimed at the leak concern
for the genuinely new allowed_fails/allowed_fails_policy fields) incorrectly dropped
that pre-existing capability too. model_info still takes priority when both are set.
* fix(router): keep the fallback-cooldown trigger in sync with #35104's review fixes
Applies the same two fixes landed on the split-out PR #35104 (which #34416
still duplicates until it's rebased onto the merged base): increment the
deployment's per-minute failure counter before evaluating cooldown, and
require the server-stamped failed_deployment_id instead of trusting a
metadata bucket, since neither "metadata" nor "litellm_metadata" can be told
apart from a caller-supplied one without knowing the call's function_name.
* fix(router): freeze the model_info fallback mapping to satisfy the type-discipline gate
* fix(router): defer f-string interpolation in fallback-cooldown debug logs
* fix(router): annotate cooldown-path locals with Final to satisfy the LIT010 budget
* fix(router): suppress reportPrivateUsage for cross-module cooldown helpers
* fix(router): don't cool down deployments for request-scoped 404s on generic API fallbacks
* fix(router): stamp the dynamic client-side-credential deployment id, not the shared static one
* fix(router): keep up with upstream typing modernization and Final-annotation ratchet
* fix(router): don't cool down deployments for a caller-supplied x-litellm-timeout
* fix(router): stamp dynamic client-side-credential id in completion fallback paths too
The generic-API-call helper already stamped the effective (dynamic-if-client-side-credential)
deployment id on exceptions, but the regular _completion/_acompletion exception handlers still
stamped the static shared deployment's id. A tenant using invalid forwarded credentials could
generate repeated failures attributed to, and eventually cooling down, the shared deployment
other tenants rely on. Extracted the stamping logic into one shared helper used by all three
call sites (generic API, sync completion, async completion) so the fix and future changes to it
stay in one place.
* fix(proxy): recognize body-supplied timeout/request_timeout/stream_timeout as caller-controlled
client_side_timeout was only set when the caller used the x-litellm-timeout header, but
Router._get_timeout also resolves the effective timeout from kwargs["timeout"],
kwargs["request_timeout"], and kwargs["stream_timeout"], all settable directly in the
request body (and x-litellm-stream-timeout wasn't marked either). A caller could set any
of those to a near-zero value, force a 408 on every deployment in a fallback chain, and
cool down deployments other tenants rely on without the guard in
_trigger_cooldown_for_failed_deployment recognizing it as caller-controlled. Also strip
any client-forged client_side_timeout from the request body so the marker is always
server-computed.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
Aggregate the ptu_flat_cost written by the rollup into SpendMetrics.flat_cost and
DailySpendMetadata.total_flat_cost, so /team/daily/activity returns flat cost
alongside per-request spend. The aggregated SQL path selects ptu_flat_cost only
for LiteLLM_DailyTeamSpend and a constant zero for the other daily tables, keeping
the response shape uniform.
Rows written under the PTU sentinel api_key add their flat cost to every parent
bucket (per-model, per-day, per-team totals) but never appear as an api_key row in
any breakdown, and are excluded from the per-request provider breakdown; the
sentinel string is not a real key alias. Both flat_cost and total_flat_cost default
to zero, so a read of any entity without PTU config is unchanged.
The sentinel row now keys on the deployment id, so the per-model breakdown keys it on
model_group instead. That breakdown key is rendered directly as a label by the Usage page
and the daily_with_models export, and a deployment id there would read as a UUID. Two
deployments sharing a public name merge under it, which is the collapse the write path
used to do by summing them into one row. Request rows are untouched and still key on
model, since their model_group is a routing concept rather than a display name.