The daily spend flush emitted one INSERT ... ON CONFLICT per aggregated key,
so every replica put hundreds of statements on the database each interval,
all contending for the same handful of hot rows and each holding its row
locks for the rest of the enclosing batch transaction. LiteLLM_DailyTagSpend
felt it worst because a request writes one row per tag, and litellm adds two
user-agent tags of its own by default.
A batch now goes out as a single multi-row statement. Rows are folded by the
conflict tuple first, and every nullable member of that tuple is normalized
to '': a NULL can never match itself in a unique index, so such a row was
re-inserted on every flush rather than aggregating, and a NULL model made
prisma reject the whole batch.
* fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill
A deployment that failed partway through a fallback chain (any attempt after
the first) was silently exempt from cooldown, because the has_logged_async_failure
dedup flag blocks the normal failure callback for every attempt past the first.
_trigger_cooldown_for_failed_deployment now explicitly evaluates cooldown for
that deployment when the dedup flag is set, using the same deployment-config >
response-header > router-default precedence as the primary failure path, and
skips advisor-orchestration failures. Deployment-ID resolution prefers the
exception's stamped failed_deployment_id, now also set from the generic-API-call
fallback path (rerank, embeddings, /v1/messages, etc.), falling back to metadata
inspection for call paths that don't stamp it yet.
CooldownCache also recomputes the remaining TTL when DualCache promotes a Redis
entry into the in-memory layer: before this, a cooldown entry restored from Redis
kept the in-memory layer's default 600s TTL regardless of the deployment's real
cooldown_time, so a deployment could stay excluded from routing for up to 10
minutes after a much shorter cooldown had already expired.
* fix(router): address Greptile review on the fallback-cooldown trigger
Two P1 findings on PR #35104:
- _trigger_cooldown_for_failed_deployment never incremented the deployment's
per-minute failure counter before evaluating cooldown, so a fallback
deployment's repeated retryable failures never accumulated toward the
default percent-fail-rate threshold that _should_cooldown_deployment checks.
- The metadata-bucket fallback (checking "metadata" before "litellm_metadata"
for a deployment_model_name marker) could be fooled by a caller with
permission to set metadata, since neither bucket's authorship can be
determined without knowing the call's function_name. Removed it entirely;
cooldown now requires the server-stamped failed_deployment_id, matching
what the primary chat-completions path and the generic-API-call path
(rerank, embeddings, /v1/messages, etc.) already set unconditionally.
* fix(router): freeze the litellm_params fallback mapping to satisfy the type-discipline gate
* fix(router): defer f-string interpolation in fallback-cooldown debug logs
* fix(router): annotate cooldown-path locals with Final to satisfy the LIT010 budget
* fix(router): don't cool down deployments for request-scoped 404s on generic API fallbacks
* fix(router): stamp the dynamic client-side-credential deployment id, not the shared static one
* fix(router): don't cool down deployments for a caller-supplied x-litellm-timeout
* fix(router): stamp dynamic client-side-credential id in completion fallback paths too
The generic-API-call helper already stamped the effective (dynamic-if-client-side-credential)
deployment id on exceptions, but the regular _completion/_acompletion exception handlers still
stamped the static shared deployment's id. A tenant using invalid forwarded credentials could
generate repeated failures attributed to, and eventually cooling down, the shared deployment
other tenants rely on. Extracted the stamping logic into one shared helper used by all three
call sites (generic API, sync completion, async completion) so the fix and future changes to it
stay in one place.
* test(router): add direct-reference unit tests for the new stamping helper
router_code_coverage.py's coverage gate flags _stamp_failed_deployment_id_with_effective_model_info
as untested because it only sees the function invoked indirectly through _completion/_acompletion's
exception handlers. Added two tests that call it directly, covering both the dynamic-id-present and
static-fallback branches.
* test(router): cover the timeout stamping branch and async active-cooldown append
_acompletion's litellm.Timeout handler and async_get_active_cooldowns' happy
path both lacked direct coverage despite their sibling branches (the generic
Exception handler, the sync get_active_cooldowns) being tested.
* test(router): remove duplicate cooldown-trigger and fallback-helper tests
#34416 landed its own TestTriggerCooldownForFailedDeployment/
TestRunAsyncFallbackTriggersCooldown classes and
test_ageneric_api_call_with_fallbacks_helper_stamps_failed_deployment_id
covering the exact same scenarios as this branch's earlier flat-function
tests, once its version of fallback_event_handlers.py was taken as-is
during the last merge. Dropping the redundant copies.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* feat(proxy): add per-deployment keepalive_seconds SSE heartbeat for long-running streams
Adds _iter_with_keepalive, _keepalive_from_deployment_config, and
_resolve_keepalive_seconds helpers to proxy_server.py. When enabled
(keepalive_seconds > 0 in request body or deployment litellm_params),
async_data_generator emits ': ping\n\n' SSE comment frames every N
seconds during idle upstream intervals, preventing load-balancer
idle-timeout drops on long chain-of-thought reasoning streams.
The hot path (keepalive_seconds absent or 0) is a plain async-for with
no per-chunk Task wrapping — zero overhead. Includes 8 new unit tests
covering sentinel emission, hot-path pass-through, early-close cleanup,
priority resolution, deployment-config lookup, and end-to-end heartbeat
emission through async_data_generator.
Registers keepalive_seconds in all_litellm_params (types/utils.py) so
the parameter is not stripped from request bodies. Adds the field to
LiteLLMParamsTypedDict and GenericLiteLLMParams (types/router.py) so
deployment YAML config is parsed and validated.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(proxy): narrow BaseException to CancelledError to fix BLE001 strict lint gate
* fix: use explicit None check instead of truthiness in keepalive_seconds extraction
`float(raw or 0)` would treat any falsy value (including the integer 0)
as absent and substitute 0.0 before float() saw it. Replace with
`float(raw) if raw is not None else 0.0` so a caller-supplied zero is
correctly passed through to the `value <= 0` guard that disables
keepalive, rather than being silently overwritten.
* fix(proxy): don't guess a deployment's keepalive_seconds when model_id is missing
When a streaming response lacks _hidden_params.model_id, the fallback that
looks up keepalive_seconds by model_name previously returned the first
configured deployment's value, which could apply the wrong interval (or
override an explicit disable) when multiple deployments share the same
model_name with different keepalive_seconds settings. Only resolve the
fallback when every deployment agrees; otherwise leave it unset.
* fix(proxy): also treat an unset keepalive_seconds as disagreement in the fallback
The model_name fallback for keepalive_seconds only compared configured
values, filtering out deployments that leave the field unset entirely.
That meant a deployment with no keepalive_seconds configured could still
inherit a sibling deployment's interval when model_id is unavailable.
Compare the raw per-deployment value (including None for unset) so an
unconfigured deployment never silently adopts another's heartbeat.
* fix(proxy): deployment-level keepalive_seconds: 0 is a hard disable clients can't override
Previously an authenticated client's request-level keepalive_seconds always
took precedence over the deployment default, including when a deployment
operator explicitly set keepalive_seconds: 0 to disable heartbeats. That let
any client re-enable heartbeats for a deployment the operator opted out of,
using them to keep an idle-looking stream alive past a load balancer's idle
timeout and hold a parallel-request slot open longer than intended.
Treat an explicit deployment-level 0 as authoritative: resolve the
deployment's configured value first, and short-circuit to disabled before
ever looking at the request body if the deployment hard-disabled it.
* fix(proxy): a stale (unresolvable) model_id must not fall through to model_name guessing
A populated _hidden_params.model_id names the specific deployment that
served a stream. If that ID no longer resolves (e.g. a deployment removed
by a config reload mid-stream), the resolver was falling through to the
model_name-based fallback, letting a currently-live sibling deployment's
keepalive_seconds silently apply to a stream it never served. Return None
once a populated model_id fails to resolve, rather than degrading to a
guess.
* fix(proxy): keepalive_seconds is operator-only by default; require deployment opt-in for client override
A security review flagged that a client's request-level keepalive_seconds
could unilaterally enable heartbeats for any deployment, even one that
never configured keepalive_seconds at all, letting an authenticated client
defeat load-balancer idle timeouts and hold a parallel-request slot open
for longer than the deployment operator ever intended, with no way for
the operator to prevent it short of explicitly setting keepalive_seconds: 0.
Add allow_client_keepalive_override (default False) to LiteLLMParamsTypedDict
and GenericLiteLLMParams. _resolve_keepalive_seconds now ignores the request
body's keepalive_seconds entirely unless the resolved deployment explicitly
grants override permission; only the deployment's own configured value (or
disabled, if unset) applies otherwise. An explicit deployment-level 0 still
takes priority over everything, including a grant of override permission.
* fix(proxy): register allow_client_keepalive_override in all_litellm_params
Caught during live proxy verification against the real Anthropic API:
allow_client_keepalive_override was added to LiteLLMParamsTypedDict and
GenericLiteLLMParams but never registered in all_litellm_params, so it
leaked straight through into the provider request body as an unrecognized
field. Anthropic rejected every call on a deployment that had this field
configured with a 400 ("Extra inputs are not permitted"), regardless of
its value. Register it alongside keepalive_seconds so it's stripped
before reaching the provider, matching what keepalive_seconds already
does.
* feat(proxy): support keepalive_seconds via x-litellm-keepalive-seconds header
Some clients (e.g. the Vercel AI SDK) can set custom headers more easily
than extra JSON body fields. Add x-litellm-keepalive-seconds, following
the existing x-litellm-timeout/x-litellm-stream-timeout/x-litellm-num-retries
convention in LiteLLMProxyRequestSetup: the header merges into the same
data["keepalive_seconds"] field the request body already populates, so it
goes through the exact same _resolve_keepalive_seconds precedence and the
allow_client_keepalive_override gate -- a header can't enable heartbeats
for a deployment that hasn't opted in any more than the body field can.
Verified live against the real Anthropic API: the header produces real
heartbeats on an opt-in deployment (88 pings over a genuine long-reasoning
stall) and is silently ignored on a deployment without override permission
(0 pings), matching the existing body-field behavior exactly.
* chore: rebase onto litellm_internal_staging, drop unrelated credential_migration.py reformat, fix budget-ratchet drift
Rebased onto the current litellm_internal_staging (merge-base was 5 days
stale). Dropped the now-redundant schema.d.ts-only regen commit entirely
(the new base's own schema.d.ts already supersedes it) and regenerated
schema.d.ts fresh against the new base.
Reverted litellm/proxy/management_endpoints/credential_migration.py to
exactly match litellm_internal_staging: it was a pure reformat with no
semantic change, unrelated to this PR, flagged by review as unnecessary
noise in an encryption-migration file.
Fixed two lint-budget-ratchet failures caused by the base's ceilings
tightening since this branch last synced (other merged work lowered
ANN401/LIT001 budgets; this code was previously under budget and didn't
change):
- _iter_with_keepalive's aiter param: Any -> AsyncIterator[Any], a real
narrowing (it's always the result of .__aiter__()).
- _keepalive_from_deployment_config/_resolve_keepalive_seconds's
request_data param: dict[str, Any] -> Mapping[str, Any], matching the
existing read-only-dict convention already used elsewhere in this file
(_apply_ssrf_general_settings, _build_redis_usage_cache, etc.) for
params that are only ever read, never mutated.
- response/raw params: dropped the explicit `Any` annotation to match
async_data_generator's own (deliberately unannotated) `response` param,
its actual caller.
- litellm_pre_call_utils.py's new headers param: dict -> Mapping[str, str],
same read-only-dict rationale.
* fix(proxy): freeze the transient collections in the keepalive helpers
_iter_with_keepalive and _keepalive_from_deployment_config built a set
literal for asyncio.wait, a set comprehension for the per-deployment
config-agreement check, and two dict-literal fallbacks, all flagged by
the LIT002 mutable-collection-construction gate. Switched to a tuple
for asyncio.wait, a frozenset-wrapped generator plus next(iter(...))
for the config check, and a shared MappingProxyType({}) empty mapping
for the fallbacks.
* fix(proxy): trust metadata.model_info.id over the stale model group after a router fallback
Greptile P1: when a streaming request falls back from model group A to
group B and the response's _hidden_params carries no model_id,
_keepalive_from_deployment_config fell straight through to guessing
via request_data["model"], which still names the pre-fallback group A
since the fallback handler mutates its own local **kwargs copy, not
this dict. request_data[metadata|litellm_metadata]["model_info"]["id"],
by contrast, is mutated on this same dict by
Router._update_kwargs_with_deployment on every attempt including
fallbacks (the same source ProxyLogging._build_litellm_call_info uses
for logging), so check it before falling through to the model-name
guess.
Added two regression tests that fail on the prior code (assert
get_model_list is never called once metadata.model_info.id resolves)
and pass with the fix.
* Revert "fix(proxy): trust metadata.model_info.id over the stale model group after a router fallback"
This reverts commit d779067864.
* fix(proxy): satisfy the new LIT010/ANN001 gates in the keepalive helpers
litellm_internal_staging picked up a LIT010 (every local/module variable
must be declared Final unless it's genuinely rebound) and tightened
ANN001 (missing parameter annotations) since this branch last synced.
Annotated every single-assignment local and module constant with
Final, suppressed pending's loop-carried reassignment with
# rebind-ok, and typed the previously-bare response/raw parameters as
object with isinstance narrowing at their use sites instead of cast
(LIT006 discourages cast; validate into a concrete type instead).
Also swapped the hand-rolled getattr(response, "_hidden_params", None)
+ isinstance(hidden, dict) check for the existing
get_hidden_params_dict() helper already used for this exact purpose
elsewhere in this file and in common_request_processing.py.
* fix(proxy): re-resolve keepalive_seconds per chunk to track mid-stream fallback
Greptile P1: the router can perform a mid-stream fallback to a
different deployment partway through a stream (MidStreamFallbackError
in router.py), and Router._apply_fallback_hidden_params_to_item merges
the fallback deployment's hidden params onto every subsequent chunk.
But _resolve_keepalive_seconds was only ever called once, before
iteration started, against the pre-fallback response wrapper, so a
stream that fell back to a deployment with a different (or disabled)
keepalive policy kept using the original deployment's interval for the
rest of the stream.
_iter_with_keepalive now takes a resolve_keepalive_seconds(item)
callback and re-resolves after every real chunk using that chunk's own
_hidden_params (which do carry the fallback deployment's identity),
rather than trusting the value picked before iteration began. Updated
the three existing timing tests to inject a constant-returning
resolver, since they pin the sentinel/cancellation mechanics rather
than re-resolution, and added two regression tests (interval lowered
and raised mid-stream) that fail against the prior static-resolve
signature and pass with the fix.
* fix(proxy): keep re-resolving keepalive even when a stream starts disabled
Greptile P1: a stream that starts on a deployment with keepalive off
(or unset) skipped _iter_with_keepalive entirely at the call site, so
a mid-stream fallback to a deployment that enables it never got a
chance to activate heartbeats for the rest of that stream, risking the
exact load-balancer idle-timeout this feature exists to prevent.
_iter_with_keepalive now has an internal fast path for
keepalive_seconds <= 0 that still re-resolves after every chunk (no
asyncio.create_task/wait overhead while inactive, same cost as a bare
async for), so activation from a disabled start works the same way
deactivation and interval changes already do. The caller now only
skips wrapping entirely when there's no router to ever fall back
through in the first place (llm_router is None), rather than whenever
the first chunk's deployment happens to start with keepalive off.
Added a regression test that starts keepalive_seconds=0, has the
resolver enable a short interval on a later chunk, and asserts
sentinels appear afterward; it fails against the prior
call-site-gated code and passes with the fix.
* perf(proxy): memoize keepalive resolution per chunk's model_id
_resolve_keepalive_seconds ran a full llm_router.get_deployment() Pydantic
rebuild after every streamed chunk, even when keepalive was unconfigured
anywhere in the deployment list, since async_data_generator wraps every
stream once a router exists. Caching the result by model_id keeps mid-stream
fallback re-resolution correct while paying the router lookup once per
deployment instead of once per token.
* fix(proxy): expire cached keepalive resolution after a bounded TTL
veria-ai flagged that caching by model_id alone lets an already-in-flight
stream keep evading a live config reload (deployment removed, keepalive
disabled, or client override revoked) for the rest of the stream. Expiring
the memo after _KEEPALIVE_CACHE_TTL_SECONDS bounds that window instead of
freezing the resolved value for the stream's full lifetime, while still
avoiding a full deployment rebuild on every chunk in the steady state.
Also fixes add_litellm_data_for_backend_llm_call's now-required request_data
kwarg in the header-merge test, picked up by rebasing onto
litellm_internal_staging.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(arize): stop MCP CallToolResult from aborting span attribute setting
`call_mcp_tool` logs the MCP SDK's `CallToolResult`, a Pydantic model with
no `.get`. `_coerce_response_obj_for_attrs` left it untouched and
`_set_request_attributes` then raised AttributeError, which aborted the rest
of the attribute block, so MCP tool spans lost their invocation params,
input messages, and outputs.
Dump Pydantic models that lack `.get` to a dict, and guard the response
id/model reads the same way `_set_response_attributes` already does so any
other uncoercible response object degrades instead of crashing.
* feat(arize): render MCP tool calls as OpenInference TOOL spans
`call_mcp_tool` spans carry neither `messages` nor `choices`, so every
generic extraction path left Input and Output blank and the span showed only
provider/model metadata.
Emit `tool.name` from `metadata.mcp_tool_call_metadata`, `input.value` from
the tool arguments, and `output.value` from the `CallToolResult` content
(text parts when present, JSON otherwise). Arguments and results are user
content, so the input/output emit is gated on the same
`should_redact_message_logging` check the passthrough normalizer uses.
Reuse `_to_plain_dict` for the Pydantic coercion instead of the local
BaseModel branch added in the previous commit.
* fix(arize): annotate the new MCP helper parameters
The strict-rule gate flagged three new ANN001 violations. Type the payload
as StandardLoggingPayload | None and the coerced response as object, which
the isinstance guards already narrow.
* fix(arize): annotate the MCP helper against the type-discipline gate
LIT001 bans mutable collections in annotations, so the kwargs parameter
becomes Mapping[str, object]. should_redact_message_logging still declares a
dict it only ever reads, and widening it would cascade into core_helpers, so
the call carries a scoped ignore instead. Narrow the payload by None rather
than isinstance now that it is typed, and annotate the values read out of the
untyped logging payload.
* fix(arize): record empty MCP arguments and results instead of dropping them
Zero-argument tools record arguments={} and successful calls can return
content=[]; both were skipped by truthiness, leaving the generic placeholder
on Input and nothing on Output. Read structuredContent when content yields
no text, and cover the list_mcp_tools response shape.
* fix(arize): keep media parts in mixed MCP results
A result mixing text and media returned the text alone, so Arize showed
text/plain and dropped the image or resource parts.
---------
Co-authored-by: Sean Lee <yihsean@gmail.com>
The prompt editor is a base-ui Dialog at z-index 50. The create form houses it in
the same base-ui Dialog, so it stacks on top, but the edit form was an antd Modal
whose portal computes to z-index 1000, so the editor opened underneath it and was
neither readable nor clickable.
Move the edit form onto the Dialog the create form already uses, which puts the
whole nesting chain in one overlay layer. A dialog opened from inside another
dialog now reads as a drill-down rather than a stack: base-ui stamps
data-nested-dialog-open on the parent while a child is open, so the parent steps
aside instead of showing its own edges around a differently sized child.
Hiding the tab behind all_admin_roles took it away from org admins, who are
entitled to it: /v2/team/list?status=deleted returns 200 for them, scoped to
their own organizations. An org admin is an organization membership rather
than a global role, so their session carries user_role "internal_user" and no
role-based gate can ever see them.
Lift the membership lookup the left nav already did into a shared
useIsOrgAdmin hook, and let a capability opt into allowing org admins.
viewDeletedTeams is the only one that opts in; the backend still refuses org
admins on /v1/tool/list, /policies/list, /prompts/list and /audit, so those
gates stay as they are. The hook also accepts a session role of org_admin, in
case a deployment maps one through SSO.
* fix(bedrock): reject Anthropic server-side web_search tool with actionable error
Bedrock's Anthropic Messages endpoints cannot execute Anthropic's
server-side web_search tool, so forwarding it returns an opaque
"The provided request is not valid" 400 from Bedrock. Fail fast in the
invoke transform with an error that names the unsupported tool, the
model, and links the web search interception docs as the fix.
* refactor(bedrock): address review nits on web_search guard typing
knip derives its entry points from vitest's `test.include`, which does not
cover `test.typecheck.include`, so the new `*.test-d.tsx` file read as an
unused file and failed the lint job. Declare the glob as an entry point.
Also drops the doc comment on `DataTableResolvedProps`; the rationale for
the resolved/public split belongs in the commit that introduced it.
* fix(proxy): treat SAML as configured in UI SSO detection
_has_user_setup_sso only checked OAuth client IDs, so SAML-only setups
left /.well-known/litellm-ui-config sso_configured=false and the login
button gray even when SAML IdP metadata was set. Include
SAML_IDP_METADATA_URL / SAML_IDP_METADATA_XML so UI discovery matches
the login redirect path.
* chore: adhere to comment policy
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix: apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
The removed comments restated the test names and the assertions directly
below them. The reasoning they carried is already recorded in the commit
that introduced the fix and in the pull request body.
Workflow Runs, Memory and Guardrails Monitor were visible to every role
while their page-load routes are proxy-admin-only, so a non-admin got a
page shell and a 401. Cost Optimization was half-broken the same way: its
Overall charts run on /user/daily/activity, which every role may call, but
tool spend, prompt caching, prompt compression and auto-router benchmarks
are all proxy-admin-only.
Add viewWorkflowRuns, viewMemory, viewGuardrailUsage and
viewProxyWideCostData, each gating the nav entry, the page and the request
together. The first three hide their page, including the direct-URL path,
since nothing on them works for a non-admin. Cost Optimization keeps its
page and drops only the parts a non-admin cannot read.
Gating both Agentic children left roles with no visible child rendering the
parent as a leaf link to ?page=agentic, which is not a route, so a parent
whose children are all filtered out is now dropped.
Role lists follow what the proxy actually grants: proxy_admin and
proxy_admin_viewer are served, and org admins are not, because
_user_is_org_admin needs an organization_id that a page-load GET never
carries.
* feat(ui): show vector store indexes on the Vector Stores page
Adds a proxy-admin-only Indexes tab listing rows from GET /v1/indexes:
index name, backing vector store, provider index, creator, and created
date. The tab is hidden for non proxy-admin roles to match the
endpoint's gate, and data loads lazily on first visit.
* feat(ui): link index rows to their vector store and creator
Vector Store cells open the store's info view when the name resolves to
a registered store, and Created By cells deep link to the users page via
a new userDetailHref, with the users page reading the user query param
through nuqs so the link is shareable.
* feat(ui): link docs and note supported providers on Indexes tab
* fix(ui): show not-found state instead of infinite loading for missing vector store
DataTable accepted any mix of its 40-odd props and rejected the incoherent
combinations at runtime, from a validator that threw during the first render.
A caller only found out it had wired server sorting without a `sorting` prop
when the page blew up in front of them.
Split the public prop type into mode-keyed unions instead, so the compiler
rejects those combinations at the call site. `validateDataTableConfig` and
`DataTableConfigError` go away; the component body reads an unchanged flat
`DataTableResolvedProps`, which every union member is assignable to, so there
is no narrowing inside it.
All 44 existing call sites typecheck against the new union unchanged, which
`next build` covers. That build only typechecks the app module graph, so the
prop type itself needed a gate of its own: `npm run test:types` runs vitest's
typecheck mode over `*.test-d.tsx`, and the unit workflow now runs it. The
four guards deleted from `DataTable.test.tsx` come back there as compile-time
assertions, and loosening the union back to the flat shape fails all five.
The Old Usage nav entry carried no role restriction, so every role saw it
and the page immediately fired eight /global/spend/* requests that the
proxy withholds from non-admins, producing a wall of 401s.
Gate the nav entry, the page, and both of its mount effects behind a
single viewGlobalSpend capability scoped to proxy_admin and
proxy_admin_viewer, matching what the backend actually serves.
Also drop the session JWT that adminspendByProvider put in the
/global/spend/provider query string; the handler never read it.
* fix(scripts): stop type-discipline checker reading Literal strings as forward refs
The checker re-parsed every string constant inside an annotation as a
forward reference, so Literal["list"] was counted as the mutable list
type. Skip Literal subtrees and ratchet the LIT001 ceiling down to the
corrected count.
* fix(proxy): keep lazy openapi snapshot fragments for transitively imported features
generate_snapshot skipped register_fn for any feature module already in
sys.modules, so a module pulled in transitively by an earlier feature
never mounted its routes and its fragment silently vanished on regen
(vector_store_management). Route collection also matched path_prefixes
only, dropping suffix-matched routes from fragments. Register every
feature and collect routes with feat.matches, mirroring the runtime
loader.
* feat(proxy): add GET /v1/indexes to list vector store indexes
/v1/indexes was POST-only, so indexes created through it could never be
viewed again. Add an admin-only list endpoint returning the stored index
rows newest first, fix the stale index_create docstring curl, and
regenerate the lazy openapi snapshot and dashboard schema types.
* chore(proxy): defer lazy openapi snapshot catch-up regen to a follow-up
Reverts _lazy_openapi_snapshot.json and schema.d.ts to the staging
versions. The snapshot was months stale, so regenerating it here buried
the actual change under ten thousand generated lines. A follow-up will
land the regen together with CI enforcement that keeps the snapshot
current. Until then GET /v1/indexes is served but absent from the
dashboard's generated types, which the UI step needs anyway.
* fix(proxy): use Annotated dependency to avoid new B008 violation
The Virtual Keys table and the Logs page team filter both asked for every
team on the proxy, which /v2/team/list and /team/list reject with a 401 for
any role below proxy admin or org admin. Both endpoints answer the same
request with the caller's own teams when it carries a user_id, so send one.
Only the two unscoped call sites change. The remaining callers either
already role-branch or render on surfaces gated to roles the endpoints
answer broadly, and scoping those would shrink the list they see: a proxy
admin scoped to their own id gets nothing back, and an org admin scoped on
/team/list loses the org teams they administer but do not belong to.
The shared helper reads the display-form session role rather than
all_admin_roles, which mixes display labels with raw role names and so does
not match the "Org Admin" value the dashboard actually holds.
The runbook still read as a fully manual flow: dispatch the publish workflow by
hand, then approve a second gate in the mirror repo. Neither is true now.
project-releaser checks the provider changelog on every release except adhoc,
nightly included, and dispatches the publish itself when the topmost released
heading has moved ahead of the mirror's tags, so cutting the version heading is
what ships the provider. The mirror's own release workflow no longer gates,
leaving one approval in project-releaser.
* fix(reset_budget_job): advance budget_reset_at atomically with the spend cascade
A postgres timeout mid-cascade previously left LiteLLM_BudgetTable rows
stamped for the next window while team member, enduser, org and tag spend
stayed at cap, so every later tick skipped them until the window rolled
over. All cascade writes and the budget_reset_at advance now share one
prisma batch transaction; a failed run persists nothing and the rows stay
due for the next ~10 minute tick. Cache and counter invalidation runs only
after commit, and the catch-all enduser log line now names the cascade.
* fix(reset_budget_job): elect one runner per tick and chunk the reset scans
Every pod and worker previously ran the reset job every ~10 minutes,
each fetching every expired row with no limit and writing one giant
transaction at the same calendar-aligned boundary; that concurrency is
what piled up postgres lock contention and timeouts. The job now takes
the shared PodLockManager redis lock (no redis keeps the old behavior),
and each phase walks its due rows in 500-row chunks, one transaction per
chunk, stopping when a chunk is short, makes no forward progress, or
hits the per-run cap; leftovers wait for the next tick.
* chore(lint): ratchet budget ceilings down for fixed violations
* fix(reset_budget_job): harden chunk loop, fail open on redis errors, heartbeat the lock
Review fixes on the two prior commits. Reset scans now skip rows with no
budget_duration, so permanently due rows can neither starve a phase nor
have a lifetime cap zeroed every tick. Chunk progress counts rows whose
new budget_reset_at actually cleared the cutoff, so a zero-length
duration cannot burn the per-run chunk cap. A failed lock acquire only
skips the run when another pod verifiably holds the lock; a broken redis
runs unguarded instead of silently disabling resets fleet-wide. Partial
row failures report real progress and fire the failure hook without
killing the phase. The leader re-asserts the lock between phases and
stops if another pod took over, and the budget window advance uses
update_many so a tier deleted mid-chunk cannot abort the transaction.
Lint budget ceilings re-ratcheted for the net-fixed violations.
* fix(reset_budget_job): renew the leader lease and reject non-positive budget durations
Bot review follow-ups. PodLockManager now extends the lock TTL when the
holding pod re-acquires, via an atomic compare-and-expire script with a
plain SET fallback, so a run longer than the TTL keeps its lease instead
of silently sharing the job with another pod. The positive-duration
validation that team member endpoints already had is hoisted to
management common_utils and applied to key, internal user, budget,
customer and team intake, so a tenant can no longer create zero-duration
budgets whose permanently due rows starve other tenants' resets. Such
durations now return 400 at intake; existing rows are untouched.
* refactor(reset_budget_job): defer leader election to a follow-up PR
* fix(reset_budget_job): satisfy strict lint gates
String defaults for the two getenv calls (PLW1508) and the chunk
outcome returns moved to try/else (TRY300).
* fix(proxy): isolate guardrail load failures per row
One DB guardrail row that fails to initialize aborted the whole
_init_guardrails_in_db loop, so a single typo'd guardrail type or a
missing required param left the proxy running with zero DB guardrails
registered and requests that should have been blocked reaching the
provider.
Catch per row around sync_guardrail_from_db, log the guardrail name, id
and error, and continue with the remaining rows. The failing row's id is
still added to db_guardrail_ids before the attempt so reconcile_db_guardrails
cannot mistake a live row for a deleted one.
* test(proxy): drop inline note and record reconcile via a handler double
Replaces the patched bound method with an InMemoryGuardrailHandler subclass
that records what reconcile_db_guardrails received, so the test injects a
double instead of swapping a method on a live object.
Merging staging's flat-cost summary work with the capability gating pushed
EntityUsage.tsx to 815 counted lines, over the 800-line eslint cap. Move the
four pure top-N/rollup helpers to entityUsageAggregations.ts and pass their
inputs explicitly.
TopKeyView and TopModelView were mocked to render static text, so nothing
asserted which breakdown fed which table. The mocks now surface their rows and
a new case pins each table to its own data source.
LITELLM_ENABLE_PTU_COST_ATTRIBUTION, read through get_secret_bool and defaulting to
false, makes the whole PTU flat-cost feature inert unless an operator opts in. The
daily rollup cron is not registered at all, so no sentinel row is ever written;
/model/new and /model/{id}/update reject a request that carries any PTU model_info
field with a 400 naming the fields and the env var rather than dropping them; the
daily activity read path reports zero flat cost; and the model add and edit forms
hide the four PTU inputs.
The read gate lives where flat cost enters SpendMetrics rather than in the aggregated
SQL select. /team/daily/activity, the endpoint the Usage page reads, is served by the
paginated find_many path and never runs that query, so forcing the select to a
constant zero would have left the reporting surface that matters still showing flat
cost.
Sentinel row filtering is deliberately not gated. An operator can enable the flag,
accrue rows under the __ptu_flat_cost__ api_key, then disable it, and those rows stay
in LiteLLM_DailyTeamSpend; gating the filter too would surface the sentinel as a bogus
api_key and mint a provider bucket for its empty provider. Response fields keep their
shape and report 0.0, so typed clients are unaffected, and the migration and the
ModelInfo field declarations are untouched.
The write gate reads the incoming request rather than the merged deployment, so a
model configured during an earlier opt-in stays editable, and the edit form drops the
PTU keys from the payload instead of sending nulls that would clear stored config.
The dashboard reads the flag from a read-only enable_ptu_cost_attribution key on
/get/ui_settings, computed from the environment on every read. It is deliberately not
an allowlisted persisted setting, and PATCH /update/ui_settings rejects it with a 400,
so an admin cannot flip an env-gated feature from the UI.
Two review findings on the gate itself. The PTU clear loop now runs only when the
feature is enabled: the write gate rejects a value but lets an explicit null through,
and a client round-tripping a model_info blob sends the PTU keys as nulls, so a
disabled proxy would have quietly erased a billing configuration set up during an
earlier opt-in. Disabling pauses PTU rather than discarding its setup. And the
dashboard flag is re-read every thirty seconds instead of the hour the other UI settings
use, since those are persisted records while this one tracks the proxy process; a
restart that flips the variable would otherwise leave the model form offering inputs
the backend now rejects. The flag is polled rather than only marked stale, since a form
that stays mounted and focused never refetches on its own.
The read gate checks the row before the flag. It runs once per metric accumulation and a
record fans out across roughly a dozen breakdowns, while the flag reads through the secret
manager uncached, so consulting it for every accumulation put thousands of lookups on a
shared endpoint that made none before. Only a row actually carrying flat cost reaches it.
The timeout contract check skipped a job whenever either budget came from
a `with:` value it could not parse, or from a matrix column no `include`
row supplied as a number. Both paths produced no pairs and no errors, so
the guard printed "invariants hold" for a caller whose budgets were never
compared at all. A caller reading `${{ matrix.timeout }}` off a mistyped
column while capping the job at 1 minute passed clean.
Unresolvable budgets now come back as the reason they could not be read
and are reported as violations, which is the whole point of a guard built
to catch checks that silently do not run. `Column` tags a matrix
reference so it stays distinguishable from that reason string, and the
report names only the columns that resolve nowhere, since a column every
row supplies is not what left the pair unchecked.
calculate_image_response_cost_from_usage read input_tokens_details with
getattr(), but OpenAI images.edit responses carry it as a plain dict, so
both fields came back None and all input tokens were priced at
input_cost_per_token instead of input_cost_per_image_token (e.g. $5/M
instead of $8/M for gpt-image-2). Read it with the dict-tolerant
_get_token_detail_value helper, as the output side of the same function
already does.
Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>