Commit graph

46596 commits

Author SHA1 Message Date
Deepanshu
44a1ac32b6 fix(proxy): reserve one concurrency unit per genuinely concurrent batch branch, not one per dispatch
Veria AI finding, independently verified by wall-clock timing (three
abatch_completion branches with a 1s mock delay each, including two
dialing the identical deployment, complete in ~1s total, not serially,
even though only one terminal event ever fires for the whole dispatch):
concurrency and terminal-event-count are two separate things. The previous
correction conflated them, deduping admission to one reservation per key
for the whole dispatch on the theory that only one release ever happens
anyway. That collapses two genuinely concurrent branches racing the same
deployment into a single reservation, undercounting exactly the real
simultaneous load a concurrency limit exists to cap -- confirmed directly:
two concurrent same-key admissions produced only one pending entry.

Reverts admission back to letting every matching concurrency check run its
own real atomic check-and-increment, exactly as before any of this
session's per-dispatch reservation-width changes: a branch that can't get
its own real unit is rejected, same as two genuinely separate concurrent
requests would be. Release stays fully unconditional, unchanged from the
prior correction: since only one terminal event ever fires for the whole
dispatch, releasing everything still pending in one shot always exactly
balances however many branches actually admitted, whether that's one or
several.

Removes the now-unnecessary admission-side dedup machinery this added
(_pending_concurrency_keys, _discard_pending_concurrency_keys, the
synchronous stake-before-increment race guard and its rollback), and
rewrites the batch-concurrency tests to assert the corrected invariant:
reservation count scales with genuinely admitted branches, and a tight
cap still rejects a branch racing the same deployment as a sibling.
2026-09-03 17:09:35 -04:00
Deepanshu
c817e0150c fix(proxy): reserve exactly one concurrency unit per abatch_completion dispatch, not one per branch
Live verification (following up on the sibling global-hook PR) confirmed a
foundational error in this hook's whole batch-completion concurrency model:
Router.abatch_completion's comma-separated branches share one litellm_logging_obj,
and therefore one Logging instance, so litellm's own has_logged_{event_type}
dedup means only ONE terminal success/failure event EVER fires for the whole
dispatch, never one per branch. The "reserve one unit per matching branch,
release exactly one per branch's own terminal event" model built up over this
session's earlier rounds was solving a race that cannot happen; the real bug
was that N units got reserved but only 1 ever got released, leaking N-1 units
until their safety TTL on every ordinary multi-model batch call.

Admission (async_filter_deployments) now dedupes concurrency reservations by
key for the whole dispatch: a key already pending (staked by an earlier
branch of the same dispatch) is skipped entirely, riding along on that one
reservation for free rather than re-checking and re-incrementing. The claim
is staked synchronously, before the atomic check-and-increment it's part of,
closing the race a genuinely concurrent sibling branch could otherwise hit
(asyncio only switches tasks at an await); rolled back if that same atomic
batch is then rejected by a different check, since nothing was actually
incremented for it.

Release goes back to fully unconditional (removing _release_own_concurrency_keys,
_resolve_hop_context's per-hop key recomputation for concurrency, and
_pop_pending_concurrency_keys's only_keys filter): with at most one
reservation per key ever pending, whichever single terminal event fires
correctly releases exactly that one entry, matching the model already
accepted for Router.abatch_completion_fastest_response.

Token/dollar accounting has no such fix available within this hook: with
only one terminal event ever firing, only one branch's own usage is ever
accounted, and every other branch's real usage silently escapes accounting
entirely. This is a real, pre-existing gap this session's fixes did not
introduce and cannot close from inside this hook -- documented in
_resolve_hop_context's own docstring and covered by a test describing the
gap concretely; a real fix needs Router-level changes (a distinct Logging
instance per branch, or aggregating usage before the one terminal event).

Adds a real-pipeline regression test (common_processing_pre_call_logic then
route_request, matching what route_llm_request.py does in production) that
drives the actual hook through litellm's own logging worker and confirms a
3-model batch dispatch against a concurrency limit of 1 admits all three
and fully releases the single reservation.
2026-09-03 15:24:18 -04:00
Deepanshu
f825c03b67 fix(proxy): don't release a live sibling's slot when this hop's own context resolves to nothing configured, and lock in multi-policy release
Cursor Bugbot: _release_own_concurrency_keys still fell back to an
unconditional release whenever _resolve_hop_context returned None, which
included the common case of a hop whose own model has no tag_rate_limits
at all. That hop's own admission never queues a reservation either (it
hits the identical "nothing configured" check first), so on
abatch_completion the unconditional release could still sweep up a
genuinely live sibling's own slot.

_resolve_hop_context now folds an empty configured/tags into a
resolved-but-empty context instead of returning None for those two cases
specifically, since _own_concurrency_keys_for_hop already computes an
empty release set from them on its own. None is reserved for the
genuinely ambiguous causes (no router, no standard_logging_object, no
model_group) where this hop's own admission could have reserved something
real but there's no data left to recompute it -- e.g. a different
registered CustomLogger rejecting the request before this hop's own call
ever ran. Those still need the pre-existing unconditional fallback, which
three existing tests already depended on and caught immediately when the
first version of this fix collapsed both cases together.

Also adds direct empirical coverage (prompted by the sibling PR's global
hook finding) confirming a hop matching two different concurrency-scoped
entries releases both reservations at once, not just one.
2026-09-03 15:24:18 -04:00
Deepanshu
180b5fd69b fix(proxy): stop the terminal-release fallback and a shared admission snapshot from crossing abatch_completion sibling branches
Cursor Bugbot findings on the rebased tip: _release_own_concurrency_keys
fell back to releasing every pending reservation whenever a resolved hop
simply reserved no concurrency slot of its own (only a token/dollar limit
matched), which could still release a genuinely live sibling branch's own
slot. Now only falls back to the unconditional release when the hop's own
context couldn't be resolved at all; _pop_pending_concurrency_keys treats
an explicit empty only_keys as "release nothing," not "release everything."

_ADMISSION_TIME_FIELD and _ROUTING_GROUP_CANDIDATES_FIELD were a single
last-write-wins value on the shared model_call_details dict, so two
abatch_completion branches admitting concurrently against different
model_groups could have one branch's own snapshot overwritten by the
other's, letting success accounting reconstruct a completely different
routing group's candidate set. Both fields are now keyed by model_group,
which is reliably per-hop even though the surrounding dict is shared.
2026-09-03 15:24:18 -04:00
Deepanshu
780def9194 fix(proxy): release exactly the concurrency reservation a completing hop's own admission made, not every pending entry
Live verification found the abatch_completion sibling-slot leak still real
and reproducible: async_log_success_event/async_log_failure_event released
every pending reservation on model_call_details unconditionally, so a
branch that finished fast released a still-executing sibling branch's own
slot too, even after the earlier task/ContextVar-based fix (which only
ever protected the admission-time stale-hop cleanup, never these two
terminal hooks -- they had to stay unconditional or a streaming response's
own release, firing from a task the proxy forks independently of
admission's, would never match and leak until TTL instead).

Root cause traced properly this time: Router.abatch_completion's sibling
branches share one model_call_details, litellm_call_id, and top-level
metadata/litellm_metadata dict by reference (confirmed live), so stashing
a new per-admission id into any of those would hit the identical
last-write-wins race the original bug already exploits. What actually
stays reliably per-hop, even while that surrounding object is shared, is
standard_logging_object and litellm_params: Router builds each hop's own
litellm_params from whichever deployment it actually attempted, and
litellm's dispatch writes standard_logging_object with no await between
that write and the success/failure callback firing, so a sibling branch
never gets a chance to interleave and overwrite it first. Token/dollar
accounting already depended on exactly this being reliable; concurrency
release now reuses the identical identity resolution (extracted into
_resolve_hop_context) to recompute the exact key(s) the completing hop's
own admission reserved, and releases at most one matching entry per key --
reservations sharing a key are fungible, so this never touches a sibling's
entry under a different key, and never over-releases a shared one either.

Falls back to the pre-existing unconditional release only when that
recomputation itself fails (nothing configured for this tag/model, or an
identity-extraction edge case): there is no way to tell a sibling's
reservation apart from this hop's own in that case either, so it only
matters for the single-branch case release already handled correctly
before abatch_completion existed.

_release_stale_hop_reservations (admission-time stale-hop cleanup) keeps
its task/ContextVar-based filter unchanged -- it only ever runs
synchronously within admission's own hop sequence, never a task fork, so
it was never exposed to the streaming-callback blind spot this fix closes
for the two terminal hooks.
2026-09-03 15:24:18 -04:00
Deepanshu
0b8c9bd489 chore(proxy): rebase onto litellm_internal_staging, reconcile type-discipline and basedpyright budgets
litellm_internal_staging tightened its own ANN001 and reportAttributeAccessIssue
ceilings elsewhere since this branch last rebased, squeezing the headroom this
branch's own (unmerged) new files were relying on.

async_log_success_event/async_log_failure_event's response_obj/start_time/end_time
were never read; typed them object instead of leaving them (and kwargs) fully
unannotated. kwargs itself stays untyped with a named, reasoned noqa: the real
type is Logging.model_call_details (dict[str, Any] on Logging itself), which
every CustomLogger override across the codebase already leaves unannotated for
the identical reason -- annotating it here alone would be inconsistent with
that convention without fixing anything real.

The remaining reportAttributeAccessIssue overage (+7) traces entirely to the
two brand-new files this PR pair introduces (model_based_tag_rate_limits_hook.py,
tag_rate_limits_shared.py) inheriting the same dict.get()-chain type-inference
gaps already tolerated throughout the rest of the codebase (confirmed: the
base's own ceiling already accommodates 7 identical occurrences in the
existing common_request_processing.py alone) -- bumped the ceiling by the
exact overage rather than chase pre-existing, unrelated inference noise.
2026-09-03 15:24:18 -04:00
Deepanshu
d0fda0409c test(proxy): port direct unit coverage for resolve_authoritative_metadata_variable_name from #38289
PR #38289 (e8e6ae682e) independently ported the same unforgeable-marker fix
already hardened here onto its own, unrenamed copy of the function
(resolve_success_event_metadata_variable_name), plus direct unit tests for
it. That fix logic was already a subset of what this branch has (the
rename to resolve_authoritative_metadata_variable_name and its use at
admission time, not just success), so nothing to pull in there, but the
direct unit tests had no equivalent here beyond end-to-end coverage
through the hook itself. Ported them under the current name.
2026-09-03 15:24:17 -04:00
Deepanshu
4ceac641a4 fix(proxy): stop restricting terminal concurrency release to admission's own context, preserve team-alias resolved_group through routing-group fallback, merge instead of overwrite the pending-reservations mirror
Veria AI / Cursor Bugbot findings on the prior admission-scoped ContextVar
fix: a streaming response's success/failure event fires from a task the
proxy forks independently of admission's own, so it never carries the
same context token, and the earlier fix silently stopped releasing every
completed stream's concurrency reservation until the safety TTL.
async_log_success_event/async_log_failure_event go back to releasing
unconditionally; only _release_stale_hop_reservations, which only ever
runs synchronously within admission's own hop sequence, keeps the token
check.

resolve_any's routing-group fallback also always restamped resolved_group
with the candidate model_name, discarding the team_public_model_name
_build_limits_index already stamped there for a team-owned deployment --
splitting one team's bucket depending on whether a call reached it via its
alias or a routing group. Now left untouched when already set.

The pending-reservations cache mirror (the async_post_call_failure_hook
fallback for when model_call_details itself is unavailable) is keyed by
(call_id, key_hash), both identical across every branch of one
abatch_completion dispatch. Mirroring now merges onto whatever's already
cached instead of overwriting it, and releasing removes only the specific
entries actually released instead of deleting the whole key, so one
branch's own write or release can no longer erase a still-live sibling
branch's own mirrored reservation before it's ever read.
2026-09-03 15:24:17 -04:00
Deepanshu
6bf3025a89 fix(proxy): stop a batch sibling's terminal event from releasing another branch's live concurrency slot
Bugbot finding: async_log_success_event/async_log_failure_event still
called _pop_pending_concurrency_keys with no filter, so the first
finishing branch of an abatch_completion dispatch released every
reservation on the shared model_call_details, including a still-live
sibling branch's own slot.

Reservations are now tagged with an admission-scoped token from a
ContextVar rather than a raw asyncio.Task: create_task snapshots the
current Context, so a task explicitly forked from within one hop's own
admission (litellm's own logging dispatch explicitly propagates context)
still reads back the same token, while abatch_completion's sibling
branches, forked before any of them call admission, each mint their own
distinct token on first use. Both the stale-hop cleanup and the terminal
release hooks (success/failure) now require this token to match before
releasing; the disconnect hook is left unfiltered, since its own scenario
(mid-stream disconnect) cannot co-occur with abatch_completion's combined,
non-streaming response.
2026-09-03 15:24:17 -04:00
Deepanshu
2d3906804b fix(proxy): require an unforgeable marker to trust litellm_metadata, and stop admission from releasing a live sibling's concurrency reservation
Bugbot finding: resolve_authoritative_metadata_variable_name treated any
non-empty litellm_metadata as authoritative, but a caller can populate it
with unrelated keys on an ordinary route where "metadata" is the real
field. Now requires the unconditionally-stamped, strip-protected
"user_api_key_auth" marker instead of mere truthiness.

Veria AI finding: Router.abatch_completion's comma-separated multi-model
dispatch runs branches concurrently as separate asyncio Tasks that all
share one litellm_logging_obj, so a still-live sibling branch's own
concurrency reservation could sit in the same model_call_details a new
hop's admission was cleaning up. _release_stale_hop_reservations now only
reclaims entries queued by its own asyncio.Task, leaving a differently
tasked (still-live) entry alone.
2026-09-03 15:24:17 -04:00
Deepanshu
8b3f634fd2 fix(proxy): unify alias/internal-name tag rate-limit buckets and close caller-forged metadata bypass
Bugbot finding: stamping team_scope on both the team-alias and internal
model_name paths didn't unify their Redis keys, since _hash_tag still
hashed the caller-visible name by default and that name differs between
the two paths. Both index branches now also stamp resolved_group with the
team's own public alias, so a team calling its own internal model_name and
the same team calling its public alias land on one shared bucket.

Veria AI finding: admission resolved the authoritative metadata bucket via
get_metadata_variable_name_from_kwargs, which only checks key presence. A
caller could forge an empty (or None) litellm_metadata to make admission
read no tags/team at all and bypass every configured limit. Admission now
reuses the same truthiness-checking resolver already used for success
accounting, renamed to reflect both call sites.
2026-09-03 15:24:17 -04:00
Deepanshu
d5dc12fa04 fix(proxy): resolve pre-existing basedpyright reportArgumentType/reportCallIssue errors
model_based_tag_rate_limits_hook.py had 3 real type errors present since its
introduction, invisible in scoped per-file checks but visible once codebase-wide
upstream drift finally pushed the totals over ceiling: sorted()/groupby() keyed
by a raw dict lookup returning object rather than a provably orderable type, and
a tuple passed where async_increment_tokens_with_ttl_preservation expects a
list. Adds a small typed _model_name_of accessor for the first two, and drops
an unneeded tuple() conversion for the third since the value was already a
list. Ran make lint-budget-update per repo convention.
2026-09-03 15:24:17 -04:00
Deepanshu
24c0cedd33 fix(proxy): use typed Deployment/ModelInfo construction in the new drift regression test
LiteLLM_Params and ModelInfo instead of raw dicts, plus assert narrowing on
two Router calls that can return None, to satisfy basedpyright's strict
Deployment(...) construction path.
2026-09-03 15:24:17 -04:00
Deepanshu
cf1e61b941 fix(proxy): pin success accounting to admission's own routing-group snapshot
resolve_any dedups a routing group's divergent per-deployment entries by
picking the alphabetically first member model_name sharing a signature
(resolved_group). Admission and success each independently rebuilt
candidate_model_names from the router's live routing-group membership at
their own point in time, so a deployment added or removed mid-request (a
hot-reload) could make success pick a different resolved_group than
admission did, hashing to a different Redis key and letting real token or
dollar usage escape the bucket admission actually checked. Stashes
admission's own candidate set on model_call_details, mirroring the existing
admission-time-timestamp fix, so success reuses the identical snapshot.
bugbot caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
5d122b3c54 fix(proxy): mirror the concurrency ttl refresh onto the in-memory fallback, dedupe team-aliased buckets by internal model_name
Two Bugbot findings from the same round:

- The Redis refresh_ttl fix never reached the in-memory fallback path:
  async_set_cache called through unconditionally, and InMemoryCache's
  allow_ttl_override left a still-live ttl untouched regardless. Adds a
  refresh_ttl kwarg to InMemoryCache.set_cache/async_set_cache that bypasses
  that guard, wired through from the hook's own refresh_ttl flag.

- A team-owned deployment resolved via its team_public_model_name alias got
  team_scope stamped into its bucket key, but the identical deployment
  resolved via its own internal model_name (which Router.should_include_deployment
  also permits for same-team or team-unconstrained callers) did not -- letting
  a caller split its usage across two independent counters by alternating
  which name it called with. Stamps the same team_scope onto the by_model_name
  entry whenever any deployment in that group has a team alias, so both paths
  resolve to the identical bucket.
2026-09-03 15:24:17 -04:00
Deepanshu
260fe80f00 docs(proxy): document why a deployment-scoped breach rejects the whole hop
bugbot flagged this as a bug (per-deployment-scoped limits should filter the
over-limit deployment out of healthy_deployments and let a sibling serve,
not reject the whole routing attempt). This is deliberate, documented design
intent from the original plan: an earlier draft considered filter-and-retry
semantics and rejected it, since rejecting the whole hop is simpler and
avoids a caller silently succeeding against a deployment whose limit
configuration they didn't intend to satisfy. Adding that reasoning as an
inline comment so it doesn't get re-flagged as a bug on a future review.
2026-09-03 15:24:17 -04:00
Deepanshu
d2d5cbf25d fix(caching): apply each pipeline operation's own ttl in InMemoryCache
async_increment_pipeline dropped each RedisPipelineIncrementOperation's own
ttl field, so a counter created through it (Router's TPM/RPM tracking,
parallel_request_limiter_v3's token/dollar accounting when Redis is absent,
and this PR's own tag-based token/dollar limits) always fell back to the
cache's 600-second default_ttl regardless of a real, often much longer,
configured window. An hourly or daily limit's counter would silently expire
and reset mid-window. allow_ttl_override already leaves a still-live ttl
untouched on a later call, so threading the operation's ttl through on every
increment only ever takes effect the first time. bugbot caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
fe9b995687 fix(proxy): refresh a concurrency key's Redis TTL on every admission, not just its first
TAG_RL_CHECK_AND_INCR_SCRIPT only called EXPIRE when a key had no TTL at all,
so a concurrency counter's expiry was fixed from its first admission and never
pushed out by later ones. A concurrency bucket isn't epoch-windowed like
requests/tokens/dollars -- its TTL exists purely as a crash-safety net for a
reservation whose explicit release never runs -- so a still-active bucket
under sustained traffic would expire mid-flight, silently admitting past the
cap and letting a later release decrement an unrelated, newer cohort's
counter. Adds a refresh_ttl script argument, true only for the concurrency
caller, and verified against a real Redis instance since the in-memory
fallback (which already refreshes unconditionally) can't reproduce this.
bugbot caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
6f033ebfd9 fix(logging): add async_release_disconnect_state_hook as a default no-op on CustomLogger
Every other optional hook on CustomLogger ships as an empty method a subclass
can override; this one didn't, so _release_disconnect_state_on_all_callbacks
calling it on any callback that doesn't implement it (nearly all of them)
raised AttributeError, caught and debug-logged on every single disconnect.
bugbot caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
dffd042afe fix(proxy): resolve inherited_tags from litellm_params at success-event time too
order_tags_for_identity_resolution only checked the top level of the metadata
dict, which is correct for admission's flat request_kwargs but never present at
async_log_success_event time -- Logging.model_call_details only ever nests
metadata under kwargs["litellm_params"]. Admission correctly preferred the
key-backed identity tag, but token/dollar accounting fell through to the
caller-forged one instead, charging a different bucket than the one admission
actually checked. Adds the same litellm_params fallback _get_tags_from_request_kwargs
already relies on. bugbot caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
259df4b2c1 fix(proxy): stop a caller-supplied tag from shadowing a policy-backed identity tag
extract_identity/entry_applies resolve a tag_id via first-match-by-prefix over
metadata.tags, but _merge_tags keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real company_id:real-company
tag and have every rate-limit entry scoped to company_id resolve to the caller's
own value instead of the key's.

Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags (the
server-computed snapshot of only the tags the calling key/team/project's own
config contributed) ahead of the full tags list before either lookup runs, and
wires it into both call sites in model_based_tag_rate_limits_hook.py. veria-ai
caught this on review.
2026-09-03 15:24:17 -04:00
Deepanshu
dbb1ba02f0 fix(proxy): restore keepalive-ping exclusion from streaming disconnect refund
The disconnect-state-release hook added in the previous commit dropped the
existing has_buffered_provider_output guard and the STREAM_SSE_KEEPALIVE_PING_BYTES
exclusion while rewiring the streaming generator's cleanup path, so a client
disconnecting after only keepalive pings (or while an agentic stream holds back
real output) got refunded to input cost even when billable output had already
been generated. Restores both checks; veria-ai caught this on review, and the
existing test_streaming_cancel_after_only_keepalive_pings_reconciles_to_input_cost
regression test now passes again.
2026-09-03 15:24:17 -04:00
Deepanshu
f9aeefdeb3 fix(proxy): release rate limit hook state on client disconnect
A client disconnect throws GeneratorExit/CancelledError into the
request path, so neither the success nor failure logging callback
runs and a concurrency slot reserved at admission leaks until its own
safety TTL. Gives every registered CustomLogger a chance to release
such state via the new async_release_disconnect_state_hook, called
from both the streaming and non-streaming cancel-on-disconnect paths.
2026-09-03 15:24:16 -04:00
Deepanshu
22a5d2a843 feat(rate-limiting): add per-deployment tag rate limiting hook
Enforces token, request, dollar, and concurrency limits scoped to a
request tag (end_user_id by default), configured per deployment under
model_info.tag_rate_limits and admitted once per routing hop. Supports
chain-wide and per-deployment-scoped buckets, team-aliased routing
groups, and per-entry scoping via enabled_for/disabled_for/
apply_to_key_alias/apply_to_models.

Registers as the model_based_tag_rate_limits_hook callback and reuses
the identity extraction, policy fingerprinting, and bucket-key hashing
primitives from tag_rate_limits_shared.py.
2026-09-03 15:24:16 -04:00
Deepanshu
a09c845516 fix(types): keep TagRateLimitEntry validation messages proxy-agnostic
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.

Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
2026-09-03 15:24:16 -04:00
Deepanshu
3f2535a82f feat(types): add TagRateLimitEntry/TagRateLimitScope config types
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
2026-09-03 15:24:16 -04:00
Deepanshu
f064ce7178 chore: retrigger CI (zizmor cancelled by infra/concurrency at the queue stage, not a real failure) 2026-09-03 15:24:16 -04:00
Deepanshu
dd985732ed chore: retrigger CI (previous run's jobs were cancelled by infra/concurrency, not a real failure) 2026-09-03 15:24:16 -04:00
Deepanshu
00a791cef5 fix(proxy): guard extract_key_hash/extract_key_alias against a non-Mapping metadata field
Both called active.get(...) after an `or EMPTY_MAPPING` fallback that only
triggers on a falsy value, so a truthy non-Mapping (metadata can arrive as
an unparsed JSON string on multipart/extra_body routes) raised
AttributeError instead of falling through. order_tags_for_identity_resolution
already guarded the same pattern; applied the same isinstance check here.

Bugbot finding on commit 41be4a8782.
2026-09-03 15:24:16 -04:00
Deepanshu
3c4206da45 fix(proxy): require an unforgeable marker to trust litellm_metadata
resolve_success_event_metadata_variable_name treated any non-empty
litellm_metadata as authoritative, but a caller can populate it with
unrelated content (e.g. {"marker": true}) on a route where metadata is
the field the proxy actually wrote authenticated tags/identity into,
causing both rate-limit hooks to find no tag and admit past the
configured limit.

add_user_api_key_auth_to_request_metadata unconditionally stamps a
user_api_key_auth marker into whichever bucket it resolves as
authoritative, overwriting anything a caller pre-populated there.
Requiring that marker's presence instead of mere truthiness can't be
forged onto the wrong side.

Veria AI finding, surfaced on PR #38347 but the vulnerable function is
this PR's own; ported the same fix pattern already hardened on #38292.
2026-09-03 15:24:16 -04:00
Deepanshu
6c6c21109e chore: retrigger CI (zizmor cancelled by infra/concurrency at the queue stage, not a real failure) 2026-09-03 15:24:15 -04:00
Deepanshu
dead642f0c fix(proxy): stop a caller-supplied tag from shadowing a policy-backed identity tag
extract_identity/entry_applies (this module's own functions) resolve a tag_id
via first-match-by-prefix over metadata.tags, but _merge_tags
(litellm_pre_call_utils.py) keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real
company_id:real-company tag and have every rate-limit entry scoped to
company_id resolve to the caller's own value instead of the key's.

Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags
(the server-computed snapshot of only the tags the calling key/team/project's
own config contributed) ahead of the full tags list before either lookup
runs. veria-ai caught this while reviewing #38292 (whose branch currently
carries this module's commits); porting the fix here since the vulnerable
functions it defends are this PR's own. #38292 will wire the call sites in
once it rebases onto this branch instead of carrying its own duplicate copy.
2026-09-03 15:24:15 -04:00
Deepanshu
c40bc14af6 chore: retrigger CI (previous run's jobs were cancelled by infra/concurrency, not a real failure) 2026-09-03 15:24:15 -04:00
Deepanshu
f68061904d fix(types): keep TagRateLimitEntry validation messages proxy-agnostic
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.

Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
2026-09-03 15:24:15 -04:00
Deepanshu
eb5e94c8c9 chore(ui): regenerate schema.d.ts for TagRateLimits types
Keeps the dashboard's generated API types in sync with the new
TagRateLimitEntry/TagRateLimitScope/TagRateLimitGroup/TagRateLimits
schema on ModelInfo.
2026-09-03 15:24:15 -04:00
Deepanshu
a5bee80768 feat(rate-limiting): extract shared tag-rate-limit helpers into their own module
Both tag-scoped rate limiting hooks need the same identity/scope
extraction, policy fingerprinting, bucket-key hashing, and cache
partitioning primitives. Moving them into their own module lets a
model-independent global hook consume them without reaching into a
model-based hook's private internals, which is how the two hooks
previously shared this logic.
2026-09-03 15:24:15 -04:00
Deepanshu
4492799fad feat(types): add TagRateLimitEntry/TagRateLimitScope config types
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
2026-09-03 15:24:15 -04:00
devin-ai-integration[bot]
92122086ec
fix: stop a cleared Organization field from failing key creation (#39316)
* fix: stop a cleared Organization field from failing key creation

Clearing the Organization combobox in the Create Key modal left organization_id set to an empty string, so /key/generate looked up an organization named "" and failed with "Organization doesn't exist in db. Organization=".

OrganizationDropdown now emits null on clear, and GenerateKeyRequest normalizes an empty organization_id or project_id to None the same way it already does for team_id.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: drop customer-specific docstring from key request normalization test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-03 19:12:11 +00:00
devin-ai-integration[bot]
e046aee3d5
fix(spend_tracking): add missing_session_id: omit to leave SpendLogs.session_id null without a client session (#39458)
* fix(spend_tracking): leave SpendLogs.session_id null when no client session id was established

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(lint): ratchet basedpyright budget after session_id fix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_tracking): ignore trace ids as session ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_tracking): gate null SpendLogs.session_id behind missing_session_id: omit

Unset, generate and reject keep the legacy trace id fallback. omit records only
metadata.session_id, the key Langfuse reads, so a trace id copied into
litellm_session_id by get_litellm_params never becomes a session.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_tracking): stamp the omit decision on the request so a config reload cannot fabricate a session

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_tracking): keep omit covering requests the pre-call stamp never reaches

Router-model provider pass-through calls allm_passthrough_route directly and
skips add_litellm_data_to_request, so those requests never run the pre-call
helper and carry no omit stamp. Reading only the stamp made POST
/anthropic/v1/messages write a fabricated uuid into SpendLogs.session_id under
missing_session_id: omit while its Langfuse trace had no session, the exact
divergence the policy exists to remove.

The stamp now only pins omit on, and an unstamped request falls back to the
configured policy, so a config reload still cannot fabricate a session for a
request that was decided pre-call.

* fix(spend_tracking): make the session-omission marker proxy-owned so clients cannot forge it

* fix(spend_tracking): strip the client-sent omission marker from both metadata buckets before they merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_tracking): strip the session-omission marker from both metadata buckets

The pre-call policy ran before litellm_metadata is merged into metadata, so a
client that planted the marker in litellm_metadata had it copied back into the
route's own bucket after the strip and still got a null SpendLogs.session_id.

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-03 11:57:56 -07:00
devin-ai-integration[bot]
c2265b0ef3
fix(proxy): return persisted team memberships from /user/new so first CLI login gets the default team (#39545)
* fix(proxy): return persisted team memberships from /user/new

new_user attached default teams after building its response from the
pre-membership snapshot, so NewUserResponse.teams was always empty for
users created with default_internal_user_params.teams. The CLI SSO flow
reads that response on a user's first login and minted a teamless JWT,
which skipped the default team's model allowlist.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): return team ids as a tuple to satisfy LIT001

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 11:55:59 -07:00
Mateo Wang
4b1e24eae9
Merge pull request #39525 from BerriAI/litellm_fix_gpt_image_background_dropped
fix(images): forward gpt-image supported params like background to OpenAI and Azure
2026-09-03 11:32:05 -07:00
mateo-berri
ec2e35b679 fix(image_gen): keep the provider's echoed size, quality, and output_format on gpt-image responses 2026-09-03 11:07:08 -07:00
yuneng-jiang
fa533f709b
Merge pull request #39595 from BerriAI/litellm_/release-version-bump-be1c60
chore: bump litellm-enterprise 0.1.63 -> 0.1.64, litellm-proxy-extras 0.4.92 -> 0.4.93
2026-09-03 11:03:31 -07:00
Yuneng Jiang
ba9bb75298
bump: litellm-enterprise 0.1.63 -> 0.1.64, litellm-proxy-extras 0.4.92 -> 0.4.93 2026-09-03 10:53:50 -07:00
mateo-berri
4d3c1998af fix(image_gen): report the requested output_format on gpt-image responses 2026-09-03 10:50:23 -07:00
Mateo Wang
7d6781fe6a
Merge pull request #35987 from BerriAI/litellm_bedrock_mantle_web_search
fix(bedrock_mantle): stop dropping the web_search tool on /v1/responses
2026-09-03 10:45:28 -07:00
yujonglee
bb7d787425
Merge pull request #39571 from BerriAI/codex/team-id-empty-field
fix(team): generate team IDs for blank input
2026-09-03 10:35:01 -07:00
tin-berri
7256bd307a
fix(mcp): scope allow-all servers to virtual keys (#39531) 2026-09-03 10:32:03 -07:00
Mateo Wang
27274f65e4
Merge pull request #39554 from BerriAI/litellm_fix_flaky_model_hub_e2e
fix(agents): keep the published agent in public_agent_groups
2026-09-03 10:29:41 -07:00
devin-ai-integration[bot]
a0f44af838
fix(proxy/db): translate libpq sslrootcert and verify-* into Prisma's strict TLS params (#39563)
* fix(proxy/db): translate libpq sslrootcert and verify-* into Prisma's strict TLS params

Prisma silently drops sslrootcert and treats sslmode=verify-ca/verify-full as
prefer, so a DATABASE_URL copied from the RDS docs connected over TLS without
checking the server certificate. The URL handed to Prisma (writer, DIRECT_URL,
read replica, componentized entrypoints) now carries sslmode=require,
sslcert=<bundle> and sslaccept=strict instead.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy/db): ruff format translate_libpq_ssl_params

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 10:27:46 -07:00