Commit graph

45217 commits

Author SHA1 Message Date
Deepanshu
7abed5a14c fix(rate-limiting): reject non-positive tag rate limits at config load
limit <= 0 makes the atomic requests/concurrency check reject every
admission and the read-only tokens/dollars check never admit, the same
silent always-block failure mode already guarded against for a
negative-infinity limit. Almost certainly a config typo rather than an
intended policy, so reject it at config load time like NaN/infinity.

Found by Cursor Bugbot on PR #36541.
2026-08-27 18:39:51 -04:00
Deepanshu
cae03ad043 fix(rate-limiting): release concurrency slots on a cross-hook rejection
Both tag-rate-limit hooks raise the identical ProxyRateLimitError shape
(detail.error == tag_rate_limit_exceeded). async_log_failure_event fires on
every registered CustomLogger regardless of which one raised, so each hook's
own skip-on-that-marker special case incorrectly suppressed releasing a
reservation the *other* hook's rejection left pending, leaking the slot
until the safety TTL. Removes the special case from both hooks: a hook's
own rejection never reaches the point where a reservation is queued, so the
existing empty-pending-keys check already no-ops correctly for that case.

Found by Veria AI on PR #36541.
2026-08-27 18:39:51 -04:00
Deepanshu
095321afe4 fix(rate-limiting): regenerate dashboard schema, close lint-budget gaps from rebase
Rebasing onto litellm_internal_staging pulled in 234 new upstream commits;
regenerates ui/litellm-dashboard/src/lib/http/schema.d.ts for the new
apply_to_key_alias field and suppresses/annotates the handful of
mutable-construction and missing-Final findings the stricter, moved
type-discipline budget now surfaces across this PR's own prior commits.
2026-08-27 18:39:51 -04:00
Deepanshu
6663924b2f feat(rate-limiting): global-scope, model-independent tag rate limits
Adds apply_to_key_alias to TagRateLimitEntry (unset applies to every
request; set, restricts an entry to specific virtual keys) and a new
global_tag_rate_limits_hook enforcing tag rate limits via a single
litellm_settings.global_tag_rate_limits config block, once per request
in async_pre_call_hook, before routing. Composes with the existing
scope_by_key_hash field to optionally split a bucket per calling key.
Generalizes the already-shipped, per-deployment model_based_tag_rate_limits_hook
to a config surface that is neither model-scoped nor key-scoped by default.
2026-08-27 18:39:50 -04:00
Deepanshu
402c73cb62 fix(rate-limiting): fold policy identity into the bucket key, reject infinite limits
_hash_tag built the Redis/in-memory counter key from name, tag_id, unit,
and deployment/team scope only. Two entries sharing a name but disagreeing
on limit, period_seconds, or any of the four scoping fields therefore
checked and charged the identical bucket, even though resolve_any and
_build_group_limits already treat that as two distinct policies for dedup
purposes. Adds a policy fingerprint (a fixed-length hash of the fields
that make two entries genuinely different) into the key.

A limit of positive or negative infinity made admission comparisons
degenerate the same way NaN did: +inf never rejects, -inf always does.
Rejected at config load time alongside the existing NaN check.
2026-08-27 18:39:50 -04:00
Deepanshu
db989da2af fix(rate-limiting): close scoping-field dedup gaps and reject NaN limits
resolve_any's own routing-group dedup key omitted the four scoping fields
(included_values, excluded_values, enabled_for, disabled_for), unlike
_build_group_limits's dedup signature which already folded them in --
members disagreeing only on scope collapsed to whichever model_name
sorted first, silently applying the wrong member's policy. Folds the same
four fields into resolve_any's own key.

included_values/excluded_values and TagRateLimitScope.values entered the
dedup signature in raw config order with no normalization, so two
deployments declaring the identical set in a different order were treated
as genuinely divergent policies instead of deduping to one chain-wide
entry. Both now normalize to a sorted, deduplicated tuple at construction.

A NaN limit made atomic requests/concurrency admission admit indefinitely
while read-only tokens/dollars checks rejected every tagged request,
since NaN compares false against every ordering operator either way.
Rejected at config load time.
2026-08-27 18:39:50 -04:00
Deepanshu
1b1c22991f feat(rate-limiting): scoped/conditional tag-rate-limit entries via included/excluded values and enabled/disabled_for
Adds included_values, excluded_values, enabled_for, and disabled_for as
optional fields on a tag rate limit entry, so a single entry can apply to
only a subset of identities: an allow/deny list on the entry's own
resolved identity, and/or a gate on a second, independent tag. Lets a
tiered override (e.g. a company-wide cap with named exceptions) be
expressed directly in config instead of pushing the membership decision
into whatever attaches request tags upstream of the proxy.

Renames the hook and its file from tag_rate_limiter to
model_based_tag_rate_limits_hook (and the matching
tag_rate_limiter_max_in_memory_cache_size setting to
model_based_tag_rate_limits_max_in_memory_cache_size), since this hook is
scoped to per-model tag_rate_limits and a distinct, more general hook
could reasonably share the tag-based-rate-limiting name later.
2026-08-27 18:39:50 -04:00
Deepanshu
840d12f814 fix(rate-limiting): satisfy stricter test-lint and type-discipline gates picked up by rebase
The base branch added a stricter ruff config for tests (PT011/B017) and this
branch's own type-discipline gate flagged a mutable dict type annotation --
neither introduced by this branch's own changes, both surfaced by rebasing
onto a newer base. Narrowed the blind pytest.raises(Exception/ValueError)
assertions to ValidationError with a match on the actual validator message,
and annotated a Redis read-only result as Mapping instead of dict.
2026-08-27 18:39:50 -04:00
Deepanshu
7117ec3a63 fix(rate-limiting): resolve success accounting's team_id the same way admission does
Success accounting read standard_logging_object.metadata.user_api_key_team_id
directly, a separately-constructed field not guaranteed to come from the
same metadata_variable_name-authoritative field admission's own
_extract_team_id uses. A mismatched team_id changes team_scope, which is
hashed into the bucket key, so a team-aliased limit could be checked
against one bucket at admission and accounted against another on success.
Success now calls the same _extract_team_id helper admission does.
2026-08-27 18:39:50 -04:00
Deepanshu
96e53c037f fix(rate-limiting): dedup admission against the full routing group, not just healthy members
Admission derived resolve_any's candidate set from healthy_deployments
(Router's cooldown-filtered list for this hop), while success accounting
reconstructs the full, static group membership. A member merely cooled
down at admission time would shrink admission's candidate set without
changing success's, so the two could dedup to different resolved_group
values and land on different buckets whenever a member was temporarily
unhealthy. Admission now derives its candidate set from the same full
group membership success does.
2026-08-27 18:39:50 -04:00
Deepanshu
d8ea14d8ff fix(rate-limiting): read token/dollar buckets straight from Redis, not a stale in-memory snapshot
Success accounting increments these buckets through a Lua script that
writes directly to Redis, bypassing DualCache's in-memory layer entirely.
Once an earlier admission read had backfilled that key into the in-memory
cache, DualCache's own batch-read treats that non-None hit as authoritative
and never rechecks Redis, so later admissions kept seeing the same frozen
snapshot while the real counter climbed underneath it, silently admitting
traffic past the configured token/dollar limit for up to the in-memory TTL.
Admission now reads these buckets directly off the Redis connection when
one is configured, skipping the in-memory layer that write path never
keeps coherent; the no-Redis, single-process fallback is unaffected since
its increments already go through the same in-memory cache these reads use.
2026-08-27 18:39:50 -04:00
Deepanshu
c2ca9ee2e9 fix(rate-limiting): charge token/dollar usage against the window admission checked
Success accounting recomputed a fresh timestamp instead of reusing
admission's own, so a call slow enough to cross a period_seconds boundary
got admitted against one window's counter but charged into the next
window's fresh, empty one. Stash admission's timestamp on the request's
model_call_details and reuse it at success time so both stages agree on
the same bucket.
2026-08-27 18:39:50 -04:00
Deepanshu
be8ecc7c26 fix(rate-limiting): close three new leak/mismatch findings from this review round
- resolve_any now dedups routing-group candidates in sorted order, not raw
  frozenset iteration order, which depends on the process's hash seed and
  could pick a different resolved_group (and therefore Redis key) across
  worker processes for the identical candidate set.
- Success-side token/dollar accounting now derives scope_by_key_hash's key
  hash the same way admission does (straight from request metadata), instead
  of standard_logging_object's own derived field, which silently drops to
  None whenever the raw key value isn't SHA-256-shaped.
- A concurrency reservation from a failed retry/fallback hop is now released
  at the very start of the next hop's own admission call: litellm only ever
  fires async_log_failure_event once per request (the first failed hop wins,
  every later hop's failure is silently deduped), so a later hop's own
  reservation would otherwise never be released before its TTL.
- The opt-in cancel_on_disconnect path now also releases callback-held
  per-request state (e.g. a concurrency slot) before converting the
  cancellation to a 499, mirroring the existing streaming-disconnect cleanup:
  asyncio.CancelledError bypasses the normal success/failure logging
  callbacks here too, since it's a BaseException, not an Exception.
2026-08-27 18:39:50 -04:00
Deepanshu
e7d5b51a39 test(rate-limiting): use monkeypatch.setattr instead of manual litellm.callbacks restore
The base added a new test-quality gate (TQ005) flagging direct
litellm.<attr> = ... module-global mutation, even when restored via
try/finally. Switches both disconnect-hook tests to monkeypatch.setattr,
which the gate doesn't flag and which pytest reverts automatically at
teardown, removing the manual restore entirely.
2026-08-27 18:39:50 -04:00
Deepanshu
0c07a26c8f fix(rate-limiting): fix routing-group bucket mismatch and unpinned token accounting task
Two real findings from Veria AI and Bugbot, both independently caught by
both bots:

Success accounting checked a different bucket than admission (High/Low):
resolve_any's dedup stamps resolved_group from whichever member
frozenset(candidate_model_names) yields first, but success accounting for
tokens/dollars only passed the one deployment that actually served as its
sole candidate -- a trivial single-candidate dedup that resolves to that
deployment's own name, which can differ from whichever member admission's
full-group view picked. Success accounting now reconstructs the full
routing-group candidate set via Router._get_routing_group_deployments, so
it lands on the identical bucket admission checked regardless of which
member actually served.

Token/dollar accounting task not retained (Medium/Low): the same
GC-before-running gap the previous commit fixed for concurrency release
also applied to this hook's other fire-and-forget task -- token/dollar
usage accounting, fired per cache partition with a bare asyncio.create_task
and no strong reference. Renamed _BACKGROUND_RELEASE_TASKS to the more
general _BACKGROUND_TASKS and wired this task through it too.

Also investigated Bugbot's "concurrency TTL never refreshes" finding
(TAG_RL_CHECK_AND_INCR_SCRIPT only sets EXPIRE when Redis reports TTL -1,
so a bucket that already has a countdown running never gets it extended by
a later reservation). Confirmed real and traces to the very first commit
introducing this file, predating this session entirely. A correct fix
needs the shared atomic check-and-increment script to distinguish
concurrency's "extend the TTL on every new reservation" semantics from
requests' "never extend, let the fixed window expire on schedule"
semantics, since both units share this same script -- flagging as a
follow-up rather than rushing a change to shared, security-sensitive
admission logic.
2026-08-27 18:39:50 -04:00
Deepanshu
533ae89c32 fix(rate-limiting): dedup routing-group unions and pin down a fire-and-forget release task
Two real findings from Bugbot:

Routing groups charge every member (High): resolve_any unions limits
across every distinct model_name reachable from a routing-group hop's
healthy_deployments, then async_filter_deployments atomically checks and
increments every returned entry for that one hop -- even though only one
member deployment ends up serving. Members declaring an identical
signature and scope now dedup to one shared entry, so the hop reserves
capacity once, not once per member. Members that genuinely disagree on
the limit stay separate, unchanged from before this fix: resolving that
ambiguity needs knowing which deployment gets picked, which isn't known
yet at this admission-time hook.

Success release can leak slots (Medium): async_log_success_event fires
its concurrency release via a bare, unreferenced asyncio.create_task to
keep the hot success path from waiting on a Redis round trip. Per
asyncio.create_task's own docs, the event loop only holds a weak
reference to a task with no other referrer, and by the time this one
would run its keys are already popped out of model_call_details, so a
collected task's release is unrecoverable, not just delayed. Added
_BACKGROUND_RELEASE_TASKS, a module-level set holding a strong reference
for exactly as long as each release task is pending, discarding it via
the task's own done-callback once it completes.
2026-08-27 18:39:50 -04:00
Deepanshu
f70e48872f fix(rate-limiting): concurrency reservations never released on real HTTP requests
Live-proxy verification of the disconnect-leak fix surfaced a far more
severe, pre-existing bug: tag_rate_limiter's pending-concurrency-key
handoff used a contextvars.ContextVar to pass reservations from admission
(async_filter_deployments) to release (async_log_success_event /
async_log_failure_event). The real proxy request pipeline forks the
streaming response through several distinct asyncio Tasks (create_response's
disconnect race, the streaming generator's own task, ...); a ContextVar only
propagates into tasks forked after a value is set, so release ran in a task
that never saw admission's write. Confirmed via task-id tracing on a live
proxy that even a normal, fully-completed streaming request never released
its concurrency slot -- not just the disconnect case.

Replaces the ContextVar with a field directly on the request's own
Logging.model_call_details dict, which is explicitly passed by object
reference through both admission's request_kwargs and release's kwargs
(confirmed identical object identity on a live request), so it survives
task boundaries by construction. Deliberately not keyed by litellm_call_id
instead: that field is caller-controlled via the x-litellm-call-id header,
and an earlier design already tried and rejected that approach for exactly
this reason (letting unrelated concurrent requests merge reservations).

async_release_disconnect_state_hook now takes request_data so it can reach
the same model_call_details. Rewrote the concurrency-release tests to wire
a shared model_call_details across admission/release (mirroring production)
instead of relying on ambient task context, and re-verified live against a
real proxy: both normal completion and disconnect-before-first-chunk now
correctly free the slot.
2026-08-27 18:39:50 -04:00
Deepanshu
9b3891ee9a fix(rate-limiting): release tag concurrency reservations on client disconnect
A client disconnecting before the first streamed chunk raises CancelledError/
GeneratorExit, which bypasses both async_log_success_event and
async_log_failure_event -- the only two places tag_rate_limiter releases a
concurrency reservation queued at admission. Without a release, the
reservation sits held until the 1-hour safety-net TTL, letting a caller
exhaust their own tag's concurrency limit for free by repeatedly opening and
dropping streaming requests.

Adds an optional, default-no-op async_release_disconnect_state_hook on
CustomLogger, wires it into the proxy's shielded streaming-disconnect
cleanup (only when no disconnect-time success event already fired), and
implements it in tag_rate_limiter to release pending reservations.
2026-08-27 18:39:50 -04:00
Deepanshu
69004900b1 style(rate-limiting): fix a LIT010 rebind the base's ratcheted-down budget now flags
configured_limit/tag_value were rebound via unpacking after already
being bound by an earlier for-loop over the same atomic_checks
sequence; only the first unpacking binding of a name is exempt from
requiring Final. Renamed to failing_limit/failing_tag_value, and used
the literal `_` placeholder (the only underscore-prefixed name this
gate's Final-exemption recognizes) for the discarded third element
instead of reusing `_key`.
2026-08-27 18:39:50 -04:00
Deepanshu
66def53981 fix(rate-limiting): revert refunding the raising key's own ambiguous outcome
The previous fix (refund through the raising index inclusive) was
itself a regression: these are shared, chain-wide buckets with no
per-request ownership tracking, so decrementing a key whose own
increment might or might not have committed is just as likely to erase
a different, legitimately-admitted concurrent request's charge on the
same key as it is to undo our own. An attacker could repeatedly cancel
requests to deliberately erase other callers' charges and exceed the
configured limit -- worse than the alternative this reverts to: a
committed-but-unrefunded key self-heals via its own TTL. Only strictly
earlier admissions in the same batch (this request's own
confirmed-successful increments, never ambiguous) are refunded now.
2026-08-27 18:39:50 -04:00
Deepanshu
23b2a750fe fix(rate-limiting): refund the current key too when its own admission call raises
The earlier exception-refund fix only rolled back indices before the
one that raised, on the assumption a raise meant nothing committed for
that key. That's not guaranteed: Redis can commit the INCRBY and still
have the call raise if the response back to us is lost (a timeout, a
dropped connection), which the caller can't tell apart from a call that
never reached Redis. Now refunds through the raising index inclusive;
a clean rejection (no exception) still only refunds the earlier ones,
since TAG_RL_CHECK_AND_INCR_SCRIPT guarantees that path never committed.
2026-08-27 18:39:50 -04:00
Deepanshu
f3078fe6f8 style(rate-limiting): shorten two more LIT002 annotations the base's ratcheted-down budget now flags
Same class of issue as the earlier LIT001 fix: ruff format wrapped a
setdefault(...)-then-append pattern across lines, separating the
mutable-list construction from its justification comment. Shortened
variable names so both fit on one line stably under both ruff format
and the gate.
2026-08-27 18:39:50 -04:00
Deepanshu
d9eee566fa fix(rate-limiting): hash the caller-controlled tag identity before it enters a cache key
tag_value has no length or content bound before this hook embeds it
directly into an in-memory dict key (bypassing max_in_memory_cache_size,
which caps item count, not key bytes) and an uncapped Redis key. Hashing
to a fixed-length digest bounds this hook's own contribution to key size
regardless of the caller's input, while preserving distinctness.

Existing tests that hand-wrote the raw tag value into an expected key
string now build it through the real key-construction helper instead of
hardcoding the (now-hashed) internal format.
2026-08-27 18:39:50 -04:00
Deepanshu
c6f940312c fix(rate-limiting): refund earlier admissions when a later one raises or is cancelled
_atomic_check_and_increment's refund loop only ran on a normal
rejection return, not when a later key's own admission raised (a
transient Redis error) or the coroutine was cancelled mid-call.
Everything admitted earlier in that batch stayed permanently charged --
for concurrency, a leaked reservation the caller never releases,
incorrectly throttling that tag for up to the 1-hour safety TTL. A
try/finally now refunds every earlier admission before the exception
propagates, covering both raised exceptions and cancellation uniformly.
2026-08-27 18:39:50 -04:00
Deepanshu
43654422f5 style(rate-limiting): shrink a LIT001 annotation the base's ratcheted-down budget now flags
The base branch's type-discipline ceiling tightened from unrelated
merged PRs, newly flagging an annotated dict declaration ruff format
had to wrap across lines. Added a _PartitionOperations type alias so
the declaration fits on one line the gate can associate its
justification comment with.
2026-08-27 18:39:50 -04:00
Deepanshu
3b0f8287f4 fix(rate-limiting): dedupe a deployment's own repeated entry before counting declarations
A deployment declaring the identical concurrency_limits entry twice
appended its own id twice, inflating len(declaring_ids) past
total_deployments. That made is_chain_wide false for an entry every
deployment actually agreed on, and a non-chain-wide concurrency entry
is silently dropped entirely rather than degraded -- disabling
enforcement instead of just scoping it.
2026-08-27 18:39:50 -04:00
Deepanshu
93590434f4 fix(rate-limiting): remove an unnecessary forward-reference quote (ruff UP037)
self.keys: T = value annotations inside a method body are never
evaluated at runtime, so the "_PartitionKey" forward-reference quoting
was unneeded despite _PartitionKey being defined later in the file.
2026-08-27 18:39:50 -04:00
Deepanshu
b4f6e49253 style(rate-limiting): satisfy ruff format without losing LIT001 annotations
A prior `ruff format` run wrapped several annotated mutable-dict/list
declarations across multiple lines, moving their `# mutable-ok`
justification off the line the type-discipline gate checks. Shortened
the annotations (a new _DedupSignature type alias, trimmed comments) so
they fit on one line under both ruff format and the gate.
2026-08-27 18:39:50 -04:00
Deepanshu
79e66d9754 fix(rate-limiting): stop leaving a permanent Redis key on an expired-reservation release
Releasing a reservation whose key had already expired made INCRBY
recreate it, and the floor-to-zero SET left that recreated key
permanently in Redis (SET clears any TTL). Verified against a real
Redis instance: DEL removes the key outright instead, which reads back
identically to 0 everywhere this key is read.
2026-08-27 18:39:50 -04:00
Deepanshu
f6f65e181b fix(rate-limiting): reject key_ttl_seconds shorter than period_seconds
A key_ttl_seconds override below period_seconds expired the bucket key
before its window rolled over, resetting the counter to zero mid-window
and letting tagged traffic exceed the configured limit. Only
period_seconds and above is now accepted.
2026-08-27 18:39:49 -04:00
Deepanshu
e83886a9d3 test(rate-limiting): verify scope_by_key_hash composes with the new per-entry overrides
Confirms _partition_key treats scope_by_key_hash as distinguishing (two
entries identical otherwise but differing only on this flag must not
share a partition), and that per-key bucket independence still holds
when key_ttl_seconds and max_in_memory_cache_size are also set, going
through the real _build_limits_index path rather than a hand-built
_ConfiguredLimit.
2026-08-27 18:39:49 -04:00
Deepanshu
9e07e685f3 feat(rate-limiting): make the in-memory cache size configurable per tag entry
Adds TagRateLimitEntry.max_in_memory_cache_size so a single
high-cardinality entry can get its own dedicated in-memory cache
partition instead of sharing the hook's single default one, keeping the
knob alongside key_ttl_seconds and the rest of that entry's config
rather than only as a proxy-wide setting. Partitions are keyed by each
entry's full signature, not the override value alone, so two unrelated
entries that happen to pick the same size don't get merged.

tokens/dollars accounting and concurrency-slot release are both
partition-aware too: each partition owns its own v3 handler, and a
concurrency reservation is released against the exact partition it was
incremented on so it can't leak onto the default partition instead.

Also fixes a bug this surfaced: _configured_limit_for_signature
rebuilt each entry from a 5-field dedup signature that excluded
key_ttl_seconds and max_in_memory_cache_size, silently resetting both
to their defaults for every entry reachable through the real indexing
path used by async_filter_deployments and async_log_success_event.
2026-08-27 18:39:49 -04:00
Deepanshu
a2a19cbdd9 fix(rate-limiting): reject invalid in-memory cache sizes, add a per-tag Redis TTL override
An unresolved os.environ/ substitution or a config typo could set
tag_rate_limiter_max_in_memory_cache_size to a negative number or a
string; InMemoryCache raises comparing its size against that value, and
DualCache.async_set_cache swallows the exception, silently disabling
every counter write for this hook without Redis. Only positive integers
are now accepted; anything else falls back to the safe default with a
warning.

Also adds TagRateLimitEntry.key_ttl_seconds so a high-cardinality tag_id
can shed its Redis (or in-memory fallback) keys sooner without
shortening period_seconds itself. Concurrency's safety-floor TTL is
still never lowered by this override.
2026-08-27 18:39:49 -04:00
Deepanshu
fe9e36a0bc feat(rate-limiting): make the tag rate limiter's in-memory cache size configurable
The isolated in-memory cache added for this hook still defaults to 200
entries. A deployment rate-limiting on a high-cardinality tag_id without
Redis can churn past that cap, evicting an active counter before its
period elapses. litellm_settings.tag_rate_limiter_max_in_memory_cache_size
lets that ceiling be raised; 0 is rejected in favor of the safe default
since it would disable the in-memory cache outright.
2026-08-27 18:39:49 -04:00
Deepanshu
e6c504645b fix(rate-limiting): isolate this hook's in-memory cache from the shared proxy cache
internal_usage_cache is the same DualCache instance the proxy's
key/team parallel-request limiter uses for its own authentication-
bound counters, and its default InMemoryCache evicts at 200 items.
Without isolation, a caller flooding this hook's own caller-controlled
tag buckets past that ceiling could evict an unrelated, authentication-
bound counter and let some other caller exceed a limit nothing here
configured. Give this hook a dedicated in-memory layer while still
sharing the real Redis connection when one is configured, so cross-
instance correctness is unaffected.
2026-08-27 18:39:49 -04:00
Deepanshu
c62d79f7df fix(rate-limiting): resolve limits by deployment model_name for routing-group calls
Router deliberately keeps a callable routing-group name (and, per its
own design, a direct deployment-id address) distinct from every member
deployment's own model_name (Router._get_routing_group_deployments).
Since _LimitsIndex only keys by model_name and team alias, a
group-addressed call previously matched neither table and the limiter
silently no-opped for both admission (async_filter_deployments) and
success accounting (async_log_success_event), even though the member
deployments carried real tag_rate_limits under their own model_name.

Add _LimitsIndex.resolve_any(), which falls back to resolving via each
candidate deployment's own model_name when the caller-visible name
matches neither table, stamping each result with the model_name it
came from (_ConfiguredLimit.resolved_group) so hashing stays
namespaced per underlying model_name -- otherwise two different
model_names sharing one routing group with an identically-named,
identically-configured limit would collide onto one Redis counter.

Direct deployment-id addressing has a related but separate, more severe
gap: Router.async_get_healthy_deployments returns a single dict (not a
list) for that path and short-circuits before Router.async_callback_
filter_deployments is ever called, so every CustomLogger.async_filter_
deployments-based hook is skipped, not just this one. That is a
Router-level structural issue affecting many hooks and is out of scope
for this PR.
2026-08-27 18:39:49 -04:00
Deepanshu
8d4528ba37 fix(rate-limiting): resolve the authoritative metadata bucket at success accounting too
async_log_success_event's kwargs is Logging.model_call_details, not the
router's flat request kwargs admission sees: metadata/litellm_metadata
are never top-level there, only nested under kwargs["litellm_params"].
Resolving the field name against kwargs itself always defaulted to
"metadata", so on LITELLM_METADATA_ROUTES (/v1/messages, /responses,
batches, bedrock, files) this read the caller's native, tag-less
metadata instead of the real, server-computed litellm_metadata.tags
admission already used -- silently skipping token/dollar accounting
for every successful call on those routes. Resolve against
kwargs["litellm_params"] instead, mirroring what admission already
does one level up.
2026-08-27 18:39:49 -04:00
Deepanshu
cdcaac75cd fix(rate-limiting): suppress reportPrivateUsage for established cross-module private-helper reuse
Rebasing onto a much-advanced litellm_internal_staging tipped the
basedpyright reportPrivateUsage budget: importing _get_parent_otel_span_from_kwargs,
_PROXY_MaxParallelRequestsHandler_v3, _get_tags_from_request_kwargs, and
_PROXY_TagRateLimiter across module boundaries is the same pattern the
sibling dynamic_rate_limiter_v3 hook already uses (its own imports
just predate this budget check), so suppress with a reason rather than
renaming widely-referenced symbols.
2026-08-27 18:39:49 -04:00
Deepanshu
552b016128 chore(ui): regenerate schema.d.ts for TagRateLimitGroup.limits tuple change 2026-08-27 18:39:49 -04:00
Deepanshu
50a11743d8 fix(rate-limiting): use an immutable tuple for TagRateLimitGroup.limits
Pre-existing mutable list[TagRateLimitEntry] field newly tripped the
type-discipline LIT001 ratchet after the rebase lowered its ceiling.
No other code relies on list mutation here (the one reader already
wraps it in tuple()), so switch to tuple[TagRateLimitEntry, ...].
2026-08-27 18:39:49 -04:00
Deepanshu
170aa0604b fix(rate-limiting): scope alias buckets by team, stop caller-forged metadata bypassing key-scoped limits
- _ConfiguredLimit now carries team_scope: two teams can publish the
  identical team_public_model_name alias, and the limits index already
  scopes lookup by (team_id, alias) correctly -- but the Redis bucket
  key itself never included team_id, so identically-named,
  identically-configured limits from two different teams collided on
  the same counter. Fold team_scope into the hash tag for both
  admission and concurrency keys.
- _extract_key_hash and _extract_team_id no longer OR across metadata
  and litellm_metadata. litellm_pre_call_utils.py writes the real,
  server-authenticated value into only whichever one field is
  authoritative for a given route, leaving the other exactly as the
  caller sent it -- so an OR-fallback let a caller-forged
  metadata.user_api_key (on a route where litellm_metadata is
  authoritative) win over the real hash and bypass every
  scope_by_key_hash=True limit by sending a fresh forged value per
  request. Both extractors now read only the field
  get_metadata_variable_name_from_kwargs names as authoritative,
  matching the pattern this file already uses correctly for tags.
2026-08-27 18:39:49 -04:00
Deepanshu
48b71a8a32 fix(rate-limiting): satisfy type-discipline, basedpyright, and eager-logging gates
- Rewrite tag_rate_limiter.py's dict/list usage to immutable equivalents
  (Mapping/Sequence params, tuple/frozenset/MappingProxyType returns and
  locals, Final everywhere) to clear the ruff-strict type-discipline
  budget (LIT001/LIT002/LIT010), keeping the pending-concurrency-keys
  holder mutable by design with a documented # mutable-ok.
- Rebuild _build_limits_index's grouping via a stable sort + groupby
  instead of a setdefault accumulator; caught and fixed a real bug in
  that rewrite where a per-unit loop re-consumed groupby's already-
  exhausted sub-iterator, silently returning empty limits for every
  group after the first unit.
- Fix a basedpyright reportIncompatibleMethodOverride: match
  async_filter_deployments's signature to CustomLogger's base exactly
  (list/dict, not Mapping/Sequence) since it's an override.
- Fix a second reportGeneralTypeIssues: two Final-annotated locals
  named `key` in sibling branches of the same function tripped
  "previously declared as Final" despite being on mutually exclusive
  paths; renamed them apart.
- Re-fix an eager-built f-string log message (%-style args instead)
  that a prior rewrite pass had inadvertently reintroduced.
2026-08-27 18:39:49 -04:00
Deepanshu
0004fe1b92 fix(rate-limiting): satisfy strict lint gate and regenerate stale schema.d.ts
- Add the missing __init__ return annotation (ANN204) and drop the
  now-unused typing.Any import in favor of a concrete object fallback
  for the optional-otel-Span type alias (TID251), both newly over the
  ruff-strict budget.
- Regenerate ui/litellm-dashboard/src/lib/http/schema.d.ts: the prior
  comment-trimming pass left it out of sync with the trimmed
  TagRateLimitEntry/TagRateLimits docstrings.
2026-08-27 18:39:49 -04:00
Deepanshu
11461f6e45 fix(rate-limiting): reject non-positive period_seconds, fix cross-task concurrency slot leak
- TagRateLimitEntry.period_seconds must be a positive integer: a
  configured 0 previously crashed bucket admission with a
  ZeroDivisionError instead of failing config validation up front.
- Replace the pending-concurrency-keys ContextVar's immutable-tuple
  rebind with a mutable holder shared by reference across every task
  forked off the admitting context. asyncio.create_task only copies
  which object a ContextVar is bound to, not that object's contents, so
  a .set(()) performed inside a detached failure-logging task (e.g.
  after a sibling pre-call-check/filter callback rejects an already-
  admitted hop) was invisible to the parent task that goes on to a
  fallback hop, leaving a stale key that got double-released once the
  fallback's own completion event fired in the parent's context.
  Release now pops an exact snapshot from the shared holder instead of
  rebinding or blanket-clearing it, so a sibling hop's own
  concurrently-appended reservation is never swept up either.
- Trim non-essential comments/docstrings added by this feature to
  match repo convention, keeping only the ones documenting a genuinely
  non-obvious invariant.
2026-08-27 18:39:49 -04:00
Deepanshu
fdf41b49c4 feat(rate-limiting): add tag-scoped token/request/dollar/concurrency rate limits
Port of the tag-based rate limiter from feature/tag-based-rate-limiting
(PR #36459), squashed to the final state of the 18 rate-limiting-specific
commits and rebased onto litellm_internal_staging.
2026-08-27 18:39:49 -04:00
tin-berri
30ff3723b2
feat(model_prices): let a map entry declare its exact reasoning_effort levels (#38481)
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.

Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.

The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
2026-08-27 15:38:01 -07:00
ryan-crabbe-berri
3c41392893
Merge pull request #38574 from BerriAI/litellm_combobox_server_search_hardening
fix(ui): stop server-searched comboboxes from clobbering picks and queries
2026-08-27 15:33:12 -07:00
Mateo Wang
a6816f0e96
Merge pull request #38486 from BerriAI/litellm_together_glm53_flash
feat(together_ai): add zai-org/GLM-5.3-Flash to the model registry
2026-08-27 15:24:30 -07:00
yuneng-jiang
9d1348c1a4
Merge pull request #38575 from BerriAI/litellm_e2e-vision-own-image
test(e2e): serve the vision image from our own fixture
2026-08-27 15:24:09 -07:00
Mateo Wang
4ef1c28877
Merge pull request #38431 from BerriAI/litellm_fix_messages_native_tools
fix(anthropic-adapter): pass provider-native and OpenAI-format tools through on /v1/messages
2026-08-27 15:24:07 -07:00