veria-ai finding on the previous commit: TAG_RL_CHECK_AND_INCR_SCRIPT is
shared with model_based_tag_rate_limits_hook, but only that hook's own
Redis and in-memory call sites were updated to carry refresh_ttl through.
global_tag_rate_limits_hook's own _check_and_increment_one still called the
script with three args and never passed refresh_ttl to InMemoryCache, so a
global concurrency bucket's ttl stayed fixed from its first admission and
could expire mid-flight under sustained traffic, admitting past the cap.
TAG_RL_CHECK_AND_INCR_SCRIPT only called EXPIRE when a key had no ttl at
all, so a concurrency counter's expiry was fixed from its first admission
and never pushed out by later ones. A concurrency bucket isn't
epoch-windowed like requests/tokens/dollars -- its ttl exists purely as a
crash-safety net for a reservation whose explicit release never runs -- so
a still-active bucket under sustained traffic would expire mid-flight,
silently admitting past the cap and letting a later release decrement an
unrelated, newer cohort's counter. Adds a refresh_ttl script argument, true
only for the concurrency caller, and mirrors the same bypass onto
InMemoryCache.set_cache for the no-Redis fallback path (allow_ttl_override
otherwise leaves a still-live ttl untouched).
Also stamps the identical team_scope onto a team-owned deployment's
by_model_name entry whenever any deployment in that group has a team
alias: Router.should_include_deployment lets a same-team caller reach the
deployment by its own internal model_name, not only its
team_public_model_name alias, and both paths must resolve to the same
bucket or a caller could split its usage across two independent counters
by alternating which name it calls with.
Verified against a real local Redis instance (redis-server v8.8.0 on a
scratch port) since the in-memory fallback can't reproduce the Redis TTL
behavior on its own.
async_increment_pipeline dropped each RedisPipelineIncrementOperation's own
ttl field, so a counter created through it (Router's TPM/RPM tracking,
parallel_request_limiter_v3's token/dollar accounting when Redis is absent,
and the tag rate limit hooks' own token/dollar limits) always fell back to
the cache's 600-second default_ttl regardless of a real, often much longer,
configured window. An hourly or daily limit's counter would silently expire
and reset mid-window. allow_ttl_override already leaves a still-live ttl
untouched on a later call, so threading the operation's ttl through on
every increment only ever takes effect the first time.
- Wire _order_tags_for_identity_resolution into this hook's own admission
and success-event tag resolution, matching the sibling model-based hook.
Without it a caller could put a forged tag ahead of the policy-backed
inherited one and dodge or mis-bucket every global limit.
- Drop tag_value from the client-facing rejection detail: it can resolve
from inherited_tags (server-assigned key/team/project metadata), and
echoing it back would disclose that identity to the rejected caller.
- Track every model an admission attempt saw for a call_id (admitted_models,
a frozenset) instead of overwriting a single field, only recording once an
attempt clears every check. Otherwise an apply_to_models-scoped entry that
matched an earlier _pre_call_with_fallbacks attempt lost its success-time
accounting once a later attempt admitted with a different model, and a
rejected attempt's model could wrongly drive later accounting.
- Add async_post_call_failure_hook, releasing a reservation from an earlier
admission attempt when _pre_call_with_fallbacks exhausts every fallback
and re-raises without ever running the real LLM call -- the only other
release paths (success/failure/disconnect) are tied to that call, which
never happens.
_extract_identity/_entry_applies resolve a tag_id by first-match-by-prefix
over the merged tags list, but _merge_tags (litellm_pre_call_utils.py)
keeps caller-supplied tags ahead of key/team/project tags in that list.
An authenticated caller could submit e.g. company_id:attacker-chosen ahead
of the calling key's real company_id:real-company tag and have every
rate-limit entry scoped to company_id resolve to the caller's own value
instead of the key's.
Adds _order_tags_for_identity_resolution, which puts metadata.inherited_tags
(the server-computed snapshot of only the tags the calling key/team/
project's own config contributed) ahead of the full tags list before
either lookup runs, checking both the flat admission-time shape and the
litellm_params-nested success-event shape. Wired into both admission and
success-event tag resolution in this hook; the sibling global hook reuses
this same function (it already imports helpers from this module).
_pre_call_with_fallbacks only checked the original exception for
cross_model_scope before starting the fallback loop; a later fallback
attempt's own cross_model_scope rejection was swallowed by the loop's
plain except-continue and the next, unlisted fallback model silently
served the request instead (baefd27677).
litellm/types/router.py is imported by plain SDK users, not just the proxy;
move the proxy rate-limit hook's internal admission mechanics out of the
raised ValueError messages and into code comments, matching Greptile's
finding on PR #38289 (f6946c67b0).
Cursor Bugbot follow-up finding on the previous commit: _release_stale_hop_reservations
popped _PENDING_REQUEST_INCREMENTS_FIELD unconditionally at the top of a new
hop's admission, but only re-queued it after that hop's own atomic batch
fully succeeded. A hop that failed earlier (a read-only check, or a
different entry in the same atomic batch) left the real counter correctly
charged (a 0.0-increment renewal rolls back to a genuine no-op) but the
bookkeeping field empty, so a later hop's own peek found nothing to renew
and charged a fresh unit on top of the one already sitting in the real
counter.
The field is now peeked, never popped, for "requests" -- concurrency still
pops, since concurrency actively releases on every outcome, but a
"requests" renewal only ever needs to know whether a key is already
charged, never to clear it. Queuing at the end of a hop's own successful
batch now skips keys already present, avoiding unbounded duplicate growth
across a long retry chain. This matches global_tag_rate_limits_hook's
already-correct pattern (an append-only list on the stash, never cleared),
which never had this bug in the first place.
Regression test reuses the same request_kwargs across three hops -- unlike
the existing undercount test, whose final probe uses a fresh, independent
context and so can't distinguish "the real counter is correct" from "the
bookkeeping a future hop needs is intact" -- confirmed red on the popping
version, green on this one.
veria-ai finding on PR #36541: ProxyBaseLLMRequestProcessing._pre_call_with_fallbacks
reruns the whole pre-call pipeline (this hook included) once per fallback
model on any ProxyRateLimitError, not only one this hook itself raised, but
keeps the same litellm_call_id across every attempt. Without this fix, a
call admitted once by this hook and then rejected by a different, later
check in the same pass would get charged again on every fallback retry for
what is still one logical client request, letting a caller consume one
shared tag-quota unit per fallback attempt.
A "requests" or "concurrency" check whose key was already charged/reserved
for this call_id now renews at zero net cost as part of the same
all-or-nothing atomic batch, mirroring model_based_tag_rate_limits_hook's
identical fix for its own per-hop retries.
Renewal requires the repeat admission to carry the same authenticated
key_hash as whichever admission first claimed this call_id's stash:
litellm_call_id is caller-controlled via the x-litellm-call-id header (the
same forgery vector that hook's pending-reservations mirror was hardened
against earlier in this PR), so two unrelated requests choosing an
identical call_id must not be able to renew each other's charge.
Not live-verified: reliably reproducing this specific path needs a second,
different check to reject after this hook has already admitted in the same
pre-call pass, which depends on callback registration order this session
didn't have a clean way to control from proxy config alone. Verified
instead with regression tests against the real hook class covering both the
legitimate renewal case and the forged-call-id case, plus four existing
tests updated to use distinct call_ids for what they model as separate
logical requests now that repeat-call_id admission has real behavior tied
to it.
Cursor Bugbot follow-up finding on the previous commit: refunding a prior
hop's "requests" charge unconditionally at the top of the next hop's
admission, before knowing whether that hop would itself be admitted, could
undercount a logical request. If the next hop went on to fail a different
check (a read-only limit, or another entry in the same atomic batch), the
refund had already committed with nothing to replace it, leaving a
genuine, real attempt charged at zero and letting a caller bypass the cap
under fallback pressure.
_release_stale_hop_reservations no longer refunds a stale "requests" entry;
it returns the stale key instead. A "requests" check whose key matches one
already charged by a superseded earlier hop now renews at zero net cost as
part of the same all-or-nothing atomic batch as every other check on that
hop, rather than as a separate refund-then-recharge sequence. If the batch
rolls back for any other reason, refunding a zero-cost renewal is a
genuine no-op, so the earlier hop's real charge is left exactly as it was.
Regression test reproduces the exact scenario: a hop fails for real, an
unrelated request claims the now-free concurrency slot, this request's
next hop is rejected on concurrency (not requests), and a fresh probe
confirms the original requests charge still stands. Live-verified the
original multi-hop scenario still works correctly end to end.
Cursor Bugbot finding on PR #36541, live-confirmed against a real proxy: a
"requests" limit is meant to cap logical client requests, not internal
routing attempts, but each hop's atomic admission increment had no refund
path of its own. A chain that failed once before succeeding burned two
units of a one-request-per-period budget for a single logical call; live
reproduction showed the retry's own admission rejected with current=1.0
limit=1.0 even though the client only made one request.
Generalizes the same next-hop-releases-a-prior-hop's-stale-reservation
pattern already used for concurrency: a queued "requests" increment still
present when a new hop's admission runs can only belong to an already-failed
earlier hop, so it gets refunded there before this hop's own check runs.
Unlike concurrency, a hop that goes on to succeed (or is the chain's own
final failure) is never refunded, so exactly one unit survives per logical
request regardless of how many hops it took.
An initial version of this fix also popped the new field in
async_log_success_event/async_log_failure_event "for hygiene." That broke
live: litellm's has_logged_async_failure dedup lets the first failing hop's
own failure event through, not only a chain's final failure, so the pop
discarded the entry before the next hop's own admission ever got a chance
to refund it, permanently stranding the charge. Both handlers now leave the
field completely untouched; a dedicated regression test simulates that
exact real failure-event sequence, not just two bare admission calls.
Also fixes a related veria-ai finding on the same file: the pending-
reservations mirror key embedded litellm_call_id (caller-controlled via the
x-litellm-call-id header) directly, with no length bound. Hashed via the
same _fixed_length_identity helper already used for tag values.
veria-ai finding on this PR: a global entry scoped with apply_to_models is
meant to cap an entire fallback chain as one shared unit, but the proxy's
generic local-rate-limit fallback retry (_pre_call_with_fallbacks) caught
that rejection and quietly served the request via any fallback model not
also listed in apply_to_models, defeating the whole point of the field.
global_tag_rate_limits_hook now marks a rejection raised by an
apply_to_models-scoped entry with detail["cross_model_scope"], and
_pre_call_with_fallbacks re-raises immediately on that marker instead of
retrying fallbacks. Every other ProxyRateLimitError caller (parallel
request limiter, budget limiters, model-local tag limits, etc.) is
untouched, since none of them ever set this marker.
Live-verified against a real proxy: an opus-chain to sonnet-chain fallback
configured with apply_to_models: [opus-chain] now gets a 429 on both legs
once the chain-wide cap is hit, instead of silently succeeding via
sonnet-chain.
Live verification against a real proxy showed common_request_processing.py
retries a rejected request against litellm_settings.fallbacks with the model
mutated, re-running this hook fresh. A narrowly scoped apply_to_models entry
can be bypassed this way if the fallback target isn't also listed; listing
every model in the chain closes it, since the fallback hits the same,
already-exhausted shared bucket. Corrects and sharpens the module docstring
accordingly.
Lets a global_tag_rate_limits (or per-deployment model_info.tag_rate_limits)
entry restrict itself to a named list of caller-facing model strings, so one
entry can rate-limit an entire fallback chain as a single shared bucket
instead of requiring the per-hop model_based_tag_rate_limits_hook mechanism.
apply_to_models composes with the existing apply_to_key_alias/enabled_for/
disabled_for gates, all AND together. It is evaluated once, at admission,
against the caller-requested model, and is not re-evaluated if Router later
falls back to an unlisted model for the same request; this is a documented
limitation, not a bug, and is covered by a dedicated regression test.
The cache mirror async_post_call_failure_hook reads to release a fallback
chain's final, chain-exhausting reservation was keyed only by litellm_call_id,
which comes from the caller-controlled x-litellm-call-id header. Two requests
choosing the identical id, even with different tags, would overwrite each
other's mirror entry; if the first request's chain then exhausted, its failure
hook released whatever reservation the mirror currently held, which could be
the second request's still-live one instead of its own. A caller could use
this to release their own concurrency reservation on demand, bypassing their
configured cap, by racing a second self-chosen request against the first.
Fixed by folding the calling virtual key's hash into the mirror key alongside
call_id. async_post_call_failure_hook now uses user_api_key_dict.api_key,
which the proxy's own auth middleware establishes before any hook runs and a
caller cannot forge; async_filter_deployments and the pop/clear path use the
same metadata["user_api_key"] extraction the existing scope_by_key_hash
feature already relies on. A caller can no longer collide with a different
authenticated key's tracked reservation. What remains is a caller colliding
with their own other concurrent request (same key, self-chosen call_id,
different tags), which only weakens enforcement of their own cap and is an
accepted, lower-severity residual not addressed here.
Adds a regression test reproducing the cross-key collision against the
pre-fix code (confirmed to fail for the right reason) and verified against a
live proxy with two real virtual keys racing a shared call_id.
The earlier async_post_call_failure_hook attempt (293aed9b2c) read
model_call_details off request_data["litellm_logging_obj"], but
proxy/utils.py's post_call_failure_hook deliberately pops that key before
invoking any callback ("Remove before callbacks iterate -- not
serialisable"), confirmed live: the hook always no-op'd.
A ContextVar-based mirror was tried next and also confirmed broken live:
its value never reached the task that calls post_call_failure_hook, since
Router's own per-hop execution does not keep that task a descendant of the
one that ran the final hop's admission. Mutating request_kwargs directly is
unsafe too, since litellm forwards unrecognized kwargs to the actual
provider call as extra_body.
litellm_call_id is the one identifier stable across every one of those
objects, so this mirrors the latest hop's own reservation in the same
external cache (Redis or in-memory) the reservations already live in, keyed
by call_id, and reads it back from there instead. Every normal release path
now also clears this same cache entry, so a hop already released the normal
way is never found "stale" and double-released. Verified live for both a
plain and a `stream: true` request whose entire fallback chain fails before
any token is produced: both now reach the real provider on a follow-up call
instead of being rejected by this hook's own stuck reservation.
_partition_key only distinguished entries by tag_id/name/limit/period_seconds/
scope_by_key_hash/max_in_memory_cache_size, but two entries can share all of
those while disagreeing on enabled_for/disabled_for/apply_to_key_alias (the
same class of collision _DedupSignature and _policy_fingerprint already
guard against for dedup and bucket keys). Two such entries setting the same
max_in_memory_cache_size to get their own dedicated partition collided onto
one shared partition instead, letting one entry's high-cardinality traffic
evict the other's active counters. Folds _policy_fingerprint into the
partition key, keeping max_in_memory_cache_size as the trailing element
since _partition_for reads it via partition_key[-1].
litellm's Logging object sets has_logged_async_failure=True after a fallback
chain's first hop fails and blocks async_log_failure_event for every later
hop, so a chain's own final, chain-exhausting failure never reaches that
callback. _release_stale_hop_reservations only cleans up a stale reservation
when a *next* hop's admission runs, which never happens after the last one,
so that hop's reservation sat held for the full safety-net TTL; a caller
repeatedly forcing failures across the whole chain could keep a tag's entire
concurrency budget pinned near-continuously. async_post_call_failure_hook
fires exactly once per proxy request, at the point the proxy gives up and
returns an error to the caller, regardless of how many hops ran or whether
the completion-level callback was suppressed for this one, so it releases
whatever reservation is still pending at that point.
Introduces a _StashByCallId alias so the mutable-ok suppression for the
per-call-id dict lands on one declaration ruff-format won't wrap, instead
of drifting off its annotation line every time the formatter reflows a
long inline type.
The ContextVar-held stash was one shared mutable instance with an
overwritable owner_litellm_call_id field. A nested LiteLLM call made inside
the request (an LLM-judge guardrail, a silent experiment) mints its own
fresh call id but inherits the same context rather than a separate one, so
claiming the stash for that nested call reassigned ownership away from the
outer call; the nested call's own success callback then released the
outer call's still-pending concurrency reservation while the outer request
was still genuinely in flight, letting extra same-tag requests through.
Keyed by litellm_call_id instead, so each call's own reservations and
admission_time are isolated regardless of nesting.
_policy_fingerprint hashed limit/period_seconds/enabled_for/disabled_for/
apply_to_key_alias but not scope_by_key_hash, even though _DedupSignature
already treats it as a distinct policy. _hash_tag's key_hash-derived suffix
is empty whenever key_hash resolves to None, so a key-hash-scoped entry and
an otherwise-identical unscoped entry collided onto the same counter for any
call with no resolvable virtual key.
TagRateLimitEntry no longer has included_values/excluded_values: enabled_for
or disabled_for targeting the entry's own tag_id is functionally identical,
so the extra fields only added surface area. Types, dedup/fingerprint
signatures, hook logic, tests, and the generated dashboard schema are
updated accordingly.
Live proxy verification of the remaining scoping fields surfaced a real bug
in both hooks' async_log_success_event: get_metadata_variable_name_from_kwargs
only checks whether a "litellm_metadata" key is present on kwargs["litellm_params"],
not whether it holds anything. For a plain chat completion, that key is
always present (set to None) alongside the real, populated "metadata" dict,
so token/dollar accounting silently read no tags and no identity, letting
per-tag token and dollar limits go unenforced. Both hooks now resolve the
authoritative field by checking it actually holds a populated dict, matching
the value-truthiness check litellm_logging.py's own tag resolution already
uses, instead of relying on key presence alone.
limit <= 0 makes the atomic requests/concurrency check reject every
admission and the read-only tokens/dollars check never admit, the same
silent always-block failure mode already guarded against for a
negative-infinity limit. Almost certainly a config typo rather than an
intended policy, so reject it at config load time like NaN/infinity.
Found by Cursor Bugbot on PR #36541.
Both tag-rate-limit hooks raise the identical ProxyRateLimitError shape
(detail.error == tag_rate_limit_exceeded). async_log_failure_event fires on
every registered CustomLogger regardless of which one raised, so each hook's
own skip-on-that-marker special case incorrectly suppressed releasing a
reservation the *other* hook's rejection left pending, leaking the slot
until the safety TTL. Removes the special case from both hooks: a hook's
own rejection never reaches the point where a reservation is queued, so the
existing empty-pending-keys check already no-ops correctly for that case.
Found by Veria AI on PR #36541.
Rebasing onto litellm_internal_staging pulled in 234 new upstream commits;
regenerates ui/litellm-dashboard/src/lib/http/schema.d.ts for the new
apply_to_key_alias field and suppresses/annotates the handful of
mutable-construction and missing-Final findings the stricter, moved
type-discipline budget now surfaces across this PR's own prior commits.
Adds apply_to_key_alias to TagRateLimitEntry (unset applies to every
request; set, restricts an entry to specific virtual keys) and a new
global_tag_rate_limits_hook enforcing tag rate limits via a single
litellm_settings.global_tag_rate_limits config block, once per request
in async_pre_call_hook, before routing. Composes with the existing
scope_by_key_hash field to optionally split a bucket per calling key.
Generalizes the already-shipped, per-deployment model_based_tag_rate_limits_hook
to a config surface that is neither model-scoped nor key-scoped by default.
_hash_tag built the Redis/in-memory counter key from name, tag_id, unit,
and deployment/team scope only. Two entries sharing a name but disagreeing
on limit, period_seconds, or any of the four scoping fields therefore
checked and charged the identical bucket, even though resolve_any and
_build_group_limits already treat that as two distinct policies for dedup
purposes. Adds a policy fingerprint (a fixed-length hash of the fields
that make two entries genuinely different) into the key.
A limit of positive or negative infinity made admission comparisons
degenerate the same way NaN did: +inf never rejects, -inf always does.
Rejected at config load time alongside the existing NaN check.
resolve_any's own routing-group dedup key omitted the four scoping fields
(included_values, excluded_values, enabled_for, disabled_for), unlike
_build_group_limits's dedup signature which already folded them in --
members disagreeing only on scope collapsed to whichever model_name
sorted first, silently applying the wrong member's policy. Folds the same
four fields into resolve_any's own key.
included_values/excluded_values and TagRateLimitScope.values entered the
dedup signature in raw config order with no normalization, so two
deployments declaring the identical set in a different order were treated
as genuinely divergent policies instead of deduping to one chain-wide
entry. Both now normalize to a sorted, deduplicated tuple at construction.
A NaN limit made atomic requests/concurrency admission admit indefinitely
while read-only tokens/dollars checks rejected every tagged request,
since NaN compares false against every ordering operator either way.
Rejected at config load time.
Adds included_values, excluded_values, enabled_for, and disabled_for as
optional fields on a tag rate limit entry, so a single entry can apply to
only a subset of identities: an allow/deny list on the entry's own
resolved identity, and/or a gate on a second, independent tag. Lets a
tiered override (e.g. a company-wide cap with named exceptions) be
expressed directly in config instead of pushing the membership decision
into whatever attaches request tags upstream of the proxy.
Renames the hook and its file from tag_rate_limiter to
model_based_tag_rate_limits_hook (and the matching
tag_rate_limiter_max_in_memory_cache_size setting to
model_based_tag_rate_limits_max_in_memory_cache_size), since this hook is
scoped to per-model tag_rate_limits and a distinct, more general hook
could reasonably share the tag-based-rate-limiting name later.
The base branch added a stricter ruff config for tests (PT011/B017) and this
branch's own type-discipline gate flagged a mutable dict type annotation --
neither introduced by this branch's own changes, both surfaced by rebasing
onto a newer base. Narrowed the blind pytest.raises(Exception/ValueError)
assertions to ValidationError with a match on the actual validator message,
and annotated a Redis read-only result as Mapping instead of dict.
Success accounting read standard_logging_object.metadata.user_api_key_team_id
directly, a separately-constructed field not guaranteed to come from the
same metadata_variable_name-authoritative field admission's own
_extract_team_id uses. A mismatched team_id changes team_scope, which is
hashed into the bucket key, so a team-aliased limit could be checked
against one bucket at admission and accounted against another on success.
Success now calls the same _extract_team_id helper admission does.
Admission derived resolve_any's candidate set from healthy_deployments
(Router's cooldown-filtered list for this hop), while success accounting
reconstructs the full, static group membership. A member merely cooled
down at admission time would shrink admission's candidate set without
changing success's, so the two could dedup to different resolved_group
values and land on different buckets whenever a member was temporarily
unhealthy. Admission now derives its candidate set from the same full
group membership success does.
Success accounting increments these buckets through a Lua script that
writes directly to Redis, bypassing DualCache's in-memory layer entirely.
Once an earlier admission read had backfilled that key into the in-memory
cache, DualCache's own batch-read treats that non-None hit as authoritative
and never rechecks Redis, so later admissions kept seeing the same frozen
snapshot while the real counter climbed underneath it, silently admitting
traffic past the configured token/dollar limit for up to the in-memory TTL.
Admission now reads these buckets directly off the Redis connection when
one is configured, skipping the in-memory layer that write path never
keeps coherent; the no-Redis, single-process fallback is unaffected since
its increments already go through the same in-memory cache these reads use.
Success accounting recomputed a fresh timestamp instead of reusing
admission's own, so a call slow enough to cross a period_seconds boundary
got admitted against one window's counter but charged into the next
window's fresh, empty one. Stash admission's timestamp on the request's
model_call_details and reuse it at success time so both stages agree on
the same bucket.
- resolve_any now dedups routing-group candidates in sorted order, not raw
frozenset iteration order, which depends on the process's hash seed and
could pick a different resolved_group (and therefore Redis key) across
worker processes for the identical candidate set.
- Success-side token/dollar accounting now derives scope_by_key_hash's key
hash the same way admission does (straight from request metadata), instead
of standard_logging_object's own derived field, which silently drops to
None whenever the raw key value isn't SHA-256-shaped.
- A concurrency reservation from a failed retry/fallback hop is now released
at the very start of the next hop's own admission call: litellm only ever
fires async_log_failure_event once per request (the first failed hop wins,
every later hop's failure is silently deduped), so a later hop's own
reservation would otherwise never be released before its TTL.
- The opt-in cancel_on_disconnect path now also releases callback-held
per-request state (e.g. a concurrency slot) before converting the
cancellation to a 499, mirroring the existing streaming-disconnect cleanup:
asyncio.CancelledError bypasses the normal success/failure logging
callbacks here too, since it's a BaseException, not an Exception.
The base added a new test-quality gate (TQ005) flagging direct
litellm.<attr> = ... module-global mutation, even when restored via
try/finally. Switches both disconnect-hook tests to monkeypatch.setattr,
which the gate doesn't flag and which pytest reverts automatically at
teardown, removing the manual restore entirely.
Two real findings from Veria AI and Bugbot, both independently caught by
both bots:
Success accounting checked a different bucket than admission (High/Low):
resolve_any's dedup stamps resolved_group from whichever member
frozenset(candidate_model_names) yields first, but success accounting for
tokens/dollars only passed the one deployment that actually served as its
sole candidate -- a trivial single-candidate dedup that resolves to that
deployment's own name, which can differ from whichever member admission's
full-group view picked. Success accounting now reconstructs the full
routing-group candidate set via Router._get_routing_group_deployments, so
it lands on the identical bucket admission checked regardless of which
member actually served.
Token/dollar accounting task not retained (Medium/Low): the same
GC-before-running gap the previous commit fixed for concurrency release
also applied to this hook's other fire-and-forget task -- token/dollar
usage accounting, fired per cache partition with a bare asyncio.create_task
and no strong reference. Renamed _BACKGROUND_RELEASE_TASKS to the more
general _BACKGROUND_TASKS and wired this task through it too.
Also investigated Bugbot's "concurrency TTL never refreshes" finding
(TAG_RL_CHECK_AND_INCR_SCRIPT only sets EXPIRE when Redis reports TTL -1,
so a bucket that already has a countdown running never gets it extended by
a later reservation). Confirmed real and traces to the very first commit
introducing this file, predating this session entirely. A correct fix
needs the shared atomic check-and-increment script to distinguish
concurrency's "extend the TTL on every new reservation" semantics from
requests' "never extend, let the fixed window expire on schedule"
semantics, since both units share this same script -- flagging as a
follow-up rather than rushing a change to shared, security-sensitive
admission logic.
Two real findings from Bugbot:
Routing groups charge every member (High): resolve_any unions limits
across every distinct model_name reachable from a routing-group hop's
healthy_deployments, then async_filter_deployments atomically checks and
increments every returned entry for that one hop -- even though only one
member deployment ends up serving. Members declaring an identical
signature and scope now dedup to one shared entry, so the hop reserves
capacity once, not once per member. Members that genuinely disagree on
the limit stay separate, unchanged from before this fix: resolving that
ambiguity needs knowing which deployment gets picked, which isn't known
yet at this admission-time hook.
Success release can leak slots (Medium): async_log_success_event fires
its concurrency release via a bare, unreferenced asyncio.create_task to
keep the hot success path from waiting on a Redis round trip. Per
asyncio.create_task's own docs, the event loop only holds a weak
reference to a task with no other referrer, and by the time this one
would run its keys are already popped out of model_call_details, so a
collected task's release is unrecoverable, not just delayed. Added
_BACKGROUND_RELEASE_TASKS, a module-level set holding a strong reference
for exactly as long as each release task is pending, discarding it via
the task's own done-callback once it completes.
Live-proxy verification of the disconnect-leak fix surfaced a far more
severe, pre-existing bug: tag_rate_limiter's pending-concurrency-key
handoff used a contextvars.ContextVar to pass reservations from admission
(async_filter_deployments) to release (async_log_success_event /
async_log_failure_event). The real proxy request pipeline forks the
streaming response through several distinct asyncio Tasks (create_response's
disconnect race, the streaming generator's own task, ...); a ContextVar only
propagates into tasks forked after a value is set, so release ran in a task
that never saw admission's write. Confirmed via task-id tracing on a live
proxy that even a normal, fully-completed streaming request never released
its concurrency slot -- not just the disconnect case.
Replaces the ContextVar with a field directly on the request's own
Logging.model_call_details dict, which is explicitly passed by object
reference through both admission's request_kwargs and release's kwargs
(confirmed identical object identity on a live request), so it survives
task boundaries by construction. Deliberately not keyed by litellm_call_id
instead: that field is caller-controlled via the x-litellm-call-id header,
and an earlier design already tried and rejected that approach for exactly
this reason (letting unrelated concurrent requests merge reservations).
async_release_disconnect_state_hook now takes request_data so it can reach
the same model_call_details. Rewrote the concurrency-release tests to wire
a shared model_call_details across admission/release (mirroring production)
instead of relying on ambient task context, and re-verified live against a
real proxy: both normal completion and disconnect-before-first-chunk now
correctly free the slot.
A client disconnecting before the first streamed chunk raises CancelledError/
GeneratorExit, which bypasses both async_log_success_event and
async_log_failure_event -- the only two places tag_rate_limiter releases a
concurrency reservation queued at admission. Without a release, the
reservation sits held until the 1-hour safety-net TTL, letting a caller
exhaust their own tag's concurrency limit for free by repeatedly opening and
dropping streaming requests.
Adds an optional, default-no-op async_release_disconnect_state_hook on
CustomLogger, wires it into the proxy's shielded streaming-disconnect
cleanup (only when no disconnect-time success event already fired), and
implements it in tag_rate_limiter to release pending reservations.
configured_limit/tag_value were rebound via unpacking after already
being bound by an earlier for-loop over the same atomic_checks
sequence; only the first unpacking binding of a name is exempt from
requiring Final. Renamed to failing_limit/failing_tag_value, and used
the literal `_` placeholder (the only underscore-prefixed name this
gate's Final-exemption recognizes) for the discarded third element
instead of reusing `_key`.
The previous fix (refund through the raising index inclusive) was
itself a regression: these are shared, chain-wide buckets with no
per-request ownership tracking, so decrementing a key whose own
increment might or might not have committed is just as likely to erase
a different, legitimately-admitted concurrent request's charge on the
same key as it is to undo our own. An attacker could repeatedly cancel
requests to deliberately erase other callers' charges and exceed the
configured limit -- worse than the alternative this reverts to: a
committed-but-unrefunded key self-heals via its own TTL. Only strictly
earlier admissions in the same batch (this request's own
confirmed-successful increments, never ambiguous) are refunded now.
The earlier exception-refund fix only rolled back indices before the
one that raised, on the assumption a raise meant nothing committed for
that key. That's not guaranteed: Redis can commit the INCRBY and still
have the call raise if the response back to us is lost (a timeout, a
dropped connection), which the caller can't tell apart from a call that
never reached Redis. Now refunds through the raising index inclusive;
a clean rejection (no exception) still only refunds the earlier ones,
since TAG_RL_CHECK_AND_INCR_SCRIPT guarantees that path never committed.
Same class of issue as the earlier LIT001 fix: ruff format wrapped a
setdefault(...)-then-append pattern across lines, separating the
mutable-list construction from its justification comment. Shortened
variable names so both fit on one line stably under both ruff format
and the gate.
tag_value has no length or content bound before this hook embeds it
directly into an in-memory dict key (bypassing max_in_memory_cache_size,
which caps item count, not key bytes) and an uncapped Redis key. Hashing
to a fixed-length digest bounds this hook's own contribution to key size
regardless of the caller's input, while preserving distinctness.
Existing tests that hand-wrote the raw tag value into an expected key
string now build it through the real key-construction helper instead of
hardcoding the (now-hashed) internal format.
_atomic_check_and_increment's refund loop only ran on a normal
rejection return, not when a later key's own admission raised (a
transient Redis error) or the coroutine was cancelled mid-call.
Everything admitted earlier in that batch stayed permanently charged --
for concurrency, a leaked reservation the caller never releases,
incorrectly throttling that tag for up to the 1-hour safety TTL. A
try/finally now refunds every earlier admission before the exception
propagates, covering both raised exceptions and cancellation uniformly.
The base branch's type-discipline ceiling tightened from unrelated
merged PRs, newly flagging an annotated dict declaration ruff format
had to wrap across lines. Added a _PartitionOperations type alias so
the declaration fits on one line the gate can associate its
justification comment with.
A deployment declaring the identical concurrency_limits entry twice
appended its own id twice, inflating len(declaring_ids) past
total_deployments. That made is_chain_wide false for an entry every
deployment actually agreed on, and a non-chain-wide concurrency entry
is silently dropped entirely rather than degraded -- disabling
enforcement instead of just scoping it.