model_based_tag_rate_limits_hook.py had 3 real type errors present since its
introduction, invisible in scoped per-file checks but visible once codebase-wide
upstream drift finally pushed the totals over ceiling: sorted()/groupby() keyed
by a raw dict lookup returning object rather than a provably orderable type, and
a tuple passed where async_increment_tokens_with_ttl_preservation expects a
list. Adds a small typed _model_name_of accessor for the first two, and drops
an unneeded tuple() conversion for the third since the value was already a
list. Ran make lint-budget-update per repo convention.
LiteLLM_Params and ModelInfo instead of raw dicts, plus assert narrowing on
two Router calls that can return None, to satisfy basedpyright's strict
Deployment(...) construction path.
resolve_any dedups a routing group's divergent per-deployment entries by
picking the alphabetically first member model_name sharing a signature
(resolved_group). Admission and success each independently rebuilt
candidate_model_names from the router's live routing-group membership at
their own point in time, so a deployment added or removed mid-request (a
hot-reload) could make success pick a different resolved_group than
admission did, hashing to a different Redis key and letting real token or
dollar usage escape the bucket admission actually checked. Stashes
admission's own candidate set on model_call_details, mirroring the existing
admission-time-timestamp fix, so success reuses the identical snapshot.
bugbot caught this on review.
Two Bugbot findings from the same round:
- The Redis refresh_ttl fix never reached the in-memory fallback path:
async_set_cache called through unconditionally, and InMemoryCache's
allow_ttl_override left a still-live ttl untouched regardless. Adds a
refresh_ttl kwarg to InMemoryCache.set_cache/async_set_cache that bypasses
that guard, wired through from the hook's own refresh_ttl flag.
- A team-owned deployment resolved via its team_public_model_name alias got
team_scope stamped into its bucket key, but the identical deployment
resolved via its own internal model_name (which Router.should_include_deployment
also permits for same-team or team-unconstrained callers) did not -- letting
a caller split its usage across two independent counters by alternating
which name it called with. Stamps the same team_scope onto the by_model_name
entry whenever any deployment in that group has a team alias, so both paths
resolve to the identical bucket.
bugbot flagged this as a bug (per-deployment-scoped limits should filter the
over-limit deployment out of healthy_deployments and let a sibling serve,
not reject the whole routing attempt). This is deliberate, documented design
intent from the original plan: an earlier draft considered filter-and-retry
semantics and rejected it, since rejecting the whole hop is simpler and
avoids a caller silently succeeding against a deployment whose limit
configuration they didn't intend to satisfy. Adding that reasoning as an
inline comment so it doesn't get re-flagged as a bug on a future review.
async_increment_pipeline dropped each RedisPipelineIncrementOperation's own
ttl field, so a counter created through it (Router's TPM/RPM tracking,
parallel_request_limiter_v3's token/dollar accounting when Redis is absent,
and this PR's own tag-based token/dollar limits) always fell back to the
cache's 600-second default_ttl regardless of a real, often much longer,
configured window. An hourly or daily limit's counter would silently expire
and reset mid-window. allow_ttl_override already leaves a still-live ttl
untouched on a later call, so threading the operation's ttl through on every
increment only ever takes effect the first time. bugbot caught this on review.
TAG_RL_CHECK_AND_INCR_SCRIPT only called EXPIRE when a key had no TTL at all,
so a concurrency counter's expiry was fixed from its first admission and never
pushed out by later ones. A concurrency bucket isn't epoch-windowed like
requests/tokens/dollars -- its TTL exists purely as a crash-safety net for a
reservation whose explicit release never runs -- so a still-active bucket
under sustained traffic would expire mid-flight, silently admitting past the
cap and letting a later release decrement an unrelated, newer cohort's
counter. Adds a refresh_ttl script argument, true only for the concurrency
caller, and verified against a real Redis instance since the in-memory
fallback (which already refreshes unconditionally) can't reproduce this.
bugbot caught this on review.
Every other optional hook on CustomLogger ships as an empty method a subclass
can override; this one didn't, so _release_disconnect_state_on_all_callbacks
calling it on any callback that doesn't implement it (nearly all of them)
raised AttributeError, caught and debug-logged on every single disconnect.
bugbot caught this on review.
order_tags_for_identity_resolution only checked the top level of the metadata
dict, which is correct for admission's flat request_kwargs but never present at
async_log_success_event time -- Logging.model_call_details only ever nests
metadata under kwargs["litellm_params"]. Admission correctly preferred the
key-backed identity tag, but token/dollar accounting fell through to the
caller-forged one instead, charging a different bucket than the one admission
actually checked. Adds the same litellm_params fallback _get_tags_from_request_kwargs
already relies on. bugbot caught this on review.
extract_identity/entry_applies resolve a tag_id via first-match-by-prefix over
metadata.tags, but _merge_tags keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real company_id:real-company
tag and have every rate-limit entry scoped to company_id resolve to the caller's
own value instead of the key's.
Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags (the
server-computed snapshot of only the tags the calling key/team/project's own
config contributed) ahead of the full tags list before either lookup runs, and
wires it into both call sites in model_based_tag_rate_limits_hook.py. veria-ai
caught this on review.
The disconnect-state-release hook added in the previous commit dropped the
existing has_buffered_provider_output guard and the STREAM_SSE_KEEPALIVE_PING_BYTES
exclusion while rewiring the streaming generator's cleanup path, so a client
disconnecting after only keepalive pings (or while an agentic stream holds back
real output) got refunded to input cost even when billable output had already
been generated. Restores both checks; veria-ai caught this on review, and the
existing test_streaming_cancel_after_only_keepalive_pings_reconciles_to_input_cost
regression test now passes again.
A client disconnect throws GeneratorExit/CancelledError into the
request path, so neither the success nor failure logging callback
runs and a concurrency slot reserved at admission leaks until its own
safety TTL. Gives every registered CustomLogger a chance to release
such state via the new async_release_disconnect_state_hook, called
from both the streaming and non-streaming cancel-on-disconnect paths.
Enforces token, request, dollar, and concurrency limits scoped to a
request tag (end_user_id by default), configured per deployment under
model_info.tag_rate_limits and admitted once per routing hop. Supports
chain-wide and per-deployment-scoped buckets, team-aliased routing
groups, and per-entry scoping via enabled_for/disabled_for/
apply_to_key_alias/apply_to_models.
Registers as the model_based_tag_rate_limits_hook callback and reuses
the identity extraction, policy fingerprinting, and bucket-key hashing
primitives from tag_rate_limits_shared.py.
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.
Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
extract_identity/entry_applies (this module's own functions) resolve a tag_id
via first-match-by-prefix over metadata.tags, but _merge_tags
(litellm_pre_call_utils.py) keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real
company_id:real-company tag and have every rate-limit entry scoped to
company_id resolve to the caller's own value instead of the key's.
Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags
(the server-computed snapshot of only the tags the calling key/team/project's
own config contributed) ahead of the full tags list before either lookup
runs. veria-ai caught this while reviewing #38292 (whose branch currently
carries this module's commits); porting the fix here since the vulnerable
functions it defends are this PR's own. #38292 will wire the call sites in
once it rebases onto this branch instead of carrying its own duplicate copy.
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.
Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
Keeps the dashboard's generated API types in sync with the new
TagRateLimitEntry/TagRateLimitScope/TagRateLimitGroup/TagRateLimits
schema on ModelInfo.
Both tag-scoped rate limiting hooks need the same identity/scope
extraction, policy fingerprinting, bucket-key hashing, and cache
partitioning primitives. Moving them into their own module lets a
model-independent global hook consume them without reaching into a
model-based hook's private internals, which is how the two hooks
previously shared this logic.
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
/v1/messages forwarded `thinking` verbatim for a Claude-family model and then returned,
carrying `output_config.effort` only when the model string started with a Bedrock prefix.
Every other bridged provider got a bare adaptive thinking block, so the caller's effort did
nothing: max and minimal produced byte-identical upstream bodies.
Send those targets the tier as `reasoning_effort`, which is the param they take. Bedrock keeps
taking `output_config`, since the two are not interchangeable there: an application inference
profile ARN resolves to no chat config, so `reasoning_effort` is dropped and the tier vanishes,
and a provider that rebuilds `output_config` from it overwrites a caller-set `thinking.display`
on the way. The tier stays a plain string, the summary already travelling inside the forwarded
`thinking` block. Adaptive with no tier, and budgeted thinking, both stay exactly as they were.
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.
Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.
The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
An MCP server behind an API gateway needs two credentials on one request: the
gateway's own token on a private header, and a separate bearer on Authorization
for the server behind it. Every arm that minted or held a token hardcoded
Authorization, and the conflict rule then dropped the operator's static
Authorization to make room, so the second credential never arrived.
ApiKeyConfig already modelled this as header_name plus value_prefix behind a
header() method. Extend that carrier to the four minted-token configs, have each
resolver arm ask its config which header to use instead of naming one, and drop
only the header the resolved credential is about to occupy.
Operators set it per server via upstream_token_header, plumbed through
config.yaml, the credentials blob, the management API and the admin form, on the
M2M, token-exchange, authorization-code and ID-JAG arms. It is non-secret so it
stays plaintext and round-trips on admin reads. Unset keeps today's behaviour.
Moving a credential off Authorization means it stops inheriting what Authorization
gets for free, so the slot now carries those protections itself. httpx drops
Authorization when a redirect crosses origin and keeps every other header, so a
custom slot is dropped by the client on the same condition, mirroring httpx's own
scheme/host/port rule with an agreement test that fails if the two ever diverge.
The v1 path also mirrors the v2 conflict rule, so an injected header cannot shadow
the credential the gateway resolved for that slot.
Which header a credential occupies, and what counts as being that header, was
answered independently in nine places by four hand-rolled comparisons. same_header,
has_header and without_header in litellm/types/mcp.py are now the one owner, shared
by both MCP stacks, and the client derives its slot once instead of three times.
The header name reaches egress verbatim, so the RFC 7230 grammar lives in one
place and is checked where servers are built: a bad value fails the config load
and the management API returns 400, rather than raising while a spec is built
and emptying the aggregate tool list for every other server. A blank means unset,
matching what the endpoint already accepts.
The two vision tests pointed at a Wikipedia-hosted cat photo, so every run
depended on upload.wikimedia.org staying up and unthrottled. It throttled,
and the 429 surfaced as a bedrock APIConnectionError, which reads as a
gateway failure rather than what it was.
The image is now a fixture in the repo, passed as a data URL. That also puts
the two providers on the same bytes: litellm downloads the image itself for
bedrock, while openai is handed the link and fetches it from its own servers,
so the hosted URL quietly meant the two tests were not testing the same thing.
The image was generated for this repo rather than borrowed, so nothing here
carries a third-party license. Also drops a stale comment about openai prompt
caching that sat above the vision helper; no caching test uses it.
Migrate the create-key user picker, add-member user search, and usage team filter onto the shared paginated selects, gate the logs error-code filter on input reasons, and add clearAllLabel, autoHighlight, and aria-required passthroughs the migrations need.
Select the picked label on focus and snapshot whether the pre-edit selection covered the whole input; when it did, the next input value is a full replacement, so skip the typedInsertion diff that mangles pastes sharing a prefix or suffix with the label.