internal_usage_cache is the same DualCache instance the proxy's
key/team parallel-request limiter uses for its own authentication-
bound counters, and its default InMemoryCache evicts at 200 items.
Without isolation, a caller flooding this hook's own caller-controlled
tag buckets past that ceiling could evict an unrelated, authentication-
bound counter and let some other caller exceed a limit nothing here
configured. Give this hook a dedicated in-memory layer while still
sharing the real Redis connection when one is configured, so cross-
instance correctness is unaffected.
Router deliberately keeps a callable routing-group name (and, per its
own design, a direct deployment-id address) distinct from every member
deployment's own model_name (Router._get_routing_group_deployments).
Since _LimitsIndex only keys by model_name and team alias, a
group-addressed call previously matched neither table and the limiter
silently no-opped for both admission (async_filter_deployments) and
success accounting (async_log_success_event), even though the member
deployments carried real tag_rate_limits under their own model_name.
Add _LimitsIndex.resolve_any(), which falls back to resolving via each
candidate deployment's own model_name when the caller-visible name
matches neither table, stamping each result with the model_name it
came from (_ConfiguredLimit.resolved_group) so hashing stays
namespaced per underlying model_name -- otherwise two different
model_names sharing one routing group with an identically-named,
identically-configured limit would collide onto one Redis counter.
Direct deployment-id addressing has a related but separate, more severe
gap: Router.async_get_healthy_deployments returns a single dict (not a
list) for that path and short-circuits before Router.async_callback_
filter_deployments is ever called, so every CustomLogger.async_filter_
deployments-based hook is skipped, not just this one. That is a
Router-level structural issue affecting many hooks and is out of scope
for this PR.
async_log_success_event's kwargs is Logging.model_call_details, not the
router's flat request kwargs admission sees: metadata/litellm_metadata
are never top-level there, only nested under kwargs["litellm_params"].
Resolving the field name against kwargs itself always defaulted to
"metadata", so on LITELLM_METADATA_ROUTES (/v1/messages, /responses,
batches, bedrock, files) this read the caller's native, tag-less
metadata instead of the real, server-computed litellm_metadata.tags
admission already used -- silently skipping token/dollar accounting
for every successful call on those routes. Resolve against
kwargs["litellm_params"] instead, mirroring what admission already
does one level up.
Rebasing onto a much-advanced litellm_internal_staging tipped the
basedpyright reportPrivateUsage budget: importing _get_parent_otel_span_from_kwargs,
_PROXY_MaxParallelRequestsHandler_v3, _get_tags_from_request_kwargs, and
_PROXY_TagRateLimiter across module boundaries is the same pattern the
sibling dynamic_rate_limiter_v3 hook already uses (its own imports
just predate this budget check), so suppress with a reason rather than
renaming widely-referenced symbols.
Pre-existing mutable list[TagRateLimitEntry] field newly tripped the
type-discipline LIT001 ratchet after the rebase lowered its ceiling.
No other code relies on list mutation here (the one reader already
wraps it in tuple()), so switch to tuple[TagRateLimitEntry, ...].
- _ConfiguredLimit now carries team_scope: two teams can publish the
identical team_public_model_name alias, and the limits index already
scopes lookup by (team_id, alias) correctly -- but the Redis bucket
key itself never included team_id, so identically-named,
identically-configured limits from two different teams collided on
the same counter. Fold team_scope into the hash tag for both
admission and concurrency keys.
- _extract_key_hash and _extract_team_id no longer OR across metadata
and litellm_metadata. litellm_pre_call_utils.py writes the real,
server-authenticated value into only whichever one field is
authoritative for a given route, leaving the other exactly as the
caller sent it -- so an OR-fallback let a caller-forged
metadata.user_api_key (on a route where litellm_metadata is
authoritative) win over the real hash and bypass every
scope_by_key_hash=True limit by sending a fresh forged value per
request. Both extractors now read only the field
get_metadata_variable_name_from_kwargs names as authoritative,
matching the pattern this file already uses correctly for tags.
- Rewrite tag_rate_limiter.py's dict/list usage to immutable equivalents
(Mapping/Sequence params, tuple/frozenset/MappingProxyType returns and
locals, Final everywhere) to clear the ruff-strict type-discipline
budget (LIT001/LIT002/LIT010), keeping the pending-concurrency-keys
holder mutable by design with a documented # mutable-ok.
- Rebuild _build_limits_index's grouping via a stable sort + groupby
instead of a setdefault accumulator; caught and fixed a real bug in
that rewrite where a per-unit loop re-consumed groupby's already-
exhausted sub-iterator, silently returning empty limits for every
group after the first unit.
- Fix a basedpyright reportIncompatibleMethodOverride: match
async_filter_deployments's signature to CustomLogger's base exactly
(list/dict, not Mapping/Sequence) since it's an override.
- Fix a second reportGeneralTypeIssues: two Final-annotated locals
named `key` in sibling branches of the same function tripped
"previously declared as Final" despite being on mutually exclusive
paths; renamed them apart.
- Re-fix an eager-built f-string log message (%-style args instead)
that a prior rewrite pass had inadvertently reintroduced.
- Add the missing __init__ return annotation (ANN204) and drop the
now-unused typing.Any import in favor of a concrete object fallback
for the optional-otel-Span type alias (TID251), both newly over the
ruff-strict budget.
- Regenerate ui/litellm-dashboard/src/lib/http/schema.d.ts: the prior
comment-trimming pass left it out of sync with the trimmed
TagRateLimitEntry/TagRateLimits docstrings.
- TagRateLimitEntry.period_seconds must be a positive integer: a
configured 0 previously crashed bucket admission with a
ZeroDivisionError instead of failing config validation up front.
- Replace the pending-concurrency-keys ContextVar's immutable-tuple
rebind with a mutable holder shared by reference across every task
forked off the admitting context. asyncio.create_task only copies
which object a ContextVar is bound to, not that object's contents, so
a .set(()) performed inside a detached failure-logging task (e.g.
after a sibling pre-call-check/filter callback rejects an already-
admitted hop) was invisible to the parent task that goes on to a
fallback hop, leaving a stale key that got double-released once the
fallback's own completion event fired in the parent's context.
Release now pops an exact snapshot from the shared holder instead of
rebinding or blanket-clearing it, so a sibling hop's own
concurrently-appended reservation is never swept up either.
- Trim non-essential comments/docstrings added by this feature to
match repo convention, keeping only the ones documenting a genuinely
non-obvious invariant.
Port of the tag-based rate limiter from feature/tag-based-rate-limiting
(PR #36459), squashed to the final state of the 18 rate-limiting-specific
commits and rebased onto litellm_internal_staging.
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.
Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.
The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
An MCP server behind an API gateway needs two credentials on one request: the
gateway's own token on a private header, and a separate bearer on Authorization
for the server behind it. Every arm that minted or held a token hardcoded
Authorization, and the conflict rule then dropped the operator's static
Authorization to make room, so the second credential never arrived.
ApiKeyConfig already modelled this as header_name plus value_prefix behind a
header() method. Extend that carrier to the four minted-token configs, have each
resolver arm ask its config which header to use instead of naming one, and drop
only the header the resolved credential is about to occupy.
Operators set it per server via upstream_token_header, plumbed through
config.yaml, the credentials blob, the management API and the admin form, on the
M2M, token-exchange, authorization-code and ID-JAG arms. It is non-secret so it
stays plaintext and round-trips on admin reads. Unset keeps today's behaviour.
Moving a credential off Authorization means it stops inheriting what Authorization
gets for free, so the slot now carries those protections itself. httpx drops
Authorization when a redirect crosses origin and keeps every other header, so a
custom slot is dropped by the client on the same condition, mirroring httpx's own
scheme/host/port rule with an agreement test that fails if the two ever diverge.
The v1 path also mirrors the v2 conflict rule, so an injected header cannot shadow
the credential the gateway resolved for that slot.
Which header a credential occupies, and what counts as being that header, was
answered independently in nine places by four hand-rolled comparisons. same_header,
has_header and without_header in litellm/types/mcp.py are now the one owner, shared
by both MCP stacks, and the client derives its slot once instead of three times.
The header name reaches egress verbatim, so the RFC 7230 grammar lives in one
place and is checked where servers are built: a bad value fails the config load
and the management API returns 400, rather than raising while a spec is built
and emptying the aggregate tool list for every other server. A blank means unset,
matching what the endpoint already accepts.
The two vision tests pointed at a Wikipedia-hosted cat photo, so every run
depended on upload.wikimedia.org staying up and unthrottled. It throttled,
and the 429 surfaced as a bedrock APIConnectionError, which reads as a
gateway failure rather than what it was.
The image is now a fixture in the repo, passed as a data URL. That also puts
the two providers on the same bytes: litellm downloads the image itself for
bedrock, while openai is handed the link and fetches it from its own servers,
so the hosted URL quietly meant the two tests were not testing the same thing.
The image was generated for this repo rather than borrowed, so nothing here
carries a third-party license. Also drops a stale comment about openai prompt
caching that sat above the vision helper; no caching test uses it.
Migrate the create-key user picker, add-member user search, and usage team filter onto the shared paginated selects, gate the logs error-code filter on input reasons, and add clearAllLabel, autoHighlight, and aria-required passthroughs the migrations need.
Select the picked label on focus and snapshot whether the pre-edit selection covered the whole input; when it did, the next input value is a full replacement, so skip the typedInsertion diff that mangles pastes sharing a prefix or suffix with the label.
Under otel_v2, a client that propagates W3C trace context in params._meta
(SEP-414) pulled the tools/call span out of the gateway's trace:
resolve_mcp_span_context parented the MCP span to the client's remote
context and demoted the gateway's own transport span to a span link. The
gateway's tracing backend only ever receives the gateway's half of such a
trace, so the span was unreachable from the trace view and the POST
transaction showed a dangling link.
Invert the anchoring: the MCP tool-call and tools/list spans now always
nest under the transport span of the request carrying the message, and the
client's propagated context is recorded as the span link instead, so the
correlation survives while every trace stays renderable. With no transport
at all the span roots its own trace and still carries the link, keeping a
single shape for the event. Both returned contexts are built on an
explicitly empty base so ambient session state can never leak in, and the
span inherits the transport's sampling decision like every other
request-level span.
The /v1/messages validator checked a tool_use block's name and id but not its
input, so a block whose location came back empty or wrong still passed, while
the chat side rejected the same damage. That gap predates this branch; it is
worth closing here because the point of the change is that every parallel call
is checked rather than counted.
AnthropicContentBlock now declares input as a typed field. It already survived
on extra="allow", but reaching it from a test needs a real field to keep the
e2e basedpyright gate at zero. Serialization is unchanged: bodies are dumped
with exclude_none, so a block without an input still replays exactly as before.
* fix(ui): let the paginated search select keep what the user types
The combobox handed Base UI a freshly built option object for the current
selection every time a page of results came back. Base UI answers a changed
value by rewriting the input with that option's label, so every search response
wiped the query mid-typing and the list never narrowed. Once a user had been
picked in the Usage page filter box, no other user could be reached.
The component now owns the input text. It holds the query while the list is
open, falls back to the selected option's label once the list closes, and
remembers the picked option so its label survives later pages that no longer
carry it, the way the multi-select sibling already does.
* refactor(ui): name the paginated select's search state instead of commenting it
* fix(ui): start a fresh query when typing lands on the selected label
Focusing the filter box without clicking it leaves the caret at the end of the
selected option's label, so the next keystroke extended that label into a query
no server could match. Only a click cleared the box first.
A keystroke that arrives while the box is showing a label is now read as the
start of a new query, wherever in the label it landed.
The together backend is picked as the cheapest chat row that supports both
tools and reasoning, which currently resolves to together_ai/openai/gpt-oss-120b.
That row is marked supports_parallel_function_calling, so one weather prompt
can legitimately come back as several get_weather calls. Both tool tests
asserted exactly one call, so a parallel answer failed them even though the
gateway handled it correctly.
They now check every returned call instead of counting them: each one has to
be a get_weather naming Paris, with an id a tool result can answer. Dropping,
misnaming, or mangling a call is still red; only the count is the model's
business. The round trips answer every call rather than just the first, which
is also what the Anthropic Messages spec asks for.
The shared SelectContent wrapper defaulted alignItemWithTrigger to true,
which puts Base UI's positioner into item-aligned mode and places the
popup so the active item sits on top of the trigger. In that mode the
side and sideOffset the wrapper passes two lines above are ignored, and
the popup reports data-side="none".
The overlap only becomes visible once the items are tall enough to
matter, which is why the autorouter Template picker shows it clearly:
its options are three-line cards, so the popup covers both the select
box and its own label.
No call site in the dashboard asked for item-aligned mode. 21 of them
across 15 files already passed alignItemWithTrigger={false} by hand to
undo the default, and the remaining 127 inherited the bug. Flipping the
default makes side and sideOffset live, so collision handling works and
a select with no room below now flips above the trigger rather than
covering it. The 21 hand-written opt-outs are deleted as redundant.