Commit graph

45193 commits

Author SHA1 Message Date
Deepanshu
43654422f5 style(rate-limiting): shrink a LIT001 annotation the base's ratcheted-down budget now flags
The base branch's type-discipline ceiling tightened from unrelated
merged PRs, newly flagging an annotated dict declaration ruff format
had to wrap across lines. Added a _PartitionOperations type alias so
the declaration fits on one line the gate can associate its
justification comment with.
2026-08-27 18:39:50 -04:00
Deepanshu
3b0f8287f4 fix(rate-limiting): dedupe a deployment's own repeated entry before counting declarations
A deployment declaring the identical concurrency_limits entry twice
appended its own id twice, inflating len(declaring_ids) past
total_deployments. That made is_chain_wide false for an entry every
deployment actually agreed on, and a non-chain-wide concurrency entry
is silently dropped entirely rather than degraded -- disabling
enforcement instead of just scoping it.
2026-08-27 18:39:50 -04:00
Deepanshu
93590434f4 fix(rate-limiting): remove an unnecessary forward-reference quote (ruff UP037)
self.keys: T = value annotations inside a method body are never
evaluated at runtime, so the "_PartitionKey" forward-reference quoting
was unneeded despite _PartitionKey being defined later in the file.
2026-08-27 18:39:50 -04:00
Deepanshu
b4f6e49253 style(rate-limiting): satisfy ruff format without losing LIT001 annotations
A prior `ruff format` run wrapped several annotated mutable-dict/list
declarations across multiple lines, moving their `# mutable-ok`
justification off the line the type-discipline gate checks. Shortened
the annotations (a new _DedupSignature type alias, trimmed comments) so
they fit on one line under both ruff format and the gate.
2026-08-27 18:39:50 -04:00
Deepanshu
79e66d9754 fix(rate-limiting): stop leaving a permanent Redis key on an expired-reservation release
Releasing a reservation whose key had already expired made INCRBY
recreate it, and the floor-to-zero SET left that recreated key
permanently in Redis (SET clears any TTL). Verified against a real
Redis instance: DEL removes the key outright instead, which reads back
identically to 0 everywhere this key is read.
2026-08-27 18:39:50 -04:00
Deepanshu
f6f65e181b fix(rate-limiting): reject key_ttl_seconds shorter than period_seconds
A key_ttl_seconds override below period_seconds expired the bucket key
before its window rolled over, resetting the counter to zero mid-window
and letting tagged traffic exceed the configured limit. Only
period_seconds and above is now accepted.
2026-08-27 18:39:49 -04:00
Deepanshu
e83886a9d3 test(rate-limiting): verify scope_by_key_hash composes with the new per-entry overrides
Confirms _partition_key treats scope_by_key_hash as distinguishing (two
entries identical otherwise but differing only on this flag must not
share a partition), and that per-key bucket independence still holds
when key_ttl_seconds and max_in_memory_cache_size are also set, going
through the real _build_limits_index path rather than a hand-built
_ConfiguredLimit.
2026-08-27 18:39:49 -04:00
Deepanshu
9e07e685f3 feat(rate-limiting): make the in-memory cache size configurable per tag entry
Adds TagRateLimitEntry.max_in_memory_cache_size so a single
high-cardinality entry can get its own dedicated in-memory cache
partition instead of sharing the hook's single default one, keeping the
knob alongside key_ttl_seconds and the rest of that entry's config
rather than only as a proxy-wide setting. Partitions are keyed by each
entry's full signature, not the override value alone, so two unrelated
entries that happen to pick the same size don't get merged.

tokens/dollars accounting and concurrency-slot release are both
partition-aware too: each partition owns its own v3 handler, and a
concurrency reservation is released against the exact partition it was
incremented on so it can't leak onto the default partition instead.

Also fixes a bug this surfaced: _configured_limit_for_signature
rebuilt each entry from a 5-field dedup signature that excluded
key_ttl_seconds and max_in_memory_cache_size, silently resetting both
to their defaults for every entry reachable through the real indexing
path used by async_filter_deployments and async_log_success_event.
2026-08-27 18:39:49 -04:00
Deepanshu
a2a19cbdd9 fix(rate-limiting): reject invalid in-memory cache sizes, add a per-tag Redis TTL override
An unresolved os.environ/ substitution or a config typo could set
tag_rate_limiter_max_in_memory_cache_size to a negative number or a
string; InMemoryCache raises comparing its size against that value, and
DualCache.async_set_cache swallows the exception, silently disabling
every counter write for this hook without Redis. Only positive integers
are now accepted; anything else falls back to the safe default with a
warning.

Also adds TagRateLimitEntry.key_ttl_seconds so a high-cardinality tag_id
can shed its Redis (or in-memory fallback) keys sooner without
shortening period_seconds itself. Concurrency's safety-floor TTL is
still never lowered by this override.
2026-08-27 18:39:49 -04:00
Deepanshu
fe9e36a0bc feat(rate-limiting): make the tag rate limiter's in-memory cache size configurable
The isolated in-memory cache added for this hook still defaults to 200
entries. A deployment rate-limiting on a high-cardinality tag_id without
Redis can churn past that cap, evicting an active counter before its
period elapses. litellm_settings.tag_rate_limiter_max_in_memory_cache_size
lets that ceiling be raised; 0 is rejected in favor of the safe default
since it would disable the in-memory cache outright.
2026-08-27 18:39:49 -04:00
Deepanshu
e6c504645b fix(rate-limiting): isolate this hook's in-memory cache from the shared proxy cache
internal_usage_cache is the same DualCache instance the proxy's
key/team parallel-request limiter uses for its own authentication-
bound counters, and its default InMemoryCache evicts at 200 items.
Without isolation, a caller flooding this hook's own caller-controlled
tag buckets past that ceiling could evict an unrelated, authentication-
bound counter and let some other caller exceed a limit nothing here
configured. Give this hook a dedicated in-memory layer while still
sharing the real Redis connection when one is configured, so cross-
instance correctness is unaffected.
2026-08-27 18:39:49 -04:00
Deepanshu
c62d79f7df fix(rate-limiting): resolve limits by deployment model_name for routing-group calls
Router deliberately keeps a callable routing-group name (and, per its
own design, a direct deployment-id address) distinct from every member
deployment's own model_name (Router._get_routing_group_deployments).
Since _LimitsIndex only keys by model_name and team alias, a
group-addressed call previously matched neither table and the limiter
silently no-opped for both admission (async_filter_deployments) and
success accounting (async_log_success_event), even though the member
deployments carried real tag_rate_limits under their own model_name.

Add _LimitsIndex.resolve_any(), which falls back to resolving via each
candidate deployment's own model_name when the caller-visible name
matches neither table, stamping each result with the model_name it
came from (_ConfiguredLimit.resolved_group) so hashing stays
namespaced per underlying model_name -- otherwise two different
model_names sharing one routing group with an identically-named,
identically-configured limit would collide onto one Redis counter.

Direct deployment-id addressing has a related but separate, more severe
gap: Router.async_get_healthy_deployments returns a single dict (not a
list) for that path and short-circuits before Router.async_callback_
filter_deployments is ever called, so every CustomLogger.async_filter_
deployments-based hook is skipped, not just this one. That is a
Router-level structural issue affecting many hooks and is out of scope
for this PR.
2026-08-27 18:39:49 -04:00
Deepanshu
8d4528ba37 fix(rate-limiting): resolve the authoritative metadata bucket at success accounting too
async_log_success_event's kwargs is Logging.model_call_details, not the
router's flat request kwargs admission sees: metadata/litellm_metadata
are never top-level there, only nested under kwargs["litellm_params"].
Resolving the field name against kwargs itself always defaulted to
"metadata", so on LITELLM_METADATA_ROUTES (/v1/messages, /responses,
batches, bedrock, files) this read the caller's native, tag-less
metadata instead of the real, server-computed litellm_metadata.tags
admission already used -- silently skipping token/dollar accounting
for every successful call on those routes. Resolve against
kwargs["litellm_params"] instead, mirroring what admission already
does one level up.
2026-08-27 18:39:49 -04:00
Deepanshu
cdcaac75cd fix(rate-limiting): suppress reportPrivateUsage for established cross-module private-helper reuse
Rebasing onto a much-advanced litellm_internal_staging tipped the
basedpyright reportPrivateUsage budget: importing _get_parent_otel_span_from_kwargs,
_PROXY_MaxParallelRequestsHandler_v3, _get_tags_from_request_kwargs, and
_PROXY_TagRateLimiter across module boundaries is the same pattern the
sibling dynamic_rate_limiter_v3 hook already uses (its own imports
just predate this budget check), so suppress with a reason rather than
renaming widely-referenced symbols.
2026-08-27 18:39:49 -04:00
Deepanshu
552b016128 chore(ui): regenerate schema.d.ts for TagRateLimitGroup.limits tuple change 2026-08-27 18:39:49 -04:00
Deepanshu
50a11743d8 fix(rate-limiting): use an immutable tuple for TagRateLimitGroup.limits
Pre-existing mutable list[TagRateLimitEntry] field newly tripped the
type-discipline LIT001 ratchet after the rebase lowered its ceiling.
No other code relies on list mutation here (the one reader already
wraps it in tuple()), so switch to tuple[TagRateLimitEntry, ...].
2026-08-27 18:39:49 -04:00
Deepanshu
170aa0604b fix(rate-limiting): scope alias buckets by team, stop caller-forged metadata bypassing key-scoped limits
- _ConfiguredLimit now carries team_scope: two teams can publish the
  identical team_public_model_name alias, and the limits index already
  scopes lookup by (team_id, alias) correctly -- but the Redis bucket
  key itself never included team_id, so identically-named,
  identically-configured limits from two different teams collided on
  the same counter. Fold team_scope into the hash tag for both
  admission and concurrency keys.
- _extract_key_hash and _extract_team_id no longer OR across metadata
  and litellm_metadata. litellm_pre_call_utils.py writes the real,
  server-authenticated value into only whichever one field is
  authoritative for a given route, leaving the other exactly as the
  caller sent it -- so an OR-fallback let a caller-forged
  metadata.user_api_key (on a route where litellm_metadata is
  authoritative) win over the real hash and bypass every
  scope_by_key_hash=True limit by sending a fresh forged value per
  request. Both extractors now read only the field
  get_metadata_variable_name_from_kwargs names as authoritative,
  matching the pattern this file already uses correctly for tags.
2026-08-27 18:39:49 -04:00
Deepanshu
48b71a8a32 fix(rate-limiting): satisfy type-discipline, basedpyright, and eager-logging gates
- Rewrite tag_rate_limiter.py's dict/list usage to immutable equivalents
  (Mapping/Sequence params, tuple/frozenset/MappingProxyType returns and
  locals, Final everywhere) to clear the ruff-strict type-discipline
  budget (LIT001/LIT002/LIT010), keeping the pending-concurrency-keys
  holder mutable by design with a documented # mutable-ok.
- Rebuild _build_limits_index's grouping via a stable sort + groupby
  instead of a setdefault accumulator; caught and fixed a real bug in
  that rewrite where a per-unit loop re-consumed groupby's already-
  exhausted sub-iterator, silently returning empty limits for every
  group after the first unit.
- Fix a basedpyright reportIncompatibleMethodOverride: match
  async_filter_deployments's signature to CustomLogger's base exactly
  (list/dict, not Mapping/Sequence) since it's an override.
- Fix a second reportGeneralTypeIssues: two Final-annotated locals
  named `key` in sibling branches of the same function tripped
  "previously declared as Final" despite being on mutually exclusive
  paths; renamed them apart.
- Re-fix an eager-built f-string log message (%-style args instead)
  that a prior rewrite pass had inadvertently reintroduced.
2026-08-27 18:39:49 -04:00
Deepanshu
0004fe1b92 fix(rate-limiting): satisfy strict lint gate and regenerate stale schema.d.ts
- Add the missing __init__ return annotation (ANN204) and drop the
  now-unused typing.Any import in favor of a concrete object fallback
  for the optional-otel-Span type alias (TID251), both newly over the
  ruff-strict budget.
- Regenerate ui/litellm-dashboard/src/lib/http/schema.d.ts: the prior
  comment-trimming pass left it out of sync with the trimmed
  TagRateLimitEntry/TagRateLimits docstrings.
2026-08-27 18:39:49 -04:00
Deepanshu
11461f6e45 fix(rate-limiting): reject non-positive period_seconds, fix cross-task concurrency slot leak
- TagRateLimitEntry.period_seconds must be a positive integer: a
  configured 0 previously crashed bucket admission with a
  ZeroDivisionError instead of failing config validation up front.
- Replace the pending-concurrency-keys ContextVar's immutable-tuple
  rebind with a mutable holder shared by reference across every task
  forked off the admitting context. asyncio.create_task only copies
  which object a ContextVar is bound to, not that object's contents, so
  a .set(()) performed inside a detached failure-logging task (e.g.
  after a sibling pre-call-check/filter callback rejects an already-
  admitted hop) was invisible to the parent task that goes on to a
  fallback hop, leaving a stale key that got double-released once the
  fallback's own completion event fired in the parent's context.
  Release now pops an exact snapshot from the shared holder instead of
  rebinding or blanket-clearing it, so a sibling hop's own
  concurrently-appended reservation is never swept up either.
- Trim non-essential comments/docstrings added by this feature to
  match repo convention, keeping only the ones documenting a genuinely
  non-obvious invariant.
2026-08-27 18:39:49 -04:00
Deepanshu
fdf41b49c4 feat(rate-limiting): add tag-scoped token/request/dollar/concurrency rate limits
Port of the tag-based rate limiter from feature/tag-based-rate-limiting
(PR #36459), squashed to the final state of the 18 rate-limiting-specific
commits and rebased onto litellm_internal_staging.
2026-08-27 18:39:49 -04:00
tin-berri
30ff3723b2
feat(model_prices): let a map entry declare its exact reasoning_effort levels (#38481)
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.

Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.

The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
2026-08-27 15:38:01 -07:00
ryan-crabbe-berri
3c41392893
Merge pull request #38574 from BerriAI/litellm_combobox_server_search_hardening
fix(ui): stop server-searched comboboxes from clobbering picks and queries
2026-08-27 15:33:12 -07:00
Mateo Wang
a6816f0e96
Merge pull request #38486 from BerriAI/litellm_together_glm53_flash
feat(together_ai): add zai-org/GLM-5.3-Flash to the model registry
2026-08-27 15:24:30 -07:00
yuneng-jiang
9d1348c1a4
Merge pull request #38575 from BerriAI/litellm_e2e-vision-own-image
test(e2e): serve the vision image from our own fixture
2026-08-27 15:24:09 -07:00
Mateo Wang
4ef1c28877
Merge pull request #38431 from BerriAI/litellm_fix_messages_native_tools
fix(anthropic-adapter): pass provider-native and OpenAI-format tools through on /v1/messages
2026-08-27 15:24:07 -07:00
Mateo Wang
649dc23d6a
Merge pull request #38465 from BerriAI/litellm_lit6103_tool_reference_passthrough
fix(anthropic): carry tool_reference tool results through the guardrail translation round trip
2026-08-27 15:23:51 -07:00
Mateo Wang
0c0a1f4eb1
Merge pull request #38457 from BerriAI/litellm_realtime_audio_output_tokens
fix(realtime): bill Gemini Live native-audio output tokens at the audio rate
2026-08-27 15:23:46 -07:00
Mateo Wang
66ea1bbbe8
Merge pull request #38458 from BerriAI/litellm_streaming_flex_service_tier
fix(streaming): preserve provider service-tier metadata so Vertex flex streams bill at flex rates
2026-08-27 15:23:43 -07:00
Mateo Wang
e5b2f5bec2
Merge pull request #38561 from BerriAI/litellm_transcription_srt_vtt_synthesis
fix(transcription): synthesize srt/vtt output for adapters without native subtitle formats
2026-08-27 15:23:38 -07:00
Mateo Wang
55c1537497
Merge pull request #38376 from BerriAI/devin_ai_bedrock_guardrail_external_id
fix(guardrails): forward aws_external_id when the bedrock guardrail assumes a role
2026-08-27 15:22:11 -07:00
Mateo Wang
dc1b847c4f
Merge pull request #38280 from BerriAI/litellm_together_cache_pricing
fix(cost): apply Together AI cache read pricing and per-model registry rates
2026-08-27 15:20:09 -07:00
mateo-berri
1665214bbd feat(together_ai): flag prompt caching on GLM-5.3-Flash like its sibling entries 2026-08-27 15:13:42 -07:00
mateo-berri
ae89f9cf74 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_together_glm53_flash 2026-08-27 15:12:35 -07:00
Mateo Wang
a2c814654e
Merge pull request #38449 from BerriAI/litellm_dashscope_qwen_image_3
feat(dashscope): support qwen-image-3.0 and qwen-image-3.0-pro image generation
2026-08-27 15:08:39 -07:00
Yuneng Jiang
49170695ce
test(e2e): drop the fixture helper docstring
The why belongs in the commit message and the PR, not above a one-line
helper whose name already says what it returns.
2026-08-27 14:50:17 -07:00
Mateo Wang
7083c47998
Merge pull request #38263 from BerriAI/litellm_together_reasoning_effort
feat(together_ai): map reasoning_effort per model class
2026-08-27 14:42:07 -07:00
mateo-berri
9e9c7e621f Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_together_reasoning_effort
# Conflicts:
#	litellm/llms/together_ai/chat/transformation.py
#	tests/test_litellm/llms/together_ai/chat/test_together_ai_chat_transformation.py
2026-08-27 14:33:47 -07:00
tin-berri
40ff01b987
feat(mcp): let a resolved OAuth token target a custom upstream header (#38456)
An MCP server behind an API gateway needs two credentials on one request: the
gateway's own token on a private header, and a separate bearer on Authorization
for the server behind it. Every arm that minted or held a token hardcoded
Authorization, and the conflict rule then dropped the operator's static
Authorization to make room, so the second credential never arrived.

ApiKeyConfig already modelled this as header_name plus value_prefix behind a
header() method. Extend that carrier to the four minted-token configs, have each
resolver arm ask its config which header to use instead of naming one, and drop
only the header the resolved credential is about to occupy.

Operators set it per server via upstream_token_header, plumbed through
config.yaml, the credentials blob, the management API and the admin form, on the
M2M, token-exchange, authorization-code and ID-JAG arms. It is non-secret so it
stays plaintext and round-trips on admin reads. Unset keeps today's behaviour.

Moving a credential off Authorization means it stops inheriting what Authorization
gets for free, so the slot now carries those protections itself. httpx drops
Authorization when a redirect crosses origin and keeps every other header, so a
custom slot is dropped by the client on the same condition, mirroring httpx's own
scheme/host/port rule with an agreement test that fails if the two ever diverge.
The v1 path also mirrors the v2 conflict rule, so an injected header cannot shadow
the credential the gateway resolved for that slot.

Which header a credential occupies, and what counts as being that header, was
answered independently in nine places by four hand-rolled comparisons. same_header,
has_header and without_header in litellm/types/mcp.py are now the one owner, shared
by both MCP stacks, and the client derives its slot once instead of three times.

The header name reaches egress verbatim, so the RFC 7230 grammar lives in one
place and is checked where servers are built: a bad value fails the config load
and the management API returns 400, rather than raising while a spec is built
and emptying the aggregate tool list for every other server. A blank means unset,
matching what the endpoint already accepts.
2026-08-27 14:32:01 -07:00
Yuneng Jiang
ff418ffb9c
test(e2e): serve the vision image from our own fixture
The two vision tests pointed at a Wikipedia-hosted cat photo, so every run
depended on upload.wikimedia.org staying up and unthrottled. It throttled,
and the 429 surfaced as a bedrock APIConnectionError, which reads as a
gateway failure rather than what it was.

The image is now a fixture in the repo, passed as a data URL. That also puts
the two providers on the same bytes: litellm downloads the image itself for
bedrock, while openai is handed the link and fetches it from its own servers,
so the hosted URL quietly meant the two tests were not testing the same thing.

The image was generated for this repo rather than borrowed, so nothing here
carries a third-party license. Also drops a stale comment about openai prompt
caching that sat above the vision helper; no caching test uses it.
2026-08-27 14:29:53 -07:00
mateo-berri
dbadee7210 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_together_cache_pricing
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
#	tests/test_litellm/test_cost_calculator.py
2026-08-27 14:29:46 -07:00
mateo-berri
10480b8bf0 refactor(dashscope): drop redundant routing comment 2026-08-27 14:27:41 -07:00
mateo-berri
c860d511db Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_realtime_audio_output_tokens 2026-08-27 14:25:08 -07:00
yuneng-jiang
852cb3abbe
Merge pull request #38567 from BerriAI/litellm_together-parallel-tool-calls
test(e2e): let the together tool tests accept parallel calls
2026-08-27 14:24:18 -07:00
Mateo Wang
5cab47ee3e
Merge pull request #38487 from BerriAI/litellm_fix_together_ai_fail_open_test
test(together_ai): assert fail-open supported params for models missing from the registry
2026-08-27 14:23:19 -07:00
ryan-crabbe-berri
141c91404f fix(ui): reason-gate the remaining server-searched comboboxes
Migrate the create-key user picker, add-member user search, and usage team filter onto the shared paginated selects, gate the logs error-code filter on input reasons, and add clearAllLabel, autoHighlight, and aria-required passthroughs the migrations need.
2026-08-27 14:22:10 -07:00
ryan-crabbe-berri
bf86eadd83 fix(ui): take whole-selection edits verbatim in the paginated search select
Select the picked label on focus and snapshot whether the pre-edit selection covered the whole input; when it did, the next input value is a full replacement, so skip the typedInsertion diff that mangles pastes sharing a prefix or suffix with the label.
2026-08-27 14:07:02 -07:00
Mateo Wang
67c7b97fd2
Merge pull request #38207 from BerriAI/litellm_registry_audit_bedrock_sol_anthropic_1hr
fix(model_prices): rolling registry audit - verified models and rates for Novita, DeepInfra, W&B, Bedrock Sol, Gemini, Fireworks, Azure gpt-5.6, Mistral, Together
2026-08-27 13:42:51 -07:00
Mateo Wang
0441faadca
Merge pull request #37833 from BerriAI/litellm_deflake_20260821
fix: roll up the open deflake fixes for the MCP logging queue, PTU rollup, license gate, and pricing test isolation
2026-08-27 13:42:26 -07:00
yucheng-berri
88cb83484b
fix(otel): anchor MCP tool-call spans to the gateway's own trace, link the client's context (#38317)
Under otel_v2, a client that propagates W3C trace context in params._meta
(SEP-414) pulled the tools/call span out of the gateway's trace:
resolve_mcp_span_context parented the MCP span to the client's remote
context and demoted the gateway's own transport span to a span link. The
gateway's tracing backend only ever receives the gateway's half of such a
trace, so the span was unreachable from the trace view and the POST
transaction showed a dangling link.

Invert the anchoring: the MCP tool-call and tools/list spans now always
nest under the transport span of the request carrying the message, and the
client's propagated context is recorded as the span link instead, so the
correlation survives while every trace stays renderable. With no transport
at all the span roots its own trace and still carries the link, keeping a
single shape for the event. Both returned contexts are built on an
explicitly empty base so ambient session state can never leak in, and the
span inherits the transport's sampling decision like every other
request-level span.
2026-08-27 13:41:30 -07:00