Commit graph

45006 commits

Author SHA1 Message Date
Deepanshu Lulla
211ecdb912
Merge f7ea7bad37 into f677292901 2026-08-27 00:36:53 +00:00
yuneng-jiang
f677292901
Merge pull request #38392 from BerriAI/litellm_/search-tools-sync-issue-e522a2
fix(proxy): sync search tools into the router on management writes
2026-08-26 17:34:51 -07:00
Deepanshu
f7ea7bad37 fix(proxy): mirror the concurrency ttl refresh onto the in-memory fallback, dedupe team-aliased buckets by internal model_name
Two Bugbot findings from the same round:

- The Redis refresh_ttl fix never reached the in-memory fallback path:
  async_set_cache called through unconditionally, and InMemoryCache's
  allow_ttl_override left a still-live ttl untouched regardless. Adds a
  refresh_ttl kwarg to InMemoryCache.set_cache/async_set_cache that bypasses
  that guard, wired through from the hook's own refresh_ttl flag.

- A team-owned deployment resolved via its team_public_model_name alias got
  team_scope stamped into its bucket key, but the identical deployment
  resolved via its own internal model_name (which Router.should_include_deployment
  also permits for same-team or team-unconstrained callers) did not -- letting
  a caller split its usage across two independent counters by alternating
  which name it called with. Stamps the same team_scope onto the by_model_name
  entry whenever any deployment in that group has a team alias, so both paths
  resolve to the identical bucket.
2026-08-26 20:30:25 -04:00
Deepanshu
e8b6b84349 docs(proxy): document why a deployment-scoped breach rejects the whole hop
bugbot flagged this as a bug (per-deployment-scoped limits should filter the
over-limit deployment out of healthy_deployments and let a sibling serve,
not reject the whole routing attempt). This is deliberate, documented design
intent from the original plan: an earlier draft considered filter-and-retry
semantics and rejected it, since rejecting the whole hop is simpler and
avoids a caller silently succeeding against a deployment whose limit
configuration they didn't intend to satisfy. Adding that reasoning as an
inline comment so it doesn't get re-flagged as a bug on a future review.
2026-08-26 20:30:25 -04:00
Deepanshu
75efd76d37 fix(caching): apply each pipeline operation's own ttl in InMemoryCache
async_increment_pipeline dropped each RedisPipelineIncrementOperation's own
ttl field, so a counter created through it (Router's TPM/RPM tracking,
parallel_request_limiter_v3's token/dollar accounting when Redis is absent,
and this PR's own tag-based token/dollar limits) always fell back to the
cache's 600-second default_ttl regardless of a real, often much longer,
configured window. An hourly or daily limit's counter would silently expire
and reset mid-window. allow_ttl_override already leaves a still-live ttl
untouched on a later call, so threading the operation's ttl through on every
increment only ever takes effect the first time. bugbot caught this on review.
2026-08-26 20:30:25 -04:00
Deepanshu
d2b32107a1 fix(proxy): refresh a concurrency key's Redis TTL on every admission, not just its first
TAG_RL_CHECK_AND_INCR_SCRIPT only called EXPIRE when a key had no TTL at all,
so a concurrency counter's expiry was fixed from its first admission and never
pushed out by later ones. A concurrency bucket isn't epoch-windowed like
requests/tokens/dollars -- its TTL exists purely as a crash-safety net for a
reservation whose explicit release never runs -- so a still-active bucket
under sustained traffic would expire mid-flight, silently admitting past the
cap and letting a later release decrement an unrelated, newer cohort's
counter. Adds a refresh_ttl script argument, true only for the concurrency
caller, and verified against a real Redis instance since the in-memory
fallback (which already refreshes unconditionally) can't reproduce this.
bugbot caught this on review.
2026-08-26 20:30:25 -04:00
Deepanshu
b08e5aa4f4 fix(logging): add async_release_disconnect_state_hook as a default no-op on CustomLogger
Every other optional hook on CustomLogger ships as an empty method a subclass
can override; this one didn't, so _release_disconnect_state_on_all_callbacks
calling it on any callback that doesn't implement it (nearly all of them)
raised AttributeError, caught and debug-logged on every single disconnect.
bugbot caught this on review.
2026-08-26 20:30:25 -04:00
Deepanshu
8b7b648e17 fix(proxy): resolve inherited_tags from litellm_params at success-event time too
order_tags_for_identity_resolution only checked the top level of the metadata
dict, which is correct for admission's flat request_kwargs but never present at
async_log_success_event time -- Logging.model_call_details only ever nests
metadata under kwargs["litellm_params"]. Admission correctly preferred the
key-backed identity tag, but token/dollar accounting fell through to the
caller-forged one instead, charging a different bucket than the one admission
actually checked. Adds the same litellm_params fallback _get_tags_from_request_kwargs
already relies on. bugbot caught this on review.
2026-08-26 20:30:25 -04:00
Deepanshu
da3eb37342 fix(proxy): stop a caller-supplied tag from shadowing a policy-backed identity tag
extract_identity/entry_applies resolve a tag_id via first-match-by-prefix over
metadata.tags, but _merge_tags keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real company_id:real-company
tag and have every rate-limit entry scoped to company_id resolve to the caller's
own value instead of the key's.

Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags (the
server-computed snapshot of only the tags the calling key/team/project's own
config contributed) ahead of the full tags list before either lookup runs, and
wires it into both call sites in model_based_tag_rate_limits_hook.py. veria-ai
caught this on review.
2026-08-26 20:30:25 -04:00
Deepanshu
7a8ff014ee fix(proxy): restore keepalive-ping exclusion from streaming disconnect refund
The disconnect-state-release hook added in the previous commit dropped the
existing has_buffered_provider_output guard and the STREAM_SSE_KEEPALIVE_PING_BYTES
exclusion while rewiring the streaming generator's cleanup path, so a client
disconnecting after only keepalive pings (or while an agentic stream holds back
real output) got refunded to input cost even when billable output had already
been generated. Restores both checks; veria-ai caught this on review, and the
existing test_streaming_cancel_after_only_keepalive_pings_reconciles_to_input_cost
regression test now passes again.
2026-08-26 20:30:25 -04:00
Deepanshu
e2a0002850 fix(proxy): release rate limit hook state on client disconnect
A client disconnect throws GeneratorExit/CancelledError into the
request path, so neither the success nor failure logging callback
runs and a concurrency slot reserved at admission leaks until its own
safety TTL. Gives every registered CustomLogger a chance to release
such state via the new async_release_disconnect_state_hook, called
from both the streaming and non-streaming cancel-on-disconnect paths.
2026-08-26 20:30:25 -04:00
Deepanshu
8b227a8ff2 feat(rate-limiting): add per-deployment tag rate limiting hook
Enforces token, request, dollar, and concurrency limits scoped to a
request tag (end_user_id by default), configured per deployment under
model_info.tag_rate_limits and admitted once per routing hop. Supports
chain-wide and per-deployment-scoped buckets, team-aliased routing
groups, and per-entry scoping via enabled_for/disabled_for/
apply_to_key_alias/apply_to_models.

Registers as the model_based_tag_rate_limits_hook callback and reuses
the identity extraction, policy fingerprinting, and bucket-key hashing
primitives from tag_rate_limits_shared.py.
2026-08-26 20:30:25 -04:00
Deepanshu
2c9dc8baa4 fix(types): keep TagRateLimitEntry validation messages proxy-agnostic
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.

Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
2026-08-26 20:30:25 -04:00
Deepanshu
015b79155a feat(types): add TagRateLimitEntry/TagRateLimitScope config types
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
2026-08-26 20:30:25 -04:00
Deepanshu
5b005fd5c6 chore: retrigger CI (zizmor cancelled by infra/concurrency at the queue stage, not a real failure) 2026-08-26 20:30:25 -04:00
Deepanshu
a116ad37fe fix(proxy): stop a caller-supplied tag from shadowing a policy-backed identity tag
extract_identity/entry_applies (this module's own functions) resolve a tag_id
via first-match-by-prefix over metadata.tags, but _merge_tags
(litellm_pre_call_utils.py) keeps caller-supplied tags ahead of key/team/
project tags in that merged list. An authenticated caller could submit e.g.
company_id:attacker-chosen ahead of the calling key's real
company_id:real-company tag and have every rate-limit entry scoped to
company_id resolve to the caller's own value instead of the key's.

Adds order_tags_for_identity_resolution, which puts metadata.inherited_tags
(the server-computed snapshot of only the tags the calling key/team/project's
own config contributed) ahead of the full tags list before either lookup
runs. veria-ai caught this while reviewing #38292 (whose branch currently
carries this module's commits); porting the fix here since the vulnerable
functions it defends are this PR's own. #38292 will wire the call sites in
once it rebases onto this branch instead of carrying its own duplicate copy.
2026-08-26 20:30:25 -04:00
Deepanshu
7f76b53fe5 chore: retrigger CI (previous run's jobs were cancelled by infra/concurrency, not a real failure) 2026-08-26 20:30:25 -04:00
Deepanshu
eb82b3cff4 fix(types): keep TagRateLimitEntry validation messages proxy-agnostic
litellm/types/router.py is imported by plain SDK users, not just the proxy;
the previous ValueError messages for limit/key_ttl_seconds explained the
proxy rate-limit hook's internal admission mechanics (atomic
check-and-increment, read-only tokens/dollars check, cache TTL rollover),
leaking implementation details across the SDK/proxy boundary. Move that
mechanistic reasoning into code comments for future maintainers and keep
the raised messages generic, per Greptile's finding on PR #38289.

Adds regression tests asserting the three affected validators reject their
invalid inputs without leaking proxy-internal enforcement jargon.
2026-08-26 20:30:25 -04:00
Deepanshu
3e666a911e chore(ui): regenerate schema.d.ts for TagRateLimits types
Keeps the dashboard's generated API types in sync with the new
TagRateLimitEntry/TagRateLimitScope/TagRateLimitGroup/TagRateLimits
schema on ModelInfo.
2026-08-26 20:30:25 -04:00
Deepanshu
bd705a852b feat(rate-limiting): extract shared tag-rate-limit helpers into their own module
Both tag-scoped rate limiting hooks need the same identity/scope
extraction, policy fingerprinting, bucket-key hashing, and cache
partitioning primitives. Moving them into their own module lets a
model-independent global hook consume them without reaching into a
model-based hook's private internals, which is how the two hooks
previously shared this logic.
2026-08-26 20:30:25 -04:00
Deepanshu
35de9ac6af feat(types): add TagRateLimitEntry/TagRateLimitScope config types
Introduces the tag-scoped rate limit config schema (TagRateLimitEntry,
TagRateLimitScope, TagRateLimitGroup, TagRateLimits) and wires it onto
ModelInfo.tag_rate_limits, giving tag-based rate limiting hooks a
config shape to validate and consume.
2026-08-26 20:30:25 -04:00
Mateo Wang
d8595cb647
Merge pull request #38423 from BerriAI/litellm_gemini_latest_cache_read_rates
fix(model_prices): bill gemini -latest/preview alias cache reads at 10% of input
2026-08-26 17:25:24 -07:00
Mateo Wang
53a607e088
Merge pull request #38411 from BerriAI/litellm_fix_prompt_patch_sync
fix(prompts): propagate PATCHed prompt templates to every worker and pod
2026-08-26 17:20:27 -07:00
Mateo Wang
1ac39b10ba
Merge pull request #38412 from BerriAI/litellm_fix_gemini_tts_native_audio_rates
fix(cost-map): correct Gemini TTS and native-audio rates
2026-08-26 17:15:46 -07:00
Mateo Wang
b54f7505a3
Merge pull request #38419 from BerriAI/litellm_gemini_live_realtime_cost
fix(cost): price gemini-live-2.5-flash-native-audio realtime sessions
2026-08-26 17:15:41 -07:00
Mateo Wang
4e295e8eb9
Merge pull request #38422 from BerriAI/litellm_gemini35_flashlite_flex_cache_price
fix(model_prices): correct gemini-3.5-flash-lite flex cache-read pricing
2026-08-26 17:11:07 -07:00
Mateo Wang
39dd46397e
Merge pull request #38379 from BerriAI/litellm_mcp_oauth_admin_entered_authorize_urls
fix(mcp): honor admin-entered OAuth URLs on authorize after issuer yield
2026-08-26 17:11:00 -07:00
yuneng-jiang
e80ba92cfa
Merge pull request #38313 from BerriAI/litellm_/hide-unhealthy-virtual-key-models-922c37
feat(proxy): hide unhealthy models from model listings, opt-in
2026-08-26 17:05:41 -07:00
mateo-berri
901bea41e7 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_gemini_latest_cache_read_rates
# Conflicts:
#	tests/test_litellm/llms/gemini/test_cost_calculator.py
2026-08-26 17:04:59 -07:00
yuneng-jiang
cbefebbd9f
Merge branch 'litellm_internal_staging' into litellm_/search-tools-sync-issue-e522a2 2026-08-26 17:04:01 -07:00
mateo-berri
5461bb3b48 fix(prompts): sync only the newest row when environments share a versioned prompt id 2026-08-26 17:00:24 -07:00
devin-ai-integration[bot]
8a9d5b15b4
feat(langfuse): support langfuse_environment as a per-key dynamic callback param (#38264)
* feat(langfuse): support langfuse_environment as a per-key dynamic callback param

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(langfuse): type the langfuse_environment constructor param

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(langfuse): only pass environment when the SDK client supports it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langfuse): drop the request-body metadata test for langfuse_environment

The proxy bans request-body callback params by default (derived from
_supported_callback_params in auth_utils), so the metadata channel this
test asserted is rejected with a 401 on the proxy. The supported channel
is admin-set key/team callback_vars, with LANGFUSE_TRACING_ENVIRONMENT
as the deployment-wide fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(langfuse): validate langfuse_environment, avoid redundant clients, honor it in langfuse_otel

Closes the review gaps on the langfuse_environment param:

- Validate values against Langfuse's environment pattern at save time
  (/key/generate, /key/update, /team callback all 400 on e.g. 'Production'
  instead of 200-then-silently-dropping every trace server-side) and at
  logger init; non-string values are str()-coerced instead of crashing
  the SDK's regex check per event.
- Treat empty/whitespace values and values equal to the deployment-wide
  LANGFUSE_TRACING_ENVIRONMENT as non-dynamic so an environment-only
  override that changes nothing no longer mints a duplicate SDK client
  against MAX_LANGFUSE_INITIALIZED_CLIENTS.
- langfuse_otel now reads the per-key/team langfuse_environment from
  standard_callback_dynamic_params instead of only the env var.
- Advertise the param on the discovery surfaces: callback_configs.json
  (langfuse + langfuse_otel), the dashboard callback registry, and the
  /team/{team_id}/callback docstring (schema.d.ts regenerated).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: ruff format langfuse files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lint): remove duplicate test import, LIT002 dict literal, and mock-echo otel test

- drop redundant in-function import of callback_config_error (F811)
- avoid the `or {}` mutable literal in _set_langfuse_specific_attributes (LIT002)
- rewrite the dynamic-env otel test to observe span.set_attribute output
  instead of patching litellm internals (TQ002/TQ008)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng-berri <yucheng@berri.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:56:55 -07:00
mateo-berri
1e1c231076 fix(model_prices): bill gemini -latest/preview alias cache reads at 10% of input 2026-08-26 16:52:09 -07:00
Mateo Wang
40005cf7f8
Merge pull request #37724 from bisma-nawaz/fix-37647-staging
fix: map Gemini ON_DEMAND_FLEX traffic type to flex service tier
2026-08-26 16:51:56 -07:00
Mateo Wang
d77eef3d11
Merge pull request #38414 from BerriAI/litellm_fix_speech_metadata_spend_tracking
fix(speech): keep proxy metadata and completion cost through the TTS completion bridge
2026-08-26 16:51:50 -07:00
mateo-berri
bd75c38e84 fix(model_prices): scope flash-lite flex cache-read cut to vertex entries 2026-08-26 16:48:32 -07:00
Mateo Wang
ad00d90b99
Merge pull request #38418 from BerriAI/litellm_gemini_maps_grounding_cost
fix(gemini): bill Google Maps grounding as its own SKU
2026-08-26 16:47:05 -07:00
tin-berri
cebf0d6f21
fix(responses): let cache-control injection reach the system prompt from instructions (#38120)
`AnthropicCacheControlHook` spends the configured injection points on the first
message list it is shown and drops the message points that matched nothing. That is
right when the messages it sees are the ones going upstream. It is wrong for
/v1/responses: the system prompt lives in `instructions`, which only becomes a system
message once the chat-completion bridge builds one, so a role-targeted point matched
nothing and was thrown away before the message it wanted existed. Injection silently
did nothing across the whole surface.

Hand those points back instead, stamped as judged, when the caller says its message
list is provisional. The stamp is what makes carrying them safe: without it the next
pass re-judges the points against messages this pass has already marked and stands the
whole configuration down. Callers holding the final messages -- /chat/completions and
/v1/messages -- do not raise the signal and keep dropping unmatched points as before.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 16:43:26 -07:00
tin-berri
d8edfb69c2
fix(proxy): derive auto-router health from its underlying models (#38174)
An auto_router deployment is a marker, not something a probe can contact, so
`_run_model_health_check` returns `{}` for it and it lands healthy whatever is
behind it. This derives its verdict from the models it actually resolves.

Rules and owners:

- `strategy_router_dependencies` is the single answer to "what does this router
  call": tier, default, classifier and embedding names per router kind, aligned
  with what init and the request path actually use.
- `_health_check_eligible` is the single probe-eligibility gate, applied to the
  requested set and to the pool a router's dependencies are drawn from alike, so
  an opted-out deployment cannot re-enter through a router that depends on it.
- `_resolved_deployment_ids` resolves names through `get_model_list`, the same
  composition of alias, routing-group and wildcard channels a request uses.
- A dependency reds its router only when *every* deployment behind the name is
  known unhealthy. A replica this run never judged, hidden from the caller or
  opted out of health checks, can still serve what the dead one drops, so
  partial evidence leaves the verdict green. Absent information never reds.
- Verdicts settle over rounds, because a marker never fails a probe of its own
  and a parent whose tier is a red router must inherit that fault. Both sweeps
  are bounded loops, so a router cycle terminates green.
- Dependency probes are added only on the targeted `/health?model_id=` path the
  dashboard uses per deployment, and are dropped from the response.

Resolves LIT-6073

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 16:41:54 -07:00
mateo-berri
885d95d71b Merge remote-tracking branch 'origin/litellm_internal_staging' into pr37724
# Conflicts:
#	tests/test_litellm/test_cost_calculator.py
2026-08-26 16:33:28 -07:00
mateo-berri
e7b843d69b test: trim realtime cost test docstrings to one line 2026-08-26 16:33:02 -07:00
Mateo Wang
b98b2d562b
Merge pull request #38416 from BerriAI/litellm_fix_lazy_openapi_stubs_for_imported_modules
fix(proxy): key lazy openapi stubs off registered features, not sys.modules
2026-08-26 16:32:02 -07:00
mateo-berri
f758e9c30a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_speech_metadata_spend_tracking 2026-08-26 16:30:30 -07:00
mateo-berri
ac3f987883 fix: thread service_tier through vertex cost_per_character fallbacks
Vertex Gemini 3.x models route through cost_per_character (the cost_router
token-path gate only matches gemini-2), and its token fallbacks dropped
service_tier, so ON_DEMAND_FLEX responses were still billed at the standard
rate. Pass the tier through the call site and all four fallbacks.
2026-08-26 16:29:45 -07:00
mateo-berri
6c07fd547b test: accept the Maps grounding rate in the intended cost map schema 2026-08-26 16:28:00 -07:00
mateo-berri
dca5144dba Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_gemini_maps_grounding_cost
# Conflicts:
#	litellm/llms/vertex_ai/gemini/vertex_and_google_ai_studio_gemini.py
2026-08-26 16:27:40 -07:00
mateo-berri
b48bff7b54 fix(cost_calculator): require real values when detecting declared realtime pricing 2026-08-26 16:27:28 -07:00
mateo-berri
93e7e8d980 fix(mcp): token exchange rejoins discovery for a clientless DCR bridge missing its registration endpoint 2026-08-26 16:26:44 -07:00
mateo-berri
0243c5dee4 fix(model_prices): correct gemini-3.5-flash-lite flex cache-read pricing 2026-08-26 16:25:15 -07:00
mateo-berri
fdeab570a1 fix(speech): forward api_key to the TTS bridge and isolate response hidden params 2026-08-26 16:23:33 -07:00