Commit graph

8121 commits

Author SHA1 Message Date
mateo-berri
688575e5bf fix(cost): bill reasoning tokens at the selected tier's reasoning rate 2026-08-14 17:26:29 -07:00
Mateo Wang
a5038661c9
Merge pull request #34850 from BerriAI/litellm_lit_4866_anthropic_geo_cache_uplift
fix(anthropic cost): apply regional geo uplift to cached tokens
2026-08-14 17:26:17 -07:00
Mateo Wang
b9d70b5ef5
Merge pull request #34860 from BerriAI/litellm_lit_4868_cache_write_split
fix(anthropic): aggregate 5m/1h cache-write split across iterations path
2026-08-14 17:25:20 -07:00
mateo-berri
7562445273 test(proxy): assert production nesting semantics for component cost headers 2026-08-14 17:24:10 -07:00
Mateo Wang
d74cb6de1b
Merge pull request #36798 from BerriAI/litellm_azure_ai_docs_index_write_grant_rc
fix(azure_ai): recognize real Search doc endpoints so teams can read/write via passthrough
2026-08-14 17:24:06 -07:00
mateo-berri
3217b8edae fix(proxy): count worker heartbeats on the primary so replica lag cannot undercount 2026-08-14 17:23:06 -07:00
mateo-berri
eafddaaa12 feat(scripts): queue heavy gates behind a machine-wide slot lock 2026-08-14 17:22:32 -07:00
mateo-berri
c9685b2a6d fix(router): honor tiered_pricing set in a deployment's litellm_params 2026-08-14 17:21:15 -07:00
yucheng-berri
f9f5c03884
fix(mcp): drop caller host and configured upstream headers from logged metadata (#36901)
* fix(mcp): drop caller host and configured upstream headers from logged metadata

The synthetic request that carries MCP client headers into
add_litellm_data_to_request forwarded the caller's Host header, and
Request.url is built from it, so a caller chose the proxy_server_request
url and the metadata endpoint that every logging callback records.

_upstream_credential_headers also only knew the configured client side
auth header and the x-mcp- prefix family, so a header name declared in
mcp_servers.<name>.extra_headers reached logging metadata in cleartext.
Those names are admin chosen, so no prefix rule can recognize them; read
them off the server registry instead. The header is still forwarded
upstream, which is what extra_headers is for. authorization is left out
because clean_headers already strips it and claiming it here would move
authenticated_with_header on the oauth passthrough config.

The Responses bridge tests stub the server manager, so their fakes gain
the registry accessor the sanitizer now reads.

* fix(mcp): drop caller host from the sanitized header mapping too

The synthetic request stopped forwarding host, but the parallel sanitizer
did not, so a forged hostname still reached the guardrail payload and the
list_tools spend row. Drop it there as well.

Exempt the configured identity headers from the upstream credential set.
get_user_from_headers resolves end user attribution off the same request
this module reconstructs, and it only fills end_user_id when auth left it
unset, so claiming user_header_name or a user_header_mappings name would
lose attribution on the MCP paths that authenticate upstream.

Drop the isinstance guard on extra_headers entries: the field is typed
list[str], so the check is dead and basedpyright scores it.

* fix(mcp): accept a bare user_header_mappings entry when exempting identity headers

get_internal_user_header_from_mapping and get_customer_user_header_from_mapping
both normalize a single mapping to a one element list, and config_settings.md
documents the key as a dict. Iterating the bare form yields its keys instead,
so the exemption silently matched nothing and an identity header also named in
an MCP server's extra_headers was dropped after all.
2026-08-14 17:21:07 -07:00
mateo-berri
ff547be3e3 fix(cost-tracking): count dict-shaped web_search_call output items 2026-08-14 17:19:37 -07:00
mateo-berri
f2a10f6331 merge: litellm_internal_staging into litellm_vertex_batch_embeddings_translation 2026-08-14 17:17:16 -07:00
mateo-berri
e8c1fe8b11 fix(bedrock): fall back to the batch deployment model for unmapped record models 2026-08-14 17:11:24 -07:00
mateo-berri
ed33687422 feat(proxy): auto-suppress the no-Redis banner for confirmed single-worker deployments 2026-08-14 17:11:05 -07:00
Ilan Chemla
f99d0a4b38
feat(search): add Nimble as a search provider (#36347)
* feat(search): add Nimble as a search provider

Adds `NimbleSearchConfig` so `search_provider: nimble` works across the SDK,
the proxy /v1/search endpoint, the Search Tools dashboard, and spend tracking.

Nimble's /v2/search already uses the Perplexity unified spec's parameter names,
so the request transform is close to a pass-through. `search_domain_filter`
splits into include_domains/exclude_domains on the spec's `-` prefix, `country`
is upper-cased to the ISO form Nimble documents, and everything else is
forwarded so focus, search_depth, time_range and the rest stay reachable. On the
response side, snippet prefers `content` and falls back to `description`, and a
malformed body raises an attributed error rather than reporting an empty search.

Also tightens `BaseSearchConfig.get_supported_perplexity_optional_params` to
return `frozenset[str]` instead of a bare mutable `set`, which every caller
already treats as read-only.

* fix(search): surface Nimble error bodies instead of empty results

Greptile flagged that a null or absent `results` degraded to a successful empty
search. A search with no hits comes back as `"results": []`, verified against the
live API, so the field is now required and anything else raises the attributed
schema error the other malformed bodies already take.

Also unwraps Nimble's second error envelope. Collection failures return
`{"success", "task_id", "message"}` rather than the `{"detail"}` shape validation
errors use, and only the latter was being read.

Drops comments that restated the adjacent code.

* docs(search): drop the Nimble param list from the transform docstring

It restated the vendor's API reference, which the module docstring already links,
and would go stale the moment Nimble adds a focus mode.
2026-08-14 17:09:58 -07:00
Shivam Rawat
4110812b52
Merge pull request #33881 from BerriAI/litellm_fix_redis_spend_buffer_requeue_33872
fix(proxy): requeue Redis spend buffer transactions when the DB commit fails
2026-08-14 17:08:05 -07:00
mateo-berri
e06e2e6c12 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_36907_head 2026-08-14 17:06:18 -07:00
mateo-berri
b3729c50b0 fix(fireworks_ai): move top-level thinking into extra_body on the text completion path 2026-08-14 17:06:14 -07:00
tin-berri
2d3c3e3098
feat(shadow_eval): add reverse-direction shadow eval jobs (#36865)
Shadow eval only answered "should this key adopt this auto-router". Once a key
is on the router it is invisible to the feature, because the sampling gate skips
any request the shadowed router already served, so post-adoption quality
regressions go unmeasured.

Reverse mode inverts the arms: sample the traffic the router did serve and
duplicate it against a fixed baseline_model, judged by the same blind pairwise
judge. Same job table, same attempt rows, same aggregates.

real_* stays the arm the caller was served and shadow_* the duplicated one, so
in reverse real_model is the router's pick and shadow_model is the baseline. The
active-job slot becomes one per (key, direction) so both directions can run at
once, and tier attribution in reverse reads the control request's routing
decision rather than the shadow call's write-back.
2026-08-14 17:05:55 -07:00
mateo-berri
0e0c3ce0d8 Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_lit_5013_web_search_cost 2026-08-14 17:05:08 -07:00
mateo-berri
4193445647 Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_lit_5013_web_search_cost
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
#	tests/test_litellm/litellm_core_utils/prompt_templates/test_bedrock_converse_strict_tools_opus_47_48.py
2026-08-14 17:05:08 -07:00
mateo-berri
2f6f5c4961 fix(cost): reach tiered pricing for models without top-level per-token rates 2026-08-14 17:04:48 -07:00
tin-berri
d4d6bc2577
fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting (#34845)
The MCP sub-app is attached with app.mount("/mcp", ...) and a Starlette
mount never matches its bare prefix, so POST /mcp fell through to the
router's redirect_slashes 307. Behind a TLS-terminating ingress whose
peer address is not in uvicorn's forwarded-allow-ips (default: loopback
only) the redirect Location is built from the socket scheme as http://,
and MCP clients strip the Authorization header on the cross-origin
follow, so reconnects fail with ECONNRESET right after a successful
OAuth flow. The redirect also fires before auth, so the bare spelling
never returns the RFC 9728 WWW-Authenticate challenge that OAuth
clients need to start the flow.

Add an explicit /mcp route beside the existing /toolset/{name}/mcp and
/{name}/mcp spellings, forwarding to handle_streamable_http_mcp with
the same scope rewrite those routes already use (path=/mcp,
_original_path preserved for OAuth challenge URL selection). When the
mcp package is unavailable the route 404s, matching what the bare
sub-app serves on /mcp/ in that state. /mcp/, /mcp/{server},
/{server}/mcp and /toolset/{name}/mcp spellings are unchanged; the
exact-match route and the mount have disjoint match sets so
registration order cannot matter.
2026-08-14 17:04:32 -07:00
mateo-berri
e4f2ea12bc fix(responses_api): map bridged chat usage on guardrail-blocked replies
Move the blocked-usage mapping for /v1/responses next to
blocked_response_usage in guardrail_translation utils, map bridged chat
prompt/completion tokens to Responses API input/output tokens, and let
raise_passthrough_exception attach the blocked response so post-call
guardrail blocks report real usage
2026-08-14 17:04:27 -07:00
mateo-berri
9079e4c47b fix(proxy): return cost breakdown header values as a named tuple 2026-08-14 17:04:26 -07:00
mateo-berri
9a1e63c9f0 fix(caching): tolerate SSE chunk splits in anthropic stream cache writer 2026-08-14 17:04:25 -07:00
daleselaji-dev
80c37bfe3a fix(bedrock): resolve aliases in batch file records 2026-08-14 17:04:24 -07:00
Devin AI
48de8106ef fix(router): stop get_router_model_info from wiping cached pricing
Merge deployment model_info into a copy of the lru_cache'd get_model_info() dict and drop unset Nones, so Deployment's mirrored pricing defaults no longer overwrite built-in prices process-wide.

Fixes #36980

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-14 23:59:39 +00:00
mateo-berri
b14c4a8d45 fix(vector_stores): classify write endpoints before reads on substring collisions 2026-08-14 16:53:48 -07:00
Yassin Kortam
eb4b847268
fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown (#36961)
Anthropic's Models API declares max_input_tokens and max_tokens as nullable, not
optional, and the live vendor endpoint returns both keys on every entry. The
merged Anthropic-native listing dropped either key whenever LiteLLM could not
resolve a limit, so a client validating against a nullable-but-required schema
saw a malformed entry for any model the cost map does not know.
2026-08-14 16:52:39 -07:00
Yassin Kortam
2959465ea0
fix(openai,azure): return a length-truncated 200 when the output budget fits no token (#36859)
OpenAI and Azure GPT-5.x answer a chat request whose output budget cannot fit a
single visible token with a 400, while the same models return a length-truncated
200 one or two tokens higher. Agents that probe a model with a hardcoded
max_tokens of 1 read that 400 as "model unavailable".

The four chat request helpers now recognise the provider's own sentence and hand
back the length-truncated response the provider gives at a slightly larger
budget: finish_reason "length", empty content, zero completion tokens. Any other
400 still raises. Streaming is covered by the same seam, and the caller's budget
is never raised on their behalf.

The provider bills the prompt it processed but sends no usage object with the
400, so the prompt tokens are estimated with the same token_counter every other
usage-less path uses. Reporting zero would let a caller send an arbitrarily
large prompt with max_tokens 1 and be charged nothing.
2026-08-14 16:51:41 -07:00
mateo-berri
9027ab9485 Merge remote-tracking branch 'origin/litellm_internal_staging' into fireworks_nim_vllm_compat 2026-08-14 16:45:49 -07:00
mateo-berri
e94a97fcfc fix(cost_calculator): mirror the anthropic geo uplift in the token-type cost breakdown 2026-08-14 16:44:22 -07:00
mateo-berri
0ab23f5ce9 fix(anthropic): bill undetailed iteration cache writes at the 5m rate 2026-08-14 16:41:34 -07:00
Mateo Wang
870a8cf764
Merge pull request #36974 from BerriAI/litellm_vllm_dropdown_labels
fix(ui): distinguish hosted and local vLLM in the provider dropdown
2026-08-14 16:28:57 -07:00
devin-ai-integration[bot]
40d999b693
fix(mcp): keep admin-entered oauth endpoints in management reads (#36888)
* fix(mcp): keep admin-entered oauth endpoints in management reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover configured oauth endpoints on the config load path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-14 16:27:37 -07:00
mateo-berri
3b2ed3c018 fix(fireworks_ai): let extra_body thinking/reasoning_effort take precedence over chat_template_kwargs 2026-08-14 16:25:32 -07:00
yuneng-jiang
652f4cb8e4
Merge pull request #36982 from BerriAI/litellm_/revert-36837-ui-regression
Revert "fix(auth): stop the team fallback from widening model access" (#36837)
2026-08-14 16:19:17 -07:00
yucheng-berri
2fc39cde18
fix(langfuse): gate update_trace_keys behind an operator setting (#36862)
update_trace_keys lets a caller name which request metadata entries get copied
onto an existing trace, and the name is unrestricted. Sending
update_trace_keys: ["user_api_key_auth"] with existing_trace_id serializes the
resolved auth object, including the team callback credentials it carries, onto
the trace through Langfuse.trace(**trace_params). TraceBody is Extra.allow, so
an unexpected key ships rather than being dropped.

Any holder of a team key can do this and read the result in the destination the
team already logs to, so the feature is now inert unless an operator turns it on
with langfuse_enable_update_trace_keys.
2026-08-14 16:05:28 -07:00
Yuneng Jiang
9592a5447f
Revert "fix(auth): stop the team fallback from widening model access (#36837)"
This reverts commit ab2333b6c4.

Every Admin UI login mints its session key against the sentinel team_id
`litellm-dashboard`, and no LiteLLM_TeamTable row is ever created for it.
That lookup is therefore a provably-absent row on every UI request, which
#36837 turned into a hard refusal with no override, so the whole dashboard
404s.

Reverting restores the token-derived fallback. The model-access widening
#36837 closed is reopened and needs a re-land that exempts the UI sentinel
team.
2026-08-14 16:03:51 -07:00
Ahmed N
29fe342ead
fix(transcription): stop a zero output rate from zeroing transcription cost (#36914)
cost_per_second treated a declared-but-zero output_cost_per_second as a real
rate, so the output branch claimed the call and the elif locked out
input_cost_per_second. Every transcription model shipping
output_cost_per_second 0.0 next to a real input rate billed $0, which covers
43 of the 55 per-second entries in the cost map: all 36 deepgram models, both
assemblyai, both elevenlabs scribe, both groq whisper and azure-stt. Custom
deployments pairing the two fields the same way billed $0 as well

Take the output branch only when that rate is actually billable, so a zero
falls through to the input rate. Entries that duplicate one rate into both
fields, whisper-1 among them, keep billing exactly what they bill today
2026-08-14 15:12:37 -07:00
devin-ai-integration[bot]
865ed96765
fix(proxy): force prisma recreate on postgres cached-plan error (#36428)
`_query_first_with_cached_plan_fallback` recovers from Postgres's "cached
plan must not change result type" by recreating the Prisma client, which
drops both the server-side plans and the engine's client-side statement-name
cache. Since #30183 the shared reconnect path probes the writer with
`SELECT 1` first and skips the recreate when it answers, which is right for
the IAM token refresh it was added for and wrong here: the connection is
healthy, it is the session's prepared statements that are stale, so the probe
always passes and always vetoes the recreate. Callers now pass
`force_recreate` to skip that probe, and only the cached-plan fallback does.

Getting past the probe is not enough on its own. Both cooldown checks would
still skip the recreate for 15 seconds after any earlier reconnect, which
outlives the 10 second auth retry window, so a migration landing in that
window kept 503ing. `force=True` would fix that but would also let every
concurrent caller of the same burst kill the engine the first one just built.
The caller instead names the engine it observed before the query, and the
cooldown is waived only while that engine is still the live one, so the first
caller repairs the pool and the rest fall back to the normal cooldown.

That engine has to be the one the query actually ran on. `query_first` is a
top-level read, so with a read replica configured it is dispatched to the
reader and it is the reader's prepared statements that go stale, while
`writer_db` names a different engine with its own counter. The observation
and the cooldown comparison both go through `read_db`, added alongside
`writer_db` and backed by a `read_target` property on the routing wrapper
that `__getattr__` now dispatches through so the two cannot drift.

The observation carries the wrapper, not just its generation. `read_db`
resolves to the reader while it is available and to the writer once it is
not, and those counters are independent and both start at zero, so comparing
a bare number across that switch pits one engine's counter against another's.
Equal by coincidence waives the cooldown for an engine already replaced;
unequal gates a caller that needs the recreate. Identity settles it, and is
sound because the engine object is never re-pointed without the generation
also moving.

Three smaller holes on the way out. The waiver is withdrawn once a repair of
that same engine has been tried and failed, so a burst collapses onto one
attempt instead of each caller running its own recreate serially; the record
is keyed per engine rather than counted globally, so an unrelated reconnect
failure cannot suppress a stale reader's recovery and a writer failure cannot
evict the reader's record. And a forced recreate that the optimistic-lock
guard declines is no longer reported as a success on either the direct or the
heavy path, since the routing wrapper leaves the reader untouched in that
case; a decline is deliberately not counted as a failure, so the caller's own
backoff still gets its waiver on the next attempt.

A decline on the heavy path clears the dead-engine flag before raising. The
clear after the cycle is skipped by any raise, which is right for a failure
and wrong here, and the non-forced path already clears it on a decline, so
this restores that policy rather than inventing one. Stranding the flag would
route the next cycle back down the probe-free heavy branch, where the
refreshed generation matches and the recreate kills the healthy engine a
refresh just spawned, which is #29176.

Clearing that flag is necessary and not sufficient. The escalation check
re-arms it whenever the consecutive-failure count sits at the threshold, so a
decline that left the count alone sent the very next attempt back down the
same path. A decline is raised only at the generation guard, and the
generation moves only after a replacement connects, so a decline is proof
that a replacement succeeded and the count is reset on it.

Fixes #36418

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-14 15:04:01 -07:00
Scott Wilson
a62798de63 test(anthropic): type the Responses tool fixtures instead of suppressing
The two new `translate_tools_to_responses_api` calls carried
`# type: ignore[arg-type]`, which CLAUDE.md bans as LIT009: pyrightconfig.json
sets enableTypeIgnoreComments to false, so the comment silently does nothing and
the reportArgumentType error stands. Annotating the fixtures as
list[AllAnthropicToolsValues] makes both calls check clean with no suppression
at all.
2026-08-14 17:56:44 -04:00
Scott Wilson
9858d021ee fix(guardrails): record MCP tool guardrail evaluations and blocks in usage monitor
MCP tool calls run their guardrails against a throwaway LLM-shaped dict
built by `ProxyLogging._convert_mcp_to_llm_format`, not against the dict
the tool call is logged from. `@log_guardrail_information` therefore
appended `standard_logging_guardrail_information` to that throwaway
dict's metadata bucket, where `get_standard_logging_object_payload`
never saw it, so the Guardrails Monitor reported zero evaluations and
zero blocks for all MCP traffic.

Thread the request's `litellm_logging_obj` into `pre_call_tool_check`
and `_create_during_hook_task` and bridge the guardrail records onto it:

- Seed `data["litellm_logging_obj"]`, which unified guardrails read and
  pass into `apply_guardrail`.
- Call `_sync_guardrail_info_to_logging_obj` in a `finally`, which is
  what native guardrails need and what makes the block path work: a
  blocked call raises straight out of `pre_call_tool_check`, so the
  record has to be attached before the exception leaves the frame.

Only the guardrail evaluation records are copied. The synthetic
request's messages and tool arguments are deliberately left behind --
they can carry end-user data and nothing in the monitor needs them.

In `call_mcp_tool`, flush the failure handlers before
`post_call_failure_hook` so the `status="failure"` standard logging
object exists when `_ProxyDBLogger.async_post_call_failure_hook` writes
the spend-log row the monitor's "Total Blocked" counts. Both handlers
gate on `should_run_logging("sync_failure")` / `("async_failure")` and
then mark it, so the `@client` wrapper's own post-raise logging is a
no-op and nothing is double-counted -- the same pattern
`_fire_mcp_tool_call_logging` already uses for `isError=True`.

Threaded through every MCP tool entry point: the managed-server path,
the local-OpenAPI registry path, the legacy registry fallback, and the
Responses API's `_execute_tool_calls`.
2026-08-14 17:52:36 -04:00
mateo-berri
08966c842b test(vector_stores): drop redundant route-map comment 2026-08-14 14:12:09 -07:00
mateo-berri
d330b64943 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_34850_head 2026-08-14 13:58:38 -07:00
mateo-berri
0a81e1b222 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_azure_ai_docs_index_write_grant_rc 2026-08-14 13:58:28 -07:00
mateo-berri
aaa619441e chore: merge litellm_internal_staging into litellm_lit_4868_cache_write_split 2026-08-14 13:57:47 -07:00
mateo-berri
ae2a5e1472 fix(ui): distinguish hosted and local vLLM in the provider dropdown 2026-08-14 13:39:37 -07:00
Scott Wilson
4c49d03732 fix(anthropic): preserve optional Responses tool properties
Translating Anthropic tools left the outbound function-tool `strict` unset,
which the Responses API does not read as non-strict. OpenAI's function-calling
docs say strict mode requires every field in `properties` to be marked
required, and with `strict` omitted the schema gets normalized to satisfy that
instead of being rejected. What users see is a tool whose `required` lists
every property, so models fill optional Anthropic tool arguments with empty
values. Send `strict` explicitly so an unset value stays non-strict and an
explicit `strict: true` still reaches the provider

On the Chat Completions adapter, `strict` was also missing from
`mapped_tool_params`, so a tool-level `strict` was merged into the OpenAI
function `parameters` schema (mutating the caller's `input_schema` along the
way) instead of being set on the function. Map it to `function.strict` and
leave it unset when the caller omits it, since Chat Completions already
defaults to non-strict
2026-08-14 16:19:45 -04:00
shivam
3838969527 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_redis_spend_buffer_requeue_33872 2026-08-14 19:07:19 +00:00