* fix(vertex): preserve Gemini Embedding 2 usageMetadata for cost tracking
* style(vertex): apply ruff format to batch_embed_content_transformation
* fix(vertex): bill files/ image refs in Gemini embedContent at per-image rate
Resolved files/... references whose mime type is an image were not detected
by _is_image_element, so image_count stayed 0 and generic_cost_per_token fell
back to the text token rate instead of input_cost_per_image. Thread the
resolved_files mapping into the usage builder so resolved image references are
counted and billed per image. Also modernize the _flatten_input return
annotation to satisfy the ruff UP006 strict gate.
* fix(vertex): bill Gemini embedding audio per-second and stop video+audio double-billing
Audio-only embedContent responses set audio_tokens, but generic_cost_per_token only
charges audio via input_cost_per_audio_token. gemini-embedding-2 prices audio via
input_cost_per_audio_per_second, so spend stayed at $0. Plumb a new
audio_length_seconds field through PromptTokensDetailsWrapper, parse it in
_parse_prompt_tokens_details, and bill it from _calculate_input_cost. The vertex
embedding transformation derives audio_length_seconds from audio_tokens using
the documented 32 tokens/sec Gemini rate.
The 1-token text floor that protects video billing only fired when no other
modality was billable, but audio presence flipped that flag, leaving text_tokens
at zero for video+audio responses. generic_cost_per_token then rewrote
text_tokens to prompt_tokens minus audio_tokens (the video token count),
charging video tokens as text on top of the per-second video cost. The rewrite
trigger is text_tokens == 0 and image_count == 0; align the floor with that
trigger and ignore audio_tokens.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Build form_data_dict in one pass with groupby instead of rescanning form_items per field name, and assert on the files list directly in the boundary regression test so repeated field names are not collapsed by dict().
Co-authored-by: Cursor <cursoragent@cursor.com>
Records main as an ancestor of internal_staging so the staging->main
promotion (#31384) merges cleanly. Resolves in staging's favor; changes 0
files (staging already supersedes every main-side hotfix). MUST be merged
as a real merge commit (not squash/rebase) or the link to main is lost.
* docs(readme): add Deploy on AWS/GCP with Terraform section
Adds a quickstart for the two published Terraform modules on the public
registry (BerriAI/litellm/aws and BerriAI/litellm/google). Copy-paste
main.tf for each cloud, the one-time GCP Artifact Registry remote-repo
command, and pointers to the registry pages for the full input surface.
Sits inside the Get Started section, between the gateway/SDK table and
Run in Developer Mode -- where someone scanning the README for "how do I
deploy this" will land.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(readme): add 1-click deploy buttons for AWS + GCP
GCP gets the real 1-click: Open in Cloud Shell badge that clones the repo
and walks through `terraform apply` via the existing DeployStack
tutorial (already shipped at terraform/litellm/gcp/examples/default/
TUTORIAL.md). User just picks a project.
AWS gets a soft 1-click: a Launch in AWS CloudShell badge that opens an
in-browser, already-authenticated shell. User runs four commands
(clone + cd + cp tfvars + terraform apply) once inside. There's no
native AWS deeplink that pre-clones a repo + runs a tutorial -- CFN
"Launch Stack" + CodeBuild would be needed for that, and that's a
separate piece of work.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(readme): move AWS + GCP deploy buttons next to Render button
* docs(readme): unify deploy button sizes and badge styles
* docs(readme): bump deploy button height to 48 to match Render/Railway
* docs(readme): bump AWS/GCP badge height to compensate for SVG padding
* docs(readme): bump AWS/GCP badge height to 72
* docs(readme): bump AWS/GCP badge height to 84
* fix(readme): make deploy buttons same height (48px)
https://claude.ai/code/session_01MxQRMHSDXbqJh74rF86UBc
* docs(readme): flag GCP project ID substitution in image_registry
* docs(readme): equalize deploy button heights and fix Cloud Shell button font
GitHub rewrites an image's height attribute to "height: auto; max-height: Npx", which only caps and never stretches, so each image renders at its intrinsic height. The AWS/GCP shields badges are intrinsically 28px while the Render/Railway buttons are 40px, leaving the row uneven regardless of the height="48" we set. Replace the two shields badges with committed 40px PNGs so all four header buttons render at the same 40px.
Also swap the Cloud Shell button from open-btn.svg to open-btn.png. The SVG renders its label as live text with font-family "Roboto, Sans" and no generic fallback; since neither font exists in GitHub's render environment, the text fell back to a serif (Times New Roman). The PNG bakes in the correct typeface.
* docs(readme): collapse Railway deploy anchor to a single line
The Railway button wrapped its img across indented lines, so the anchor contained leading and trailing whitespace. GitHub underlines link content, rendering that whitespace as a small blue underline beside the button. Put the anchor on one line like the other three buttons so there is no inner whitespace to underline.
* Add Claude Fable 5 cost map entries as a data-only hotfix
Backports only the model map changes from #30064 so deployments on
released litellm versions pick up Fable 5 pricing, context window, and
the adaptive thinking flag through the hosted cost map fetch without
upgrading. Includes the supports_sampling_params flag on the 28
Fable 5 / Opus 4.7 / Opus 4.8 entries (ignored by released code, read
by the gating that ships with the next release) and the matching
one-line schema declaration so the map validation test passes.
https://claude.ai/code/session_01MZarYYT3aS7DxaNjoax6Gm
* feat: make rust OCR async-first
* docs: clarify rust provider call flow
* docs: clarify OCR provider transform contract
* docs: note Tokio route contract
* fix: address OCR bridge review comments
* docs: bound rust OCR HTTP exception
* feat: generate rust providers from registry
* chore: move rust provider registry into core
* chore: source rust providers from endpoint registry
* fix: satisfy OCR lint budget
* fix: reduce OCR basedpyright argument errors
* fix: address OCR greptile feedback
* fix: align rust OCR request preparation
* fix: resolve OCR CodeQL alerts
* fix: avoid duplicate Rust OCR authorization header
* ci: rerun CircleCI
---------
Co-authored-by: shin-berri <shin-laptop@berri.ai>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
Co-authored-by: ishaan-berri <155045088+ishaan-berri@users.noreply.github.com>
Passthrough multipart uploads used form.items() and a files dict, so only the last file under a repeated field name reached the upstream. Read multi_items() and send httpx a list of file tuples instead.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(mcp): add require_key_mcp_access_defined to stop keys inheriting team MCP servers
By default a virtual key that grants no MCP servers of its own inherits its
team's full MCP server list. The new general_settings flag
require_key_mcp_access_defined (default false) flips this so the team list
acts purely as a ceiling: a key reaches only the servers it grants explicitly
(or via an access group), and inherits none. This mirrors the existing
require_end_user_mcp_access_defined setting.
The default is unchanged, so existing deployments keep today's behavior until
they opt in. The no-mcp-servers sentinel and key access-group grants are
unaffected.
* docs(mcp): note require_key_mcp_access_defined effect in resolver docstring
* fix(cost-map): retarget mistral-medium-latest to Medium 3.5 and add date-pinned aliases
Mistral repointed the rolling mistral-medium-latest alias from Medium 3.1
to Medium 3.5, but the static cost map still carried Medium 3.1 specs,
showing wrong pricing/context in the model hub and undercharging spend by
about 3.75x (LIT-3883).
Update mistral/mistral-medium-latest to Medium 3.5 ($1.50/$7.50 per 1M,
256K context, reasoning + vision), add the bare date-pinned aliases
mistral/mistral-medium-2604 (Medium 3.5) and mistral/mistral-medium-2508
(Medium 3.1) that match Mistral's real API model ids, and add
supports_reasoning to mistral/mistral-medium-3-5.
Apply every change to both model_prices_and_context_window.json and the
bundled litellm/model_prices_and_context_window_backup.json so the two
stay in sync, and extend the regression tests to lock the resolved
get_model_info values and the main/backup parity for all touched models.
* test(cost-map): force local cost map in mistral-medium-latest resolution test
get_model_info reads litellm.model_cost, which is fetched from the remote
main branch at import time when LITELLM_LOCAL_MODEL_COST_MAP is unset. Until
this PR lands on main, that remote map still carries the pre-merge Medium 3.1
pricing, so the assertion was only passing when the remote fetch happened to
fail and fell back to the bundled backup. Force the local cost map (the same
fixture pattern the other get_model_info tests use) so the alias resolution is
verified deterministically against the in-repo file.
* feat(spend): store litellm_call_id on spend logs for DB-to-trace correlation
Successful spend logs keyed request_id to the provider response id while
tracing uses x-litellm-call-id, so a DB row could not be correlated with its
trace; this only worked for failures, where request_id already fell back to
the call id. Add a nullable litellm_call_id column to LiteLLM_SpendLogs,
populate it in get_logging_payload, and surface it in the spend logs read
endpoints so correlation works both directions for successful calls
Fixes LIT-3868
* chore: sync schema.prisma copies from root
* test(spend): cover cache-hit and missing-response-id paths for litellm_call_id
Lock the intended behavior surfaced in review: on a cache hit request_id gets
the uniqueness suffix while litellm_call_id stays the raw call id, and when the
provider returns no id request_id falls back to the call id so both columns
match. Both assertions fail when the populate line is reverted
* test(spend): ignore litellm_call_id in spend logs payload comparisons
get_logging_payload now always writes litellm_call_id, so the full-payload
comparisons in test_spend_management_endpoints.py saw an unexpected key and
failed. litellm_call_id is a per-request runtime uuid like request_id, which
is already ignored, so add it to ignored_keys
* test(logging): ignore litellm_call_id in gcs pubsub spend logs comparison
The gcs pubsub spend logs payload comparison flags any key present in the
actual payload but absent from the golden snapshot. get_logging_payload now
always emits litellm_call_id, a per-request runtime uuid like request_id which
is already ignored, so add it to ignored_keys
* refactor(spend): store litellm_call_id in spend log metadata, drop column
Switch DB-to-trace correlation off a dedicated column and onto the existing
metadata JSON, avoiding a schema migration entirely. litellm_call_id is now
written into spend log metadata (already selected and re-hydrated on the read
paths) instead of a new LiteLLM_SpendLogs column, so the three schema.prisma
copies and the migration are reverted and the read SELECTs go back to their
original form. Correlation is queryable via metadata->>'litellm_call_id'
Trade-off: an unindexed JSON lookup rather than an indexed column; acceptable
for this use case and removes all migration risk
* refactor(spend): thread litellm_call_id into _get_spend_logs_metadata
Set litellm_call_id beside the other computed metadata values inside
_get_spend_logs_metadata rather than mutating clean_metadata back in the
caller, matching how applied_guardrails, cost_breakdown and the rest are
threaded. No behavior change; the value still comes from kwargs with a
litellm_params fallback
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(proxy/client): redact api key from key/info client error messages
The keys management client builds GET /key/info?key=<key> and lets the
requests HTTPError propagate. str(HTTPError) renders the failing request URL
verbatim ("... for url: .../key/info?key=sk-..."), so any caller that logs the
exception leaks the full key; the 401 branch leaked the same way through
UnauthorizedError(str(orig_exception))
Redact both branches with the existing redact_secrets helper so the
secret-bearing query param is scrubbed to ?REDACTED while the status code,
reason, and response object are preserved. Server-side responses already mask
the key, so this closes the remaining client-side surface
* fix: preserve key info unauthorized response
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* feat(mistral): support Mistral OCR 4 (mistral-ocr-4-0)
Add the mistral/mistral-ocr-4-0 model to the cost map and reprice
mistral/mistral-ocr-latest, which now resolves to OCR 4 server-side,
at $4 / 1000 pages. Add the include_blocks param so callers can request
OCR 4's paragraph-level bounding boxes and typed content blocks.
OCR 4's new per-page response fields (blocks, confidence_scores, tables,
hyperlinks, header, footer) already pass through transform_ocr_response
via the extra="allow" config on OCRPage; add a regression test pinning
that behavior alongside cost and param coverage.
* fix(mistral): revert unverified OCR 4 annotation_cost_per_page bump
Mistral's published OCR 4 pricing lists $4/1000 pages for the API and no
separate annotation rate; the $5/1000 figure is the distinct Document AI
(Studio) tier. The earlier 0.003 -> 0.005 bump on annotation_cost_per_page
had no cited source, and ocr_cost() never reads that field (it bills off
ocr_cost_per_page), so the value is documentation-only.
Revert annotation_cost_per_page to the existing 0.003 convention for both
mistral-ocr-latest and mistral-ocr-4-0, keeping only the verified, tested
ocr_cost_per_page: 0.004 change.
* fix(mistral): set OCR 4 annotation_cost_per_page to verified $5/1000 rate
Verified against Mistral's authoritative sources: the pricing page, the
OCR 4 announcement, and the ocr-4-0 model card all list OCR 4 at $4/1000
pages for basic OCR and $5/1000 for annotated pages (Document AI). The
$5/1000 figure is the annotated-pages rate, which is exactly what
annotation_cost_per_page encodes, mirroring the original OCR entry's
0.001 basic / 0.003 annotated split.
Restore annotation_cost_per_page to 0.005 for mistral-ocr-latest and
mistral-ocr-4-0; the earlier revert to 0.003 was based on an incomplete
reading that treated Document AI as a separate product. ocr_cost_per_page
stays 0.004, which is the value billed by ocr_cost().
* fix(mistral-rust): include_blocks in Rust OCR supported params
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* feat(aiml): add openai/gpt-image-2 image model
Adds aiml/openai/gpt-image-2 to the cost map and teaches AimlImageGenerationConfig
to route OpenAI-style image models through the upstream OpenAI request schema
instead of the AI/ML flux schema. Without this, size, n, and response_format would
be remapped to image_size/num_images/output_format, which the gpt-image-2 endpoint
on api.aimlapi.com does not accept.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* chore(aiml): note gpt-image-2 flat-rate pricing basis; apply ruff format
Documents in the cost-map notes that output_cost_per_image is AI/ML's
published medium-quality rate, billed as a flat per-image price like the
other aiml image entries. Reformats the touched files under the repo's
ruff formatter (migrated from black in #31317).
* fix(aiml): drop /v1/images/edits from gpt-image-2 supported_endpoints
LiteLLM only implements an image generation transformer for AIML, so
listing /v1/images/edits overclaimed support. Align with every other
aiml image entry, which lists only /v1/images/generations.
* style(aiml): format transformation.py at line-length 88
The repo formats litellm/ with ruff at line-length 88 (Makefile/CI call
sites), while ruff.toml's global 120 only governs E501/import sorting.
Reformat the transformer to 88 so make format-check / CI lint pass, and
restore the test files to their original layout since tests/ is not part
of the auto-formatted tree.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* ci(image-scan): add Grype image scan for OS + library CVEs
Builds each of the 6 Dockerfiles via a matrix and scans the resulting image
with Grype (pinned v0.114.0, sha256 verified), failing on fixable HIGH or
CRITICAL across both OS/apk and language packages. This catches the layer
osv-scan is structurally blind to (Wolfi/apk OS packages and vendored deps
like prisma's node engine), which is the structural reason the openssl CVE
slipped past CI and a customer's image scanner flagged it.
Skipped on fork PRs so an outside contributor cannot run arbitrary code on
our hosted runner via a malicious Dockerfile RUN line. The same pattern is
used by guard-fork-dependencies.yml.
Grype runs as a pinned binary with a verified checksum, so there is no
mutable-tag GitHub Action in the dependency chain and no vendor credentials
in the scan job. The job uses read-only contents permissions and an empty
top-level permissions block.
* ci(image-scan): scan only Dockerfile.non_root (rootless target)
All Dockerfile variants share the same wolfi base and apk set today, so a single scan of Dockerfile.non_root gives the same OS-layer coverage at one-sixth the build cost. Dockerfile.non_root is the rootless variant we ship (USER 65534), so the scan tracks the image customers actually run. Matrix-scan if the variants ever diverge.
* ci: retrigger checks (proxy_pass_through_endpoint_tests flaked on prior run)
Fixes#29794. Adds bare, gemini/, and vertex_ai/ entries copied from preview models so proxy cost tracking works for GA model names.
Co-authored-by: Cursor <cursoragent@cursor.com>
The namespace configured under cache_params was only applied to get/set/
increment paths. Operations that take keys through other code paths (the Lua
scripts registered via async_register_script, delete, scan_iter, rpush, lpop,
get_ttl, and the sync increment_cache) hit raw keys. With a namespace set, the
rate limiter ({key}:tokens/requests/window), pod-lock release, and budget
limiters wrote keys outside the configured prefix, breaking multi-tenant key
isolation and leaving those operations reading keys the namespaced writes never
created.
check_and_fix_namespace is now applied uniformly across every key-taking
RedisCache operation. It is a no-op when no namespace is configured, so
deployments without a namespace are unaffected. The prefix is prepended ahead of
any {hash-tag}, so Redis Cluster slotting is preserved.
Resolves LIT-3374
* chore(lint): widen ruff budget slack to 10% of baseline for high-volume ANN rules and PLR0913
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* chore(lint): drop PLR0913 from strict gate to roll out rules gradually
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(lint): ratchet-guard rising baselines even when slack is cut to mask them
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Ignore the compiled, platform-specific Rust extension output (litellm/rust_bridge/_native*.so/.pyd) and the litellm-rust/target/ build dir so local maturin/cargo builds don't show up as untracked files.
Also drop the two stale self-referential .gitignore entries; .gitignore is tracked, so ignoring it did nothing except add confusion.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(otel): hashable scope for _emit_once when guardrail_mode is list
`_emit_once` keys `spans_logged` by `(class, id, *scope)`. When a
guardrail entry's `guardrail_mode` arrives as a `List[GuardrailEventHooks]`
(the shape Presidio expands to with `output_parse_pii: true`, and the
shape `event_hook` carries for any `mode: [...]` in config), the tuple
contains a list and `spans_logged.get(dedupe_key)` raises
`TypeError: unhashable type: 'list'`. On the post-call path this fires
inside the logging callback and is swallowed; the request returns 200 but
the OTEL `guardrail` span is silently dropped. On the blocking path the
same error surfaces as HTTP 500.
Adds `_freeze_for_dedupe`, a small recursive normalizer that turns lists
and tuples into tuples, sets into frozensets, dicts into frozensets of
`(key, value)` pairs, and falls back to `repr` for arbitrary
unhashables. Applied inside `_emit_once` before the dict lookup, so all
three callsites are protected without touching the guardrail-specific
callsite. Helper assumes acyclic input; `guardrail_mode` values are
built fresh from config (str enums, lists of str enums, TypedDict of
str/list-of-str), so no cycle can arise in practice.
Regression tests in `TestOpenTelemetrySpanDedupe` cover the list crash,
distinct-list-scope collision, dict and set scope parts, and an
end-to-end `_create_guardrail_span` exercise that confirms exactly one
`guardrail` span is emitted across repeated lifecycle entrypoints. Each
new test fails on a reverted helper (4/4 mutation kill)
* fix(otel): cap _freeze_for_dedupe recursion depth and ignore in recursive detector
CI's recursive_detector blocks new recursive functions in litellm/ unless they
are in the allowlist with a documented bound. Cap the helper at 16 levels and
return repr(value) past the cap; this is well past the realistic depth of
guardrail_mode (1-3 levels) and means a future caller passing a cyclic
container can no longer push the proxy logging path into a RecursionError.
Add a regression test that exercises the cycle path.
* refactor(otel): annotate _freeze_for_dedupe return as a HashableScope union
Per review feedback from @mateo-berri: replace the loose `-> object` annotation
with a recursive `HashableScope` union (str | int | float | bool | bytes | None
| Tuple[HashableScope, ...] | FrozenSet[HashableScope]) so the helper's contract
is visible at the signature. Replace the `try/except hash(value); return value`
passthrough with an explicit isinstance check over the hashable-scalar types so
the type checker can narrow without requiring `cast(Hashable, value)` on the
return. Symmetric: dict keys also flow through the freezer (a TypedDict key is
already a string in practice, so behaviorally identical). All 16 regression
tests still pass; mutation kill behavior preserved
* fix: avoid explicit casting
---------
Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>
* feat(mcp): add mcp_xff_num_trusted_hops to harden XFF client IP resolution
MCP per-server IP access control reads the client IP from X-Forwarded-For
and trusts the leftmost entry. Behind an append-style proxy or load
balancer (AWS ALB, nginx with $proxy_add_x_forwarded_for, HAProxy, Envoy,
Cloudflare), a client can prepend an arbitrary value to the header, so the
leftmost entry is attacker-controllable even when the direct peer is a
trusted proxy. An attacker can therefore spoof an internal IP and reach
servers marked available_on_public_internet=false.
This adds an optional mcp_xff_num_trusted_hops general setting modelled on
Envoy's xff_num_trusted_hops. When set to N, the client IP is read N entries
from the right of the chain (where N is the number of trusted appending
proxies in front of the gateway) instead of the leftmost value, so any
entries a client prepends are ignored. It composes with mcp_trusted_proxy_ranges,
which still validates the direct peer, and only takes effect once that check
passes; without a validated direct peer the gateway keeps failing closed, so
hop counting cannot be abused by a direct-to-pod attacker. The chain must
contain at least N valid entries or resolution fails closed.
Default is unset, preserving existing behaviour.
* chore(ui): regenerate dashboard schema for mcp_xff_num_trusted_hops
* fix(mcp): warn when mcp_xff_num_trusted_hops is below the minimum
A 0 or negative value is silently treated as disabled, which could leave
an operator believing they enabled append-style X-Forwarded-For hardening
while client IP resolution stays on the spoofable leftmost value. Emit a
warning, consistent with how the module already surfaces invalid CIDR
config, so the misconfiguration is visible in logs.
* fix(mcp): reject mcp_xff_num_trusted_hops < 1 at config-parse time
Add a ge=1 bound to the ConfigGeneralSettings field so the
update_config_general_settings path rejects 0 and negative values with a
clear validation error instead of accepting them, and self-documents the
valid range. The runtime warning stays as defense-in-depth for raw-dict
config that bypasses model validation.
* style(mcp): black-format ip_address_utils.py
* fix(mcp): fail closed when mcp_xff_num_trusted_hops is set but invalid
A present-but-invalid mcp_xff_num_trusted_hops (non-integer, or below 1)
previously made _resolve_num_trusted_hops return None, which the caller
treated identically to "unset" and silently fell back to the legacy
leftmost X-Forwarded-For value. An operator who set the value to harden
client IP resolution but typo'd it would get weaker security than before,
with no fail-closed signal.
Model the setting as a tagged union (_HopCountUnset, _HopCountInvalid,
_HopCount) so the three states are distinct: unset keeps the legacy path,
a valid count drives hop-counting, and an invalid value fails closed
(returns "") instead of reverting to the spoofable leftmost address. The
caller matches on the union exhaustively.
Add a parametrized regression test asserting get_mcp_client_ip returns ""
for 0, -1, "abc", and 1.5 even with a spoofed internal leftmost entry,
and update the resolver unit tests for the new return type.
* fix(streaming): word-sliced cache replay for stream=true cache hits
* fix(streaming): align mypy and replay happy-path test with word-sliced cache replay
* fix(streaming): short-circuit whitespace-only content in cache replay splitter
* fix(streaming): emit tool_calls/function_call only on first replay slice
* refactor(streaming): drop dead delattr guard in cache replay
A non-None usage on the replay base object always lives in
__pydantic_extra__ (it is attached via setattr earlier in the same
function), so delattr can never raise here; the try/except AttributeError
that silently swallowed a failure was dead defensive code that could only
ever hide a real regression, so it is removed in both the async and sync
generators.
Also switches the new replay annotations from typing.List to the builtin
list to satisfy the strict ruff UP006 gate and drops the unused
PLR0915 noqa directives (the rule is not enabled in this repo's ruff
config, so RUF100 flagged them).
* fix(streaming): drop carried-over metadata from later cache replay slices
The word-sliced cache replay deep-copies the full ModelResponseStream per
slice, so reasoning_content, thinking_blocks, logprobs, enhancements,
annotations and the rest of the per-message metadata rode on every slice, not
just the first. Downstream handlers that accumulate streamed deltas would
collect each one once per slice, e.g. duplicating a cached reasoning trace N
times on a stream=true cache hit.
Later slices are now rebuilt as a content-only delta with choice-level logprobs
and enhancements stripped, so the whole metadata class stays on the first slice.
Adds async (logprobs) and sync (reasoning_content/thinking_blocks/logprobs/
enhancements, plus annotations) regression tests
---------
Co-authored-by: Mateo <277851410+mateo-berri@users.noreply.github.com>
* fix(ci): point OSS contributor workflows to litellm_oss_staging
Workflow triggers and guard error messages incorrectly referenced litellm_oss_branch; update them to the branch we actually use for external contributions.
* fix(ci): include test-rust.yml in litellm_oss_staging rename
Missed test-rust.yml when updating OSS contributor target branch references.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
An oauth2 MCP server with delegate_auth_to_upstream=true never prompted the
user to sign in. On an unauthenticated initialize the gateway answered locally
(200, no tools) and emitted no WWW-Authenticate, so clients like Claude Desktop
either connected empty or hit "OAuth probe timeout after 10000ms".
#30124 added a bare `continue` in _raise_preemptive_401_for_unauthenticated_servers
to stop sending LiteLLM's gateway authorization_uri challenge for delegate-auth
servers, expecting the upstream to emit its own challenge. On initialize the
gateway never probes upstream, so no challenge ever reached the client.
Replace the `continue` with a preemptive 401 carrying the proxied
resource_metadata (RFC 9728) challenge, the same form passthrough servers and
MCPUpstreamAuthError already use. This keeps #29770 fixed (still no
authorization_uri) while restoring the upstream PKCE sign-in prompt.
* fix(mcp): resolve toolset tools by the server's known prefix
Toolsets store {server_id, bare tool_name} and reconcile that against the
live prefixed tool name at list time. The reconciliation chopped the live
name at the first MCP_TOOL_PREFIX_SEPARATOR with no server context, so a
server whose prefix contains the separator (a hyphenated alias, or the
UUID server_id used as the prefix when a server has no alias) had its
tools silently dropped from /toolset/<name>/mcp while listing fine
everywhere else. Strip the exact known prefix for the tool's server_id
instead of guessing the boundary, on both the resolve and filter sides
Also render toolset tools as {server-prefix}-{tool} in the dashboard
picker result and chips; this is display only, the persisted record
stays {server_id, bare tool_name}
Resolves LIT-3419
* test(mcp): add focused unit tests for strip_known_server_prefix
Cover the LIT-3419 cases directly on the helper with real MCPServer
objects: clean prefix round-trip, hyphenated alias, UUID server_id
fallback, unprefixed passthrough, and the server=None legacy fallback
* fix(mcp): warn loudly when X-Forwarded-For is present but use_x_forwarded_for is off
When a request carries an X-Forwarded-For header but use_x_forwarded_for is
unset, get_mcp_client_ip silently falls back to the direct peer's IP (the load
balancer / reverse proxy). That peer almost always sits inside
mcp_internal_ip_ranges, so the 'Internal network only'
(available_on_public_internet: false) restriction trusts every external caller
as internal and effectively exposes those servers.
Emit a one-shot loud error pointing the operator at use_x_forwarded_for instead
of hard-failing: on a deployment with no load balancer, a crafted
X-Forwarded-For header must not be able to take the service down, and a one-shot
log keeps a flood of crafted headers from spamming the logs.
* fix(mcp): re-arm XFF-disabled warning on config change and harden test assertion
Address PR review: tie the one-shot warning flag to the observed
use_x_forwarded_for value so it re-arms whenever the setting is seen enabled,
restoring the diagnostic on a later rollback to disabled. Also assert against
str(call_args) so the test survives a positional-to-keyword logger refactor.
* fix(mcp): correct misleading no-trusted-proxy warning for XFF access control
* test(mcp): assert the no-trusted-ranges warning was logged instead of relying on StopIteration
* fix(proxy): stop double-decrypting email/slack alerting env vars in get_config
proxy_config.get_config() already returns environment_variables decrypted
(the DB overlay decrypts them in _update_config_fields, and YAML values are
plaintext), so the /get/config/callbacks slack and email blocks were running
decrypt_value_helper() a second time on plaintext. That second decrypt always
failed and the helper swallowed the error and returned None, so every SMTP_*
field came back blank when the Admin UI reloaded the email settings, and the
proxy logged a misleading "Did your master_key/salt key change recently?"
error even when nothing changed.
Consume the already-decrypted values directly, matching process_callback's
handling of the same dict for langfuse/datadog/etc. Sensitive-value masking
is preserved.
Fixes#19221
* fix(proxy): preserve a cleared slack webhook instead of falling back to OS env
Use an explicit is-not-None guard rather than truthiness when deciding whether
to fall back to os.getenv for SLACK_WEBHOOK_URL. With `or`, a webhook the admin
cleared (stored as "") is falsy and would surface a stale SLACK_WEBHOOK_URL from
the OS environment; only a truly absent key should trigger the OS lookup. No
decryption is reintroduced.
Both tests were xfail(strict=True) for known proxy bugs: /team/new writing
budget_limits as a raw list (Prisma 500) and custom per-token pricing leaking into
the shared cost map for sibling deployments. Both are fixed, so the tests pass and
strict mode reports the unexpected pass as a failure. Remove the markers (as their
reasons instructed) so they run as plain regression guards; docstrings updated to
describe the regression each now pins.
The App Router migration moved pages to deeper path segments and the proxy
can be mounted under a sub-path (e.g. /litellm behind a reverse proxy). Local
logo asset paths were emitted without the server root prefix, so they resolved
off the origin root and 404'd. Route every local logo src through a single
resolver that prefixes the live server root path and leaves external URLs
untouched, fixing provider, guardrail, vector store, callback, MCP and
audit-log logos at any route depth and root path.