Commit graph

39895 commits

Author SHA1 Message Date
Tin Chi Lo
9927d458bc refactor(mcp): delete the unwired V1PerUserTokenStore adapter (step 1b piece 5)
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
Piece 4 replaced V1PerUserTokenStore with the v2-native chain at the composition root, leaving the
adapter with no callers, so remove it and its test. The shared v1 read/refresh core
(resolve_user_oauth_access_token and friends) stays - delegate's egress in server.py still uses it -
and comes out with the delegate migration.
2026-06-25 23:51:46 -07:00
Tin Chi Lo
a5dee6a05c feat(mcp): wire the v2-native per-user OAuth store into the resolver (step 1b piece 4)
Assemble Cached(Refreshing(V2PerUserTokenStore)) at the composition root and replace
V1PerUserTokenStore in mcp_server_manager. The chain is built lazily on first fetch (its cache/DB/
Redis collaborators are LiteLLM globals not ready at import); when Redis is wired it uses the
cross-replica path (DualCache cache + SET NX PX coordinator), else the in-process defaults. The DB
read, refresh-grant POST, and persist acquire their globals per call like v1. authorization_code
resolution now reads/refreshes through the v2-native lifecycle, not v1's core.
2026-06-25 23:48:58 -07:00
Tin Chi Lo
a81242d202 feat(mcp): Redis SET NX PX distributed lock (step 1b §1.5)
The concrete DistributedLock the RedisRefreshCoordinator elects refreshers with: acquire is an atomic
SET key NX PX ttl (first caller wins, entry self-expires so a crashed holder can't wedge refresh),
release is DEL, is_held is EXISTS. The async Redis client is injected (the client from LiteLLM's
RedisCache in prod), so it is unit-testable with a fake. Any Redis error degrades to not-acquired /
not-held so a cache blip causes an extra refresh, never a crash on the resolve path.
2026-06-25 23:30:53 -07:00
Tin Chi Lo
f8b07a089f feat(mcp): Redis SET NX PX refresh coordinator (step 1b §1.5)
The cross-replica RefreshCoordinator that plugs into the foundation's RefreshingTokenStore seam: a SET
NX PX lock elects one worker to refresh per (user, server) while the rest wait for it and re-read the
token it persisted, so a rotating refresh_token is used once across the fleet, not once per worker. The
lock self-expires (PX) so a crashed holder can't wedge refresh; a loser falls back to a bounded re-read
and the surrounding store re-checks expiry next fetch, so a crash self-heals. The lock (a thin Redis
SET NX/DEL/EXISTS wrapper in prod) is injected, so the single-flight logic is testable without Redis.
2026-06-25 23:14:43 -07:00
Tin Chi Lo
b46733e5b9 feat(mcp): DualCache-backed token cache backend (step 1b §1.5)
The cross-replica TokenCacheBackend implementation that plugs into the foundation's
CachedOAuthTokenStore seam: encrypts+serializes the token via the codec and stores it in LiteLLM's
shared DualCache under the same per-(user,server) key v1 used, so workers share one refresh and a
token cached by v1 or v2 is readable by the other across the cutover. Cache and codec are injected;
a non-positive TTL (already-expired token) is not cached, and a missing/corrupt entry reads as a miss.
2026-06-25 23:11:27 -07:00
Tin Chi Lo
9de7157f5d feat(mcp): encrypt+serialize codec for caching OAuth tokens in Redis (step 1b §1.5)
The serialize+encrypt boundary a cross-replica cache needs: a plaintext bearer in Redis is a leak, so
encode() encrypts (NaCl in prod via the injected encrypt, identity in tests). Caches only access_token
and expires_at, never the refresh_token - the hot path needs just the bearer, and the long-lived
refresh_token stays in the DB (the refresh path is always a cache miss), matching v1. A decoded token
always has refresh_token=None. Undecryptable (key rotation) or corrupt entries read as a miss.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
cec32b4574 feat(mcp): v2-native authorization_code token refresher (step 1b)
The refresh_token grant for the authorization_code mode: POSTs the RFC 6749 refresh_token grant to
the server's token endpoint, persists the rotated triple, and returns the new typed OAuthToken for
RefreshingTokenStore to cache. HTTP post and persist are injected so the grant + response parsing
are testable without a live IdP/DB. Also extends the TokenRefresher seam with (user_id, server_id),
which the foundation's refresh(token) lacked but the grant (server config) and persist (key) need.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
ea1c6135c2 feat(mcp): v2-native per-user token read store (step 1b inner store)
Reads the user's persisted authorization_code credential and returns a typed OAuthToken (access
token, epoch expiry, refresh token), validating the decoded blob at this boundary so no Any leaks
past it. The raw inner store that RefreshingTokenStore/CachedOAuthTokenStore wrap; the DB read +
decode collaborator is injected so it stays testable. Not yet wired - V1PerUserTokenStore is still
the composition-root store until the refresher and cross-worker cache land.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
b9a2589124 fix(mcp): emit the canonical WWW-Authenticate header name in the OAuth challenge
raise_user_oauth_challenge emitted the header lowercase while the sibling raise_public and every
resource_metadata (RFC 9728) emitter use the canonical WWW-Authenticate; align it. HTTP header names
are case-insensitive on the wire so this is cosmetic for compliant clients, but it keeps the challenge
builders consistent and matches RFC 6750.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
d2f9917f02 refactor(mcp): extract the authorization_code arm into a helper
Mirror the api_key arm's structure: the inline AuthorizationCodeConfig body moves into
_authorization_code(subject, server), keeping resolve_credentials a flat one-line-per-arm dispatch.
The helper is annotated with the concrete StaticHeaderAuth it returns rather than the abstract
httpx.Auth (which api_key uses) because a new method carrying the unresolved httpx.Auth return
would add reportUnknownMemberType; the concrete type is both precise and budget-neutral.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
9437ce11fe feat(mcp): route the preemptive 401 existence check through the v2 resolver
The discovery-phase 401 no longer calls v1's _get_user_oauth_extra_headers_from_db to decide
whether a migrated server has a token; it asks the v2 resolver via a new has_user_oauth_token
manager method (to_server_spec + to_subject + resolve_credentials, Ok means a token exists). With
this, every authorization_code resolution runs through the v2 resolver: the call_tool egress, the
listing connection, and the discovery challenge. Delegate servers short-circuit before the check
(the client completes PKCE with the upstream). The challenge itself still emits the RFC 8414
authorization_uri form; the format unification stays a follow-up.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
5ca5e517e8 feat(mcp): cut the tools/list connection over to v2 for authorization_code servers
The listing connection's per-user OAuth header is no longer built by v1 for migrated servers; the
v2 resolver drives it at connect time, ending the double-resolution where v1 built the token into
extra_headers and the v2 graft then deferred to it. Safe because the preemptive 401 (in the
streamable-http and SSE handlers) already challenges a missing token before the listing connection
runs, so the connection is only reached with a token present. Non-migrated oauth2 (delegate) and
the rest still build their header on v1. With this, resolve_credentials' result is honored on every
authorization_code upstream path: tool calls and listing.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
f0cd3b884a feat(mcp): cut the call_tool egress over to v2 for authorization_code servers
_resolve_oauth2_headers_for_tool_call steps aside (builds no header) when to_server_spec maps the
server, so the v2 resolver drives the token-present case instead of being shadowed by a token v1
places in extra_headers. Non-migrated oauth2 (delegate, client_credentials) and BYOK still build
their header on v1. With this, v2 owns the authorization_code egress end to end: inject the
refreshed per-user token when present, raise the per-server fail-closed 401 when absent.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
40f4a00282 feat(mcp): per-server fail-closed OAuth challenge at the v2 egress
When an authorization_code server has no usable per-user token, the arm returns a semantic
unauthorized and the graft builds the 401 where the full MCPServer is in hand: a relative,
per-server RFC 9728 resource_metadata pointer (/.well-known/oauth-protected-resource/mcp/{name})
that names the server's own authorization server, instead of the resolver's earlier root pointer
which resolved to the gateway's generic PRM. Relative, so it is correct behind a reverse proxy
without request context. The listing-phase 401 still emits the RFC 8414 authorization_uri form;
both now target the same server, so the remaining difference is cosmetic and unifies in a later PR.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
8e8a5d4f1e feat(mcp): route oauth2 per-user (authorization_code) servers through the v2 resolver
to_server_spec maps an oauth2 server to AuthorizationCodeConfig when it relies on per-user tokens
(needs_user_oauth_token and not delegate_auth_to_upstream); client_credentials (M2M), delegated
upstream OAuth, token exchange, and SigV4 still defer to v1. The manager injects V1PerUserTokenStore
(resolving through v1's shared egress core) into the credential provider. The v2 path is live but
still defers to a token v1 places in extra_headers; the cutover that makes v1 step aside lands next,
alongside the unified challenge.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
13b1dc18fb refactor(mcp): share v1's OAuth egress core; make V1PerUserTokenStore refresh-capable
Extract v1's per-user OAuth egress (Redis cache, else DB read with the refresh_token grant, then
re-cache) from _get_user_oauth_extra_headers_from_db into resolve_user_oauth_access_token in db.py;
the v1 header builder is now a thin wrapper over it and its callers are unchanged.
V1PerUserTokenStore (the v2 OAuthTokenStore adapter) resolves through that same core via an injected
server lookup, so the authorization_code arm injects exactly the token v1 would, with the same silent
refresh, rather than a Redis-only read that can never refresh. One resolution implementation, two thin
adapters (header dict and OAuthToken). Behavior-preserving: the existing v1 egress tests pass
unchanged, and the arm is not wired into the live path yet (that lands with to_server_spec + the
manager).
2026-06-25 22:19:50 -07:00
Tin Chi Lo
f406b3a626 style(mcp): modern type annotations in the authorization_code arm and source 2026-06-25 22:19:50 -07:00
Tin Chi Lo
6710a8e49c feat(mcp): v1-backed OAuth token source for authorization_code
V1PerUserTokenStore reads the user's stored access token through v1's mcp_per_user_token_cache
(Redis-backed, encrypted) and wraps it in an OAuthToken. v1 holds only the access token (its
cache TTL is the lifetime), so no expires_at/refresh_token yet; the v2 cache holds it for its
default TTL and the OAuth challenge drives re-auth once v1's cache drops it. Additive: nothing
wires it yet, so no behavior change. Step 1b swaps it for a v2-native token store behind the
OAuthTokenStore seam.
2026-06-25 22:19:50 -07:00
Tin Chi Lo
e284981c95 feat(mcp): implement the authorization_code resolver arm
Resolve a user's authorization_code token through the injected OAuthTokenStore: present ->
Authorization: Bearer <access_token>; absent -> the RFC 9728 WWW-Authenticate OAuth challenge;
store unavailable -> the same challenge (not a 500), since a transient outage is not a definite
absence. UpstreamCredentialProvider gains the oauth_token_store collaborator (fail-closed null
default); per-subject isolation comes from keying the fetch on subject_id. Not live until
to_server_spec maps authorization_code and a v1-backed token source is wired (next steps).
2026-06-25 22:19:50 -07:00
Tin Chi Lo
83df471b05 feat(mcp): inject cache-backend and refresh-coordinator seams (cross-replica token caching)
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
Make CachedOAuthTokenStore's storage and RefreshingTokenStore's single-flight injectable so a
cross-replica deployment can back them with Redis without touching the resolver. The defaults preserve
today's behavior exactly: InMemoryTokenCacheBackend (the bounded per-process dict) and
InProcessRefreshCoordinator (the asyncio single-flight). A distributed deployment injects a shared
DualCache-backed backend and a SET NX PX coordinator. invalidate() is now async (the backend may be).
The cache stores via the backend with a TTL derived from the token's expiry; the coordinator threads a
reread callback for the cross-replica case (losers re-read the persisted token) that the in-process
default ignores.
2026-06-25 22:17:10 -07:00
Tin Chi Lo
eaf7932d95 refactor(mcp): thread user_id/server_id through the TokenRefresher seam
The refresh seam took only the OAuthToken, but a refresher needs the server's
config (token endpoint, client credentials, scopes) to run the grant and the
(user_id, server_id) key to persist the minted token, neither of which is
derivable from the token. Widen TokenRefresher.refresh to (user_id, server_id,
token) and pass them through from RefreshingTokenStore so each stacked mode PR
plugs into the final seam rather than forcing a later signature change across
the stack.
2026-06-25 21:43:33 -07:00
Tin Chi Lo
54414ffe29 fix(mcp): default OAuth expiry skew to 60s, the industry standard
The proactive token-refresh / cache-expiry buffer defaulted to 30s, which is
an outlier among OAuth clients. Spring Security uses 60s as both its JWT
clock-skew tolerance and its refresh buffer, and 60s sits inside RFC 7519's
"a few minutes" leeway while preserving nearly all of a typical token's life;
30s was untested, so pin the default with two boundary-probe regression tests.
2026-06-25 17:22:38 -07:00
Tin Chi Lo
92b2b8c26e refactor(mcp): cache positive tokens only, matching v1 (no negative caching)
CachedOAuthTokenStore no longer caches the "not authorized" None result; every miss re-reads the
inner store. v1's per-user token cache never caches misses, so a token written by the OAuth flow
is visible on the next request without an invalidation hook, and uniformly across replicas since
the in-process cache holds no stale None to clear. invalidate() now only covers rotation or
revocation of a cached token. Negative caching (with distributed invalidation) can return later
if a slow DB-backed v2-native source makes per-miss reads expensive.
2026-06-25 17:22:37 -07:00
Tin Chi Lo
6ed9ecfaa7 refactor(mcp): FIFO cache eviction, fix stale single-flight comment + refresh_token docstring 2026-06-25 17:22:37 -07:00
Tin Chi Lo
701f79604d style(mcp): modern type annotations (dict/tuple/X | None) + sorted imports in the token modules 2026-06-25 17:22:37 -07:00
Tin Chi Lo
fd2a3003b5 feat(mcp): proactive token refresh with self-cleaning single-flight
Add TokenRefresher (a mode-supplied seam: mint a fresh token from an expired one and persist it)
and RefreshingTokenStore: when the stored token is near expiry, the first caller refreshes while
concurrent callers await the same in-flight task and share its result, so the IdP is not
stampeded. The task self-cleans (a done-callback drops its entry), so the map is bounded by
in-flight refreshes rather than by distinct users/servers, and is detached from the caller so a
cancelled caller does not abort the refresh. An expired token the refresher cannot renew surfaces
as None so the arm challenges, never a stale bearer; it composes under CachedOAuthTokenStore.
OAuthToken's repr masks the access/refresh tokens so a stray log cannot leak them. Cross-replica
single-flight (Redis) and reactive-401 refresh are the later distributed hardening.
2026-06-25 17:22:37 -07:00
Tin Chi Lo
1d417a8cae feat(mcp): OAuth token store seam + expiry-aware cache for authorization_code
Lay the foundation for the authorization_code resolver arm: OAuthToken (access_token,
expires_at, refresh_token), the OAuthTokenStore Protocol seam, TokenStoreUnavailable for
outages, and CachedOAuthTokenStore, an expiry-aware cache that serves a token only while
unexpired, caches the "not authorized" None for a default TTL, and propagates a store outage
without caching it. Mirrors the BYOK store/cache pattern, adapted for tokens. Refresh and
distributed single-flight are deferred to the hardening step.
2026-06-25 17:22:37 -07:00
Tin Chi Lo
a42fb2fd11 fix(mcp): make Unauthorized a frozen dataclass to keep the type budget flat
CredError's unauthorized payload was a pydantic BaseModel, whose base resolves as unknown in
this repo's basedpyright (every model in the file trips reportUntypedBaseClass plus an unknown
model_config), so the tagged-union case read as unknown and the public edge's challenge access
added reportUnknownMemberType errors over the per-rule ceiling. A frozen dataclass is fully
typed here, so error.unauthorized resolves directly with no cast or accessor and the per-rule
basedpyright counts match base.
2026-06-25 17:22:37 -07:00
Tin Chi Lo
d8b3d6d345 feat(mcp): let CredError.of_unauthorized carry a 401 challenge
The unauthorized case becomes a structured Unauthorized (detail + optional WWW-Authenticate
header + optional structured body) instead of a bare string, and raise_public emits the header
and body when present. This lets a mode reproduce a rich 401 challenge (e.g. BYOK's
provisioning prompt) through the generic resolver edge. of_unauthorized's new params are
keyword-only and default to None, so existing callers and the summary string are unchanged.
2026-06-25 17:22:37 -07:00
Mateo Wang
c5833a9d70
fix: inverted rule in CLAUDE.md (#31370) 2026-06-25 17:00:12 -07:00
Mateo Wang
e0e920d80e
feat(mistral): support Mistral OCR 4 (mistral-ocr-4-0) (#31353)
* feat(mistral): support Mistral OCR 4 (mistral-ocr-4-0)

Add the mistral/mistral-ocr-4-0 model to the cost map and reprice
mistral/mistral-ocr-latest, which now resolves to OCR 4 server-side,
at $4 / 1000 pages. Add the include_blocks param so callers can request
OCR 4's paragraph-level bounding boxes and typed content blocks.

OCR 4's new per-page response fields (blocks, confidence_scores, tables,
hyperlinks, header, footer) already pass through transform_ocr_response
via the extra="allow" config on OCRPage; add a regression test pinning
that behavior alongside cost and param coverage.

* fix(mistral): revert unverified OCR 4 annotation_cost_per_page bump

Mistral's published OCR 4 pricing lists $4/1000 pages for the API and no
separate annotation rate; the $5/1000 figure is the distinct Document AI
(Studio) tier. The earlier 0.003 -> 0.005 bump on annotation_cost_per_page
had no cited source, and ocr_cost() never reads that field (it bills off
ocr_cost_per_page), so the value is documentation-only.

Revert annotation_cost_per_page to the existing 0.003 convention for both
mistral-ocr-latest and mistral-ocr-4-0, keeping only the verified, tested
ocr_cost_per_page: 0.004 change.

* fix(mistral): set OCR 4 annotation_cost_per_page to verified $5/1000 rate

Verified against Mistral's authoritative sources: the pricing page, the
OCR 4 announcement, and the ocr-4-0 model card all list OCR 4 at $4/1000
pages for basic OCR and $5/1000 for annotated pages (Document AI). The
$5/1000 figure is the annotated-pages rate, which is exactly what
annotation_cost_per_page encodes, mirroring the original OCR entry's
0.001 basic / 0.003 annotated split.

Restore annotation_cost_per_page to 0.005 for mistral-ocr-latest and
mistral-ocr-4-0; the earlier revert to 0.003 was based on an incomplete
reading that treated Document AI as a separate product. ocr_cost_per_page
stays 0.004, which is the value billed by ocr_cost().

* fix(mistral-rust): include_blocks in Rust OCR supported params

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-06-25 16:42:37 -07:00
Mateo Wang
b7f28bd89f
feat(aiml): add openai/gpt-image-2 image model (#31323)
* feat(aiml): add openai/gpt-image-2 image model

Adds aiml/openai/gpt-image-2 to the cost map and teaches AimlImageGenerationConfig
to route OpenAI-style image models through the upstream OpenAI request schema
instead of the AI/ML flux schema. Without this, size, n, and response_format would
be remapped to image_size/num_images/output_format, which the gpt-image-2 endpoint
on api.aimlapi.com does not accept.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

* chore(aiml): note gpt-image-2 flat-rate pricing basis; apply ruff format

Documents in the cost-map notes that output_cost_per_image is AI/ML's
published medium-quality rate, billed as a flat per-image price like the
other aiml image entries. Reformats the touched files under the repo's
ruff formatter (migrated from black in #31317).

* fix(aiml): drop /v1/images/edits from gpt-image-2 supported_endpoints

LiteLLM only implements an image generation transformer for AIML, so
listing /v1/images/edits overclaimed support. Align with every other
aiml image entry, which lists only /v1/images/generations.

* style(aiml): format transformation.py at line-length 88

The repo formats litellm/ with ruff at line-length 88 (Makefile/CI call
sites), while ruff.toml's global 120 only governs E501/import sorting.
Reformat the transformer to 88 so make format-check / CI lint pass, and
restore the test files to their original layout since tests/ is not part
of the auto-formatted tree.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-06-25 16:41:43 -07:00
yucheng-berri
e64cec5add
ci(image-scan): add Grype image scan for OS + library CVEs (#31151)
* ci(image-scan): add Grype image scan for OS + library CVEs

Builds each of the 6 Dockerfiles via a matrix and scans the resulting image
with Grype (pinned v0.114.0, sha256 verified), failing on fixable HIGH or
CRITICAL across both OS/apk and language packages. This catches the layer
osv-scan is structurally blind to (Wolfi/apk OS packages and vendored deps
like prisma's node engine), which is the structural reason the openssl CVE
slipped past CI and a customer's image scanner flagged it.

Skipped on fork PRs so an outside contributor cannot run arbitrary code on
our hosted runner via a malicious Dockerfile RUN line. The same pattern is
used by guard-fork-dependencies.yml.

Grype runs as a pinned binary with a verified checksum, so there is no
mutable-tag GitHub Action in the dependency chain and no vendor credentials
in the scan job. The job uses read-only contents permissions and an empty
top-level permissions block.

* ci(image-scan): scan only Dockerfile.non_root (rootless target)

All Dockerfile variants share the same wolfi base and apk set today, so a single scan of Dockerfile.non_root gives the same OS-layer coverage at one-sixth the build cost. Dockerfile.non_root is the rootless variant we ship (USER 65534), so the scan tracks the image customers actually run. Matrix-scan if the variants ever diverge.

* ci: retrigger checks (proxy_pass_through_endpoint_tests flaked on prior run)
2026-06-25 16:35:50 -07:00
Mateo Wang
0a92734691
fix: clarify further that customer names shouldn't be made public (#31365)
* fix: make it clearer that customer names should not be put in PR descriptions

* fix: revise the policy to be stricter
2026-06-25 16:27:27 -07:00
milan-berri
7ffce15766
Add GA pricing for gemini-3-pro-image and gemini-3.1-flash-image. (#30022)
Fixes #29794. Adds bare, gemini/, and vertex_ai/ entries copied from preview models so proxy cost tracking works for GA model names.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 00:40:54 +02:00
Yassin Kortam
01035499da
fix(cache): apply Redis namespace to all key operations (#31288)
The namespace configured under cache_params was only applied to get/set/
increment paths. Operations that take keys through other code paths (the Lua
scripts registered via async_register_script, delete, scan_iter, rpush, lpop,
get_ttl, and the sync increment_cache) hit raw keys. With a namespace set, the
rate limiter ({key}:tokens/requests/window), pod-lock release, and budget
limiters wrote keys outside the configured prefix, breaking multi-tenant key
isolation and leaving those operations reading keys the namespaced writes never
created.

check_and_fix_namespace is now applied uniformly across every key-taking
RedisCache operation. It is a no-op when no namespace is configured, so
deployments without a namespace are unaffected. The prefix is prepended ahead of
any {hash-tag}, so Redis Cluster slotting is preserved.

Resolves LIT-3374
2026-06-25 15:39:07 -07:00
ishaan-berri
62f93a3343
feat: add Rust OCR providers (#31272)
* feat: port OCR providers to Rust gateway

* chore(deps): update langgraph checkpoint lock

* ci: scope ruff format check to changed files

* ci: fix OCR lint and patch coverage

* fix(ocr): block mapped IPv6 fetch targets

* test(ocr): include rust bridge coverage in OCR shard

* ci: rerun responses shard
2026-06-25 15:12:30 -07:00
Mateo Wang
92d0788da2
chore(lint): widen ANN slack to 10% of baseline and drop PLR0913 from the strict gate (#31335)
* chore(lint): widen ruff budget slack to 10% of baseline for high-volume ANN rules and PLR0913

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

* chore(lint): drop PLR0913 from strict gate to roll out rules gradually

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

* fix(lint): ratchet-guard rising baselines even when slack is cut to mask them

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-06-25 14:43:45 -07:00
ishaan-berri
a2d04ccdbb
ci: harden cargo fetches during maturin builds (#31348) 2026-06-25 14:31:05 -07:00
Mateo Wang
f98e935504
chore: gitignore rust bridge build artifacts (#31349)
Ignore the compiled, platform-specific Rust extension output (litellm/rust_bridge/_native*.so/.pyd) and the litellm-rust/target/ build dir so local maturin/cargo builds don't show up as untracked files.

Also drop the two stale self-referential .gitignore entries; .gitignore is tracked, so ignoring it did nothing except add confusion.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-06-25 14:28:49 -07:00
ishaan-berri
d8ef1da49d
feat: package Rust OCR bridge in LiteLLM wheel (#31267)
* feat: package rust ocr bridge in litellm wheel

* Install Rust in Windows CircleCI job

* Address Rust wheel review feedback

* Pin Windows rustup installer hash
2026-06-25 12:32:55 -07:00
yucheng-berri
a545c493d7
fix(otel): hashable scope for _emit_once when guardrail_mode is list (#31262)
Some checks are pending
LiteLLM Rust / rustfmt, clippy, test (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(otel): hashable scope for _emit_once when guardrail_mode is list

`_emit_once` keys `spans_logged` by `(class, id, *scope)`. When a
guardrail entry's `guardrail_mode` arrives as a `List[GuardrailEventHooks]`
(the shape Presidio expands to with `output_parse_pii: true`, and the
shape `event_hook` carries for any `mode: [...]` in config), the tuple
contains a list and `spans_logged.get(dedupe_key)` raises
`TypeError: unhashable type: 'list'`. On the post-call path this fires
inside the logging callback and is swallowed; the request returns 200 but
the OTEL `guardrail` span is silently dropped. On the blocking path the
same error surfaces as HTTP 500.

Adds `_freeze_for_dedupe`, a small recursive normalizer that turns lists
and tuples into tuples, sets into frozensets, dicts into frozensets of
`(key, value)` pairs, and falls back to `repr` for arbitrary
unhashables. Applied inside `_emit_once` before the dict lookup, so all
three callsites are protected without touching the guardrail-specific
callsite. Helper assumes acyclic input; `guardrail_mode` values are
built fresh from config (str enums, lists of str enums, TypedDict of
str/list-of-str), so no cycle can arise in practice.

Regression tests in `TestOpenTelemetrySpanDedupe` cover the list crash,
distinct-list-scope collision, dict and set scope parts, and an
end-to-end `_create_guardrail_span` exercise that confirms exactly one
`guardrail` span is emitted across repeated lifecycle entrypoints. Each
new test fails on a reverted helper (4/4 mutation kill)

* fix(otel): cap _freeze_for_dedupe recursion depth and ignore in recursive detector

CI's recursive_detector blocks new recursive functions in litellm/ unless they
are in the allowlist with a documented bound. Cap the helper at 16 levels and
return repr(value) past the cap; this is well past the realistic depth of
guardrail_mode (1-3 levels) and means a future caller passing a cyclic
container can no longer push the proxy logging path into a RecursionError.
Add a regression test that exercises the cycle path.

* refactor(otel): annotate _freeze_for_dedupe return as a HashableScope union

Per review feedback from @mateo-berri: replace the loose `-> object` annotation
with a recursive `HashableScope` union (str | int | float | bool | bytes | None
| Tuple[HashableScope, ...] | FrozenSet[HashableScope]) so the helper's contract
is visible at the signature. Replace the `try/except hash(value); return value`
passthrough with an explicit isinstance check over the hashable-scalar types so
the type checker can narrow without requiring `cast(Hashable, value)` on the
return. Symmetric: dict keys also flow through the freezer (a TypedDict key is
already a string in practice, so behaviorally identical). All 16 regression
tests still pass; mutation kill behavior preserved

* fix: avoid explicit casting

---------

Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>
2026-06-25 11:59:35 -07:00
Mateo Wang
17bfd415ae
chore: migrate Python formatter from black to ruff format (#31317) 2026-06-25 11:27:43 -07:00
Mateo Wang
6db55e0aa5
feat(mcp): add mcp_xff_num_trusted_hops to harden X-Forwarded-For client IP resolution (#31257)
* feat(mcp): add mcp_xff_num_trusted_hops to harden XFF client IP resolution

MCP per-server IP access control reads the client IP from X-Forwarded-For
and trusts the leftmost entry. Behind an append-style proxy or load
balancer (AWS ALB, nginx with $proxy_add_x_forwarded_for, HAProxy, Envoy,
Cloudflare), a client can prepend an arbitrary value to the header, so the
leftmost entry is attacker-controllable even when the direct peer is a
trusted proxy. An attacker can therefore spoof an internal IP and reach
servers marked available_on_public_internet=false.

This adds an optional mcp_xff_num_trusted_hops general setting modelled on
Envoy's xff_num_trusted_hops. When set to N, the client IP is read N entries
from the right of the chain (where N is the number of trusted appending
proxies in front of the gateway) instead of the leftmost value, so any
entries a client prepends are ignored. It composes with mcp_trusted_proxy_ranges,
which still validates the direct peer, and only takes effect once that check
passes; without a validated direct peer the gateway keeps failing closed, so
hop counting cannot be abused by a direct-to-pod attacker. The chain must
contain at least N valid entries or resolution fails closed.

Default is unset, preserving existing behaviour.

* chore(ui): regenerate dashboard schema for mcp_xff_num_trusted_hops

* fix(mcp): warn when mcp_xff_num_trusted_hops is below the minimum

A 0 or negative value is silently treated as disabled, which could leave
an operator believing they enabled append-style X-Forwarded-For hardening
while client IP resolution stays on the spoofable leftmost value. Emit a
warning, consistent with how the module already surfaces invalid CIDR
config, so the misconfiguration is visible in logs.

* fix(mcp): reject mcp_xff_num_trusted_hops < 1 at config-parse time

Add a ge=1 bound to the ConfigGeneralSettings field so the
update_config_general_settings path rejects 0 and negative values with a
clear validation error instead of accepting them, and self-documents the
valid range. The runtime warning stays as defense-in-depth for raw-dict
config that bypasses model validation.

* style(mcp): black-format ip_address_utils.py

* fix(mcp): fail closed when mcp_xff_num_trusted_hops is set but invalid

A present-but-invalid mcp_xff_num_trusted_hops (non-integer, or below 1)
previously made _resolve_num_trusted_hops return None, which the caller
treated identically to "unset" and silently fell back to the legacy
leftmost X-Forwarded-For value. An operator who set the value to harden
client IP resolution but typo'd it would get weaker security than before,
with no fail-closed signal.

Model the setting as a tagged union (_HopCountUnset, _HopCountInvalid,
_HopCount) so the three states are distinct: unset keeps the legacy path,
a valid count drives hop-counting, and an invalid value fails closed
(returns "") instead of reverting to the spoofable leftmost address. The
caller matches on the union exhaustively.

Add a parametrized regression test asserting get_mcp_client_ip returns ""
for 0, -1, "abc", and 1.5 even with a spoofed internal leftmost entry,
and update the resolver unit tests for the new return type.
2026-06-25 07:31:29 -07:00
michelligabriele
0a8a87afe0
fix(streaming): word-sliced cache replay for stream=true cache hits (#30216)
* fix(streaming): word-sliced cache replay for stream=true cache hits

* fix(streaming): align mypy and replay happy-path test with word-sliced cache replay

* fix(streaming): short-circuit whitespace-only content in cache replay splitter

* fix(streaming): emit tool_calls/function_call only on first replay slice

* refactor(streaming): drop dead delattr guard in cache replay

A non-None usage on the replay base object always lives in
__pydantic_extra__ (it is attached via setattr earlier in the same
function), so delattr can never raise here; the try/except AttributeError
that silently swallowed a failure was dead defensive code that could only
ever hide a real regression, so it is removed in both the async and sync
generators.

Also switches the new replay annotations from typing.List to the builtin
list to satisfy the strict ruff UP006 gate and drops the unused
PLR0915 noqa directives (the rule is not enabled in this repo's ruff
config, so RUF100 flagged them).

* fix(streaming): drop carried-over metadata from later cache replay slices

The word-sliced cache replay deep-copies the full ModelResponseStream per
slice, so reasoning_content, thinking_blocks, logprobs, enhancements,
annotations and the rest of the per-message metadata rode on every slice, not
just the first. Downstream handlers that accumulate streamed deltas would
collect each one once per slice, e.g. duplicating a cached reasoning trace N
times on a stream=true cache hit.

Later slices are now rebuilt as a content-only delta with choice-level logprobs
and enhancements stripped, so the whole metadata class stays on the first slice.
Adds async (logprobs) and sync (reasoning_content/thinking_blocks/logprobs/
enhancements, plus annotations) regression tests

---------

Co-authored-by: Mateo <277851410+mateo-berri@users.noreply.github.com>
2026-06-25 07:13:05 -07:00
Sameer Kankute
c712c20d0f
fix(ci): point OSS contributor workflows to litellm_oss_staging (#31270)
* fix(ci): point OSS contributor workflows to litellm_oss_staging

Workflow triggers and guard error messages incorrectly referenced litellm_oss_branch; update them to the branch we actually use for external contributions.

* fix(ci): include test-rust.yml in litellm_oss_staging rename

Missed test-rust.yml when updating OSS contributor target branch references.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-24 21:07:59 -07:00
tin-berri
0f5603895c
fix(mcp): challenge delegate-auth OAuth servers with upstream resource_metadata (#31255)
An oauth2 MCP server with delegate_auth_to_upstream=true never prompted the
user to sign in. On an unauthenticated initialize the gateway answered locally
(200, no tools) and emitted no WWW-Authenticate, so clients like Claude Desktop
either connected empty or hit "OAuth probe timeout after 10000ms".

#30124 added a bare `continue` in _raise_preemptive_401_for_unauthenticated_servers
to stop sending LiteLLM's gateway authorization_uri challenge for delegate-auth
servers, expecting the upstream to emit its own challenge. On initialize the
gateway never probes upstream, so no challenge ever reached the client.

Replace the `continue` with a preemptive 401 carrying the proxied
resource_metadata (RFC 9728) challenge, the same form passthrough servers and
MCPUpstreamAuthError already use. This keeps #29770 fixed (still no
authorization_uri) while restoring the upstream PKCE sign-in prompt.
2026-06-24 20:50:21 -07:00
tin-berri
f426912ba1
fix(mcp): resolve toolset tools by the server's known prefix (#31254)
* fix(mcp): resolve toolset tools by the server's known prefix

Toolsets store {server_id, bare tool_name} and reconcile that against the
live prefixed tool name at list time. The reconciliation chopped the live
name at the first MCP_TOOL_PREFIX_SEPARATOR with no server context, so a
server whose prefix contains the separator (a hyphenated alias, or the
UUID server_id used as the prefix when a server has no alias) had its
tools silently dropped from /toolset/<name>/mcp while listing fine
everywhere else. Strip the exact known prefix for the tool's server_id
instead of guessing the boundary, on both the resolve and filter sides

Also render toolset tools as {server-prefix}-{tool} in the dashboard
picker result and chips; this is display only, the persisted record
stays {server_id, bare tool_name}

Resolves LIT-3419

* test(mcp): add focused unit tests for strip_known_server_prefix

Cover the LIT-3419 cases directly on the helper with real MCPServer
objects: clean prefix round-trip, hyphenated alias, UUID server_id
fallback, unprefixed passthrough, and the server=None legacy fallback
2026-06-24 20:50:16 -07:00
Mateo Wang
9c41077786
fix(mcp): warn loudly when X-Forwarded-For is present but use_x_forwarded_for is off (#31266)
* fix(mcp): warn loudly when X-Forwarded-For is present but use_x_forwarded_for is off

When a request carries an X-Forwarded-For header but use_x_forwarded_for is
unset, get_mcp_client_ip silently falls back to the direct peer's IP (the load
balancer / reverse proxy). That peer almost always sits inside
mcp_internal_ip_ranges, so the 'Internal network only'
(available_on_public_internet: false) restriction trusts every external caller
as internal and effectively exposes those servers.

Emit a one-shot loud error pointing the operator at use_x_forwarded_for instead
of hard-failing: on a deployment with no load balancer, a crafted
X-Forwarded-For header must not be able to take the service down, and a one-shot
log keeps a flood of crafted headers from spamming the logs.

* fix(mcp): re-arm XFF-disabled warning on config change and harden test assertion

Address PR review: tie the one-shot warning flag to the observed
use_x_forwarded_for value so it re-arms whenever the setting is seen enabled,
restoring the diagnostic on a later rollback to disabled. Also assert against
str(call_args) so the test survives a positional-to-keyword logger refactor.
2026-06-24 20:49:32 -07:00
Mateo Wang
257d67167f
fix(mcp): correct misleading no-trusted-proxy warning for XFF access control (#31264)
* fix(mcp): correct misleading no-trusted-proxy warning for XFF access control

* test(mcp): assert the no-trusted-ranges warning was logged instead of relying on StopIteration
2026-06-24 20:49:29 -07:00