litellm/tests/test_litellm/proxy/_experimental/mcp_server
tin-berri 5963b9320f
feat(mcp): cross-replica single-flight refresh for the v2 per-user OAuth store [2/2] (#31493)
* feat(mcp): encrypt+serialize codec for caching OAuth tokens in Redis (step 1b §1.5)

The serialize+encrypt boundary a cross-replica cache needs: a plaintext bearer in Redis is a leak, so
encode() encrypts (NaCl in prod via the injected encrypt, identity in tests). Caches only access_token
and expires_at, never the refresh_token - the hot path needs just the bearer, and the long-lived
refresh_token stays in the DB (the refresh path is always a cache miss), matching v1. A decoded token
always has refresh_token=None. Undecryptable (key rotation) or corrupt entries read as a miss.

* feat(mcp): DualCache-backed token cache backend (step 1b §1.5)

The cross-replica TokenCacheBackend implementation that plugs into the foundation's
CachedOAuthTokenStore seam: encrypts+serializes the token via the codec and stores it in LiteLLM's
shared DualCache under the same per-(user,server) key v1 used, so workers share one refresh and a
token cached by v1 or v2 is readable by the other across the cutover. Cache and codec are injected;
a non-positive TTL (already-expired token) is not cached, and a missing/corrupt entry reads as a miss.

* feat(mcp): Redis SET NX PX refresh coordinator (step 1b §1.5)

The cross-replica RefreshCoordinator that plugs into the foundation's RefreshingTokenStore seam: a SET
NX PX lock elects one worker to refresh per (user, server) while the rest wait for it and re-read the
token it persisted, so a rotating refresh_token is used once across the fleet, not once per worker. The
lock self-expires (PX) so a crashed holder can't wedge refresh; a loser falls back to a bounded re-read
and the surrounding store re-checks expiry next fetch, so a crash self-heals. The lock (a thin Redis
SET NX/DEL/EXISTS wrapper in prod) is injected, so the single-flight logic is testable without Redis.

* feat(mcp): Redis SET NX PX distributed lock (step 1b §1.5)

The concrete DistributedLock the RedisRefreshCoordinator elects refreshers with: acquire is an atomic
SET key NX PX ttl (first caller wins, entry self-expires so a crashed holder can't wedge refresh),
release is DEL, is_held is EXISTS. The async Redis client is injected (the client from LiteLLM's
RedisCache in prod), so it is unit-testable with a fake. Any Redis error degrades to not-acquired /
not-held so a cache blip causes an extra refresh, never a crash on the resolve path.

* feat(mcp): wire the cross-replica cache + coordinator into the per-user store (step 1b §1.5)

Upgrade the composition root to use the DualCache-backed cache and SET NX PX refresh
coordinator when Redis is wired, falling back to the foundation's in-process defaults on a
single replica. Layers the cross-replica path on top of the single-replica dispatch store.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): refresh on lock-backend error instead of serving a stale token

The cross-replica refresh coordinator elected refreshers with a boolean acquire:
a Redis transport error was caught and returned as False, which is
indistinguishable from "another worker holds the lock". On a total Redis
outage every worker therefore took the wait-then-reread branch and served the
still-expired token upstream (the upstream then 401s), even though the lock and
coordinator docstrings claimed a Redis blip "degrades to an extra refresh".

Make acquire tristate (LockAcquisition: ACQUIRED / HELD / ERROR) so the
coordinator can tell a busy holder from a dead backend, and refresh anyway on
ERROR. This single-flight lock is a load optimization, not a correctness mutex,
so failing open is correct: it degrades a lock-backend outage to the
no-coordinator behavior (an extra refresh), never a stale bearer.

Add a regression test asserting an acquire error refreshes rather than
re-reading the expired token, and update the docstrings to match.

* style(mcp): wrap redis lock signatures at line-length 88 for CI ruff format

* fix(mcp): a refresh loser surfaces None, not a stale token, when the winner failed

The cross-replica coordinator's losers re-read the token the winner persisted.
If the winner's refresh failed, the store still holds the expired token, so the
loser re-read it and RefreshingTokenStore handed that expired bearer to the
caller (the upstream then 401s) instead of the re-auth challenge the winner
returned via None.

Make the loser's re-read expiry-aware, mirroring refresh_latest_token: a
re-read that is still expired surfaces None so the arm challenges. This only
affects the loser path; the winner's freshly refreshed token is returned
directly by the coordinator and is unaffected.

* fix(mcp): log per-user token decrypt failures at debug, matching v1

When a cached blob cannot be decrypted (e.g. after a salt or master-key rotation) the codec logged a full traceback at error level, since decrypt_value_helper defaults to exception_type=error. v1's MCPPerUserTokenCache passed exception_type=debug on the same path. The blob is ciphertext so this is log noise only, but matching v1 avoids error-level traceback spam on stale entries after a key rotation

* fix(mcp): namespace the refresh lock key and fence its release with a token

The Redis lock wrote its key through the raw client from init_async_client(), bypassing RedisCache's namespace, so two deployments sharing one Redis collided on mcp:refresh_lock:<user>:<server> for any overlapping (user, server) and a colliding deployment skipped the refresh and challenged its own users. The lock now runs every key through an injected namespace_key wired to RedisCache.check_and_fix_namespace, matching the namespace its token cache already uses

release() also deleted the key unconditionally, so a holder whose lock PX-expired and was re-acquired by another worker could delete the new holder's lock and let a third worker run a duplicate refresh, recreating the rotating refresh_token race. acquire now writes a unique per-acquisition token generated by the coordinator and release deletes only when the key still holds that token, via a compare-and-delete Lua script

Adds regression tests: release with a stale token is a no-op while the owner's release deletes; keys are namespaced before reaching Redis; the coordinator acquires and releases with the same token

* fix(mcp): fail open when the per-user token cache delete errors

DualCache swallows get/set errors internally but not delete, and the Redis
layer underneath re-raises through its circuit breaker. So a Redis outage on
the delete() path escaped CachedOAuthTokenStore.fetch()'s unauthorized branch
(which deletes before returning None) and invalidate(), turning a cache blip
into a 500 instead of the v1-style fallback. Catch in the backend so delete
degrades to the TTL-bounded stale entry like get/set already do.

* style(mcp): reformat outbound-credentials files to line-length 120

The merge from staging brought in ruff's line-length 120, but these two
PR-authored files were still wrapped at the old width, so the diff-scoped
ruff format --check in CI flagged them. Pure reformatting; no behavior change.

* fix: harden mcp oauth redis refresh coordination

* fix(mcp): make the per-user token cache backend airtight on boundary failures

get/set now degrade a cache or codec failure to the safe value (miss / no-op)
in the backend itself rather than relying on DualCache and decrypt_value_helper
happening to swallow internally, matching delete() and v1's MCPPerUserTokenCache.
This upholds the layer's boundary-failure-is-a-miss contract regardless of the
injected collaborators, so a Redis outage or an undecryptable entry reads as a
cache miss that re-reads the DB instead of a 500. Adds contract tests for the
cache raising on get/set/delete and the codec raising on encode.

* test(mcp): pin per-user cache get() to a miss when decrypt raises

Greptile's out-of-diff repro had the decrypt reject a blob with ValueError
(bad ciphertext after key rotation); cover that exact raise path, not just the
decrypt-returns-None case, so get() is regression-locked to read it as a miss.

* refactor(mcp): use frozen dataclasses for the trivial DI constructors

Replace the hand-written self._<arg> = arg constructors on OAuthTokenCacheCodec,
RedisRefreshCoordinator, RedisDistributedLock, and DualCacheTokenCacheBackend
with frozen slotted dataclasses, matching the rest of this layer. Fields take the
former parameter names so the constructor API (and the tests' keyword args) are
unchanged; KW_ONLY preserves the keyword-only collaborators.

* fix: serialize lazy per-user oauth store rebuild

* fix(mcp): stop losers challenging mid-refresh by decoupling wait from lease TTL

wait_timeout_seconds defaulted to the same 10s as lock_ttl_seconds, but the
holder renews its lease while a slow token endpoint runs, so a loser waiting
past 10s bailed and re-read the still-expired DB token, challenging the user
even though a valid refresh was in flight. Bound the holder's renewal with a
refresh budget so its lock-hold is finite, and set the loser's wait to outlast
that budget (refresh_budget_seconds + one lease tail) so a loser only re-reads
once the holder has finished or its bounded lease has lapsed, never mid-refresh.

* fix: allow concurrent lazy OAuth fetches without Redis

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-06-27 16:27:34 -07:00
..
auth feat(mcp): opt-in least-privilege default for team key MCP access (#31380) 2026-06-25 18:49:15 -07:00
guardrail_translation fix(tests): Add missing mocks for MCP IP filtering and updated APIs (#20652) 2026-02-07 11:30:49 -08:00
outbound_credentials feat(mcp): cross-replica single-flight refresh for the v2 per-user OAuth store [2/2] (#31493) 2026-06-27 16:27:34 -07:00
test_byok_oauth_endpoints.py feat(mcp): allow native MCP OAuth support for cursor (#28327) 2026-05-20 15:28:44 -07:00
test_callback_oauth_error_responses.py Litellm oss staging 250526 (#28770) 2026-05-26 11:57:39 -07:00
test_db_credentials.py fix(ui): persist Tools-tab MCP OAuth token to DB (#29809) 2026-06-05 22:29:56 -07:00
test_discoverable_endpoints.py changing expires_in default to use actual slack return details (#29951) 2026-06-08 18:13:06 -07:00
test_is_tool_name_prefixed.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_jwt_mcp_enforcement.py fix(mcp): resolve team.access_group_ids → MCP servers (#28997) 2026-05-27 12:36:50 -07:00
test_jwt_mcp_simple.py fix(mcp): resolve team.access_group_ids → MCP servers (#28997) 2026-05-27 12:36:50 -07:00
test_mcp_cost_calculator.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_mcp_custom_fields.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_mcp_debug.py chore: litellm oss staging (#30968) 2026-06-23 07:31:44 -07:00
test_mcp_discovery.py fix(mcp): default Linear MCP registry entry to streamable HTTP (#30396) 2026-06-13 14:45:47 -07:00
test_mcp_elicitation_handler.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_env_vars.py fix(mcp): drop orphaned per-user credential rows when an MCP server is deleted (#30141) 2026-06-10 15:56:58 -07:00
test_mcp_header_alias_utils.py feat(mcp): Add tool call and tool list support via UI for Oauth mcps (#28454) 2026-05-22 09:04:04 -07:00
test_mcp_hook_extra_headers.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_metadata_preservation.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_mcp_oauth_passthrough.py [internal copy of #28008] Support MCP OAuth passthrough and issuer-scoped JWT auth (#28356) 2026-06-02 12:22:04 -07:00
test_mcp_oauth_passthrough_cold_start.py [internal copy of #28008] Support MCP OAuth passthrough and issuer-scoped JWT auth (#28356) 2026-06-02 12:22:04 -07:00
test_mcp_oauth_passthrough_tools.py feat: litellm oss 110626 (#30202) 2026-06-11 22:30:26 -07:00
test_mcp_partial_update.py fix(mcp): clear allowed_tools and tool overrides on MCP server edit (#29411) 2026-06-01 21:28:29 -07:00
test_mcp_sampling_completion_flow.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_model_access.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_model_resolution.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_priority_selection.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_request_builder.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_response_conversion.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_sampling_tool_conversion.py Litellm oss staging 040626 (#29671) 2026-06-04 11:07:20 -07:00
test_mcp_server.py feat(mcp): migrate authorization_code MCP to the v2 resolver (single-replica) [1/2] (#31473) 2026-06-26 21:19:57 -07:00
test_mcp_server_identity_env.py chore: litellm oss 170626 (#30637) 2026-06-17 21:11:12 -07:00
test_mcp_server_manager.py feat(mcp): migrate authorization_code MCP to the v2 resolver (single-replica) [1/2] (#31473) 2026-06-26 21:19:57 -07:00
test_mcp_session_logging.py Add MCP semantic conventions to otelv2 (#29468) 2026-06-02 11:45:36 -07:00
test_mcp_sigv4_auth.py feat(mcp): per-server env vars with global + per-user scopes (#28917) 2026-06-05 20:15:11 -07:00
test_mcp_stale_session.py feat(mcp): migrate authorization_code MCP to the v2 resolver (single-replica) [1/2] (#31473) 2026-06-26 21:19:57 -07:00
test_mcp_toolset_scope.py fix(mcp): resolve toolset tools by the server's known prefix (#31254) 2026-06-24 20:50:16 -07:00
test_oauth2_token_cache.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_openapi_to_mcp_generator.py fix(mcp): forward extra_headers for OpenAPI MCP tools (#27383) 2026-05-09 15:10:54 -04:00
test_openapi_tool_auth.py fix(mcp): use canonical proxy_logging_obj, deny when MCP server is unresolvable 2026-05-01 22:28:46 +00:00
test_rest_endpoints.py feat(mcp): migrate authorization_code MCP to the v2 resolver (single-replica) [1/2] (#31473) 2026-06-26 21:19:57 -07:00
test_semantic_tool_filter.py chore: litellm oss staging (#30968) 2026-06-23 07:31:44 -07:00
test_short_mcp_tool_prefix.py fix(mcp): resolve toolset tools by the server's known prefix (#31254) 2026-06-24 20:50:16 -07:00
test_ui_session_utils.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00