Previously returned the full auth header value (including scheme prefix
like 'token abc') when no 'Bearer ' prefix was found, causing that
verbatim value to be sent to the IDP as the subject_token. Now returns
None, which correctly skips OBO token exchange for non-Bearer callers.
Per-user lock entries are now eligible for GC once no coroutine holds
a reference, so rotating JWTs across many users don't accumulate entries
indefinitely. Also moved IDP error body to debug log only to avoid
leaking partial token values or internal metadata into server error replies.
Example YAML config for testing the oauth2_token_exchange auth type
with mock MCP and token exchange servers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds the oauth2_token_exchange case to _get_auth_headers() so the
exchanged token is sent as a Bearer token to the MCP server.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds _extract_bearer_token() helper, updates _create_mcp_client() and
_call_regular_mcp_tool() to extract the user's JWT and pass it as the
subject token for OBO exchange. Also updates load_servers_from_config()
and build_mcp_server_from_table() to read token exchange config fields.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Updates resolve_mcp_auth() to accept a subject_token parameter and
route through TokenExchangeHandler when the server is configured for
oauth2_token_exchange.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
New module that handles OAuth 2.0 Token Exchange (RFC 8693). Exchanges
a user's JWT for a scoped token at the configured IDP endpoint, with
per-user per-server caching and async lock protection against duplicate
fetches.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sets cache max size to 500 for token exchange (higher than
client_credentials since it's per-user).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds `token_exchange_endpoint`, `audience`, and `subject_token_type`
fields plus `has_token_exchange_config` property to MCPServer for
determining when OBO token exchange should be used.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds the new `oauth2_token_exchange` value to MCPAuth enum and MCPAuthType
literal, plus `audience`, `token_exchange_endpoint`, and `subject_token_type`
fields to MCPCredentials TypedDict to support RFC 8693 token exchange.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Route publisher/model ids (e.g. xai/grok) to .../endpoints/openapi; keep model in JSON body
- Add model_prices keys for vertex_ai/openai/xai/grok-*
- Document xAI Grok on vertex_partner (aligned with GPT-OSS)
- Add tests for create_vertex_url and body-model heuristic
Made-with: Cursor
Success handlers already run when CustomStreamWrapper or
CachedResponsesAPIStreamingIterator finishes replay. Logging at
cache-hit time for acompletion/completion streaming duplicated spend
and callbacks. Align tests with deferred behavior.
Made-with: Cursor
_base_process_chunk only encodes response IDs when parsed_chunk contains
a top-level "response" key. Align test_process_chunk_completed_response_
updates_id_and_usage_cost with that contract and test_base_responses_api_streaming_iterator.
Made-with: Cursor
The test supplies a minimal PDF base64 payload but expected the wrong
constant (base64 for "test"). Assert against the same pdf_b64 value
and drop the unused import.
Made-with: Cursor
* fix(bedrock): handle document content blocks in Converse API message conversion
Document content blocks (used for PDF support) were silently dropped
during message conversion for Bedrock's Converse API. The content block
processing loop only handled text, image_url, and file types — document
blocks were skipped without warning, causing the model to respond as if
no document was provided.
Adds document block handling in three locations:
- Sync user message processing (_bedrock_converse_messages_pt)
- Async user message processing (_bedrock_converse_messages_pt_async)
- Tool result conversion (_convert_to_bedrock_tool_call_result)
Fixes#24641
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: use _validate_format for proper MIME type to Bedrock format mapping
Address Greptile review: naive media_type.split("/")[1] produced invalid
Bedrock format names for complex MIME types (e.g. OOXML → docx, text/plain
→ txt, text/markdown → md). Now reuses BedrockImageProcessor._validate_format
which handles all MIME types correctly via mimetypes + fallback.
Also fixes test assertions to expect correct Bedrock format values and adds
text/plain and text/markdown test cases.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: reject non-base64 document sources with a clear error
URL-type document sources (e.g. {"type": "url", "url": "..."}) would
crash with an opaque KeyError on missing 'media_type'. Guard at the top
of _process_document_message and raise a clear ValueError since Bedrock
Converse only supports base64-encoded document sources.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The /v2/model/info endpoint (used by the UI's Models + Endpoints page)
was not resolving access group names when filtering models by team.
When a team has models: ["Group-A"] where "Group-A" is an access group,
_filter_models_by_team_id() passed it as a literal model name to
get_model_list(), which found no deployments with that name. This caused
the UI to show all models instead of only team-accessible ones.
The request-time auth path (model_in_access_group in auth_checks.py)
correctly resolves access groups via get_model_access_groups(). This
fix applies the same resolution in _filter_models_by_team_id() for both
the in-memory router lookup and the database fallback query.
Tests added:
- test_filter_resolves_access_group_names
- test_filter_resolves_mix_of_access_groups_and_literal_names
- test_filter_excludes_models_from_other_access_group
- test_filter_db_fallback_receives_resolved_model_names
When model_info.id equals model_name (common for batch models), the router
resolves via has_model_id and returns one deployment dict instead of a list.
The dict branch incorrectly iterated deployment keys (model_name,
litellm_params, model_info), producing non-string values that broke
LiteLLM_ManagedFileTable validation on managed file upload.
Normalize list vs dict by wrapping single deployments and extracting
model_info.id for each response pair.
Add regression tests including the batch model id == model_name case.
Made-with: Cursor