* fix(mcp): Add standard MCP URL pattern support for OAuth discovery (#17272)
OAuth discovery endpoints now support both URL patterns:
- Standard MCP pattern: /mcp/{server_name} (new)
- Legacy LiteLLM pattern: /{server_name}/mcp (backward compatible)
The standard pattern is required by MCP-compliant clients like
mcp-inspector and VSCode Copilot, which expect resource URLs
following the /mcp/{server_name} convention per RFC 9728.
Changes:
- Add _build_oauth_protected_resource_response() helper
- Add oauth_protected_resource_mcp_standard() endpoint
- Add oauth_authorization_server_mcp_standard() endpoint
- Keep legacy endpoints for backward compatibility
- Add tests for both URL patterns
Fixes#17272
* fix(mcp): Add standard MCP URL pattern support for OAuth discovery (#17272)
OAuth discovery endpoints now support both URL patterns:
- Standard MCP pattern: /mcp/{server_name} (new)
- Legacy LiteLLM pattern: /{server_name}/mcp (backward compatible)
The standard pattern is required by MCP-compliant clients like
mcp-inspector and VSCode Copilot, which expect resource URLs
following the /mcp/{server_name} convention per RFC 9728.
Changes:
- Add _build_oauth_protected_resource_response() helper
- Add oauth_protected_resource_mcp_standard() endpoint
- Add oauth_authorization_server_mcp_standard() endpoint
- Keep legacy endpoints for backward compatibility
- Add tests for both URL patterns
Fixes#17272
* Test was relocated
* refactor(mcp): Extract helper methods from run_with_session to fix PLR0915
Split the large run_with_session method (55 statements) into smaller
helper methods to satisfy ruff's PLR0915 rule (max 50 statements):
- _create_transport_context(): Creates transport based on type
- _execute_session_operation(): Handles session lifecycle
Also changed cleanup exception handling from Exception to BaseException
to properly catch asyncio.CancelledError (which is a BaseException subclass
in Python 3.8+).
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* test(mcp): Fix flaky test by mocking health_check_server
The test_mcp_server_manager_config_integration_with_database test was
making real network calls to fake URLs which caused timeouts and
CancelledError exceptions.
Fixed by mocking health_check_server to return a proper
LiteLLM_MCPServerTable object instead of making network calls.
* test(mcp): Fix skip condition to properly detect claude model names
The skip condition for missing API keys was checking for "anthropic" in
the model name, but the test uses "claude-haiku-4-5" which doesn't match.
Updated to check for both "anthropic" and "claude" model patterns.
Also added skip condition for OpenAI models when OPENAI_API_KEY is not set.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* test(mcp): Fix skip condition to properly detect claude model names
The skip condition for missing API keys was checking for "anthropic" in
the model name, but the test uses "claude-haiku-4-5" which doesn't match.
Updated to check for both "anthropic" and "claude" model patterns.
Also added skip condition for OpenAI models when OPENAI_API_KEY is not set.
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* fix(proxy): support slashes in google route params
* fix(proxy): extract google model ids with slashes
* test(proxy): cover google model ids with slashes
* [Feat] Add model parameter to Generic Guardrail API
Add model information to guardrail requests, allowing guardrails to make
model-specific security decisions.
Changes:
- Add `model` field to GenericGuardrailAPIInputs TypedDict
- Add `model` field to GenericGuardrailAPIRequest Pydantic model
- Update OpenAI and Anthropic handlers to pass model from request/response
- Add unit tests for model parameter handling
* [Feat] Add model parameter to all guardrail_translation handlers
Extend model parameter support to all guardrail handlers for consistent
implementation across all endpoint types:
- OpenAI Responses API (input/output + streaming)
- OpenAI Image Generation (input only)
- OpenAI Text Completion (input/output)
- OpenAI Text-to-Speech (input only)
- OpenAI Audio Transcription (output only)
- Cohere Rerank (input only)
- Pass-through Endpoints (input/output)
- MCP Server (input only)
This addresses the review feedback requesting consistent model parameter
handling across all guardrail_translation/handler.py files.
---------
Co-authored-by: Igal Boxerman <igal@pillar.security>
- Add @lru_cache decorator to get_model_info() and _cached_get_model_info_helper()
- Update _invalidate_model_cost_lowercase_map() to clear these caches when model_cost changes
- Update test to call cache invalidation after modifying litellm.model_cost
Reduces get_model_cost_information from 46% to <1% of request handling time.
Add support for /embeddings endpoint via Vercel AI Gateway.
Closes#19658
Changes:
- Add VercelAIGatewayEmbeddingConfig in litellm/llms/vercel_ai_gateway/embedding/
- Register provider in utils.py and main.py
- Add unit tests for embedding transformation
- Update documentation with embeddings examples
Usage:
```python
from litellm import embedding
response = embedding(
model="vercel_ai_gateway/openai/text-embedding-3-small",
input="Hello world",
api_key="your-api-key"
)
```
* Fix: SSO user roles are not updated for existing users
Fixes#19620
* Refactor: Remove redundant user_info retrieval in SSOAuthenticationHandler
* Test: add new tests for user creation and updates in get_user_info_from_db