Orgs with no MCP permissions configured (the common default) previously
returned None without writing to cache, meaning every subsequent MCP
request triggered a fresh find_unique against litellm_organizationtable.
Cache a sentinel string on the negative path so the DB is queried at most
once per cache TTL per org, regardless of whether the org has an
object_permission or not.
Made-with: Cursor
- Cache org object_permission in user_api_key_cache to avoid a DB hit on
every MCP request (was: raw find_unique on every call).
- Expand org mcp_servers list via expand_permission_list() so name-based
entries resolve to canonical IDs, consistent with key/team/end-user path.
- Expand org mcp_tool_permissions via expand_tool_permissions() in both
_get_allowed_mcp_servers_for_org and get_allowed_tools_for_server,
closing the silent name-vs-ID mismatch that could let restricted tools
through.
- The second _get_org_object_permission call in get_allowed_tools_for_server
now hits the cache (warm from the earlier server-list check), resolving
the double DB round-trip without changing the call structure.
Made-with: Cursor
Apply organization object_permission as a ceiling on allowed MCP servers
and tool permissions, consistent with vector store org checks.
Includes unit tests for org ceiling, intersection, and tool filtering.
Made-with: Cursor
The async/sync delete_response_api_handler always passed json=data into
httpx.delete, where data is {} from the transformer. httpx serializes that
to a 2-byte body. The Azure Responses DELETE endpoint now rejects any
request body with code: unexpected_body, breaking
test_basic_openai_responses_delete_endpoint on the llm_responses_api_testing
job. Build the kwargs dict and only set json= when data is truthy.
Add unit tests that patch httpx.delete and assert json/data are not in the
captured kwargs for the Azure DELETE path (sync and async).
The proxy's ingress hardening (commit 842eea0131) now strips client-supplied
`mock_response` from the request body unless the calling key or team has the
`allow_client_mock_response: true` admin-metadata flag set. The e2e model
access tests rely on `mock_response` to short-circuit the LLM call, so without
the flag they hit real backends — the bedrock wildcard route fakes out to a
shared example endpoint that now 404s on unsupported paths, causing
`test_model_access_patterns[key_models2-bedrock/anthropic.claude-3-True]`
(and the bedrock/anthropic.* row that pytest -x never reaches) to fail.
Set `allow_client_mock_response: true` on every key and team this test file
provisions so `mock_response` is preserved end-to-end.
Cover the full litellm.rerank()/arerank() path with HTTP mocked, asserting
metadata.requester_metadata reaches the Discovery Engine :rank body as
userLabels (and stays absent when no metadata is set). Catches plumbing
regressions that unit tests on transform_rerank_request alone would miss.
The wrapper had no production callers after transform_parsed_response
was refactored to call _resolve_json_mode_non_streaming directly.
Updated the parametrized test to call the underlying method.