* fix(proxy): honor object_permission for managed vector store access
* perf(proxy): preload team object_permission on UserAPIKeyAuth
Populate team_object_permission during virtual-key and JWT auth when the
team is loaded, so can_user_access_vector_store uses it in memory first
and only falls back to get_object_permission by id when missing.
Made-with: Cursor
Three concerns raised by bot reviewers, all addressed:
1. CodeQL cyclic-import warning
``experimental_pass_through/transformation.py`` imported from the
parent ``..transformation`` module, which CodeQL flagged as a
potential cycle. Extracted the helper into a new leaf module
``vertex_ai_partner_models/anthropic/output_params_utils.py`` that
has no heavy imports of its own. Both transformation files now
import from it cleanly. Renamed the helper from the underscore-
prefixed ``_sanitize_vertex_anthropic_output_params`` to the
public ``sanitize_vertex_anthropic_output_params`` since it is now
shared across modules.
2. Greptile P2: redundant ``None`` guard on ``extra_kwargs``
``handler.py`` had two ``extra_kwargs = extra_kwargs if ... else {}``
coercions; the second was a no-op because line 220 already
coerced. Removed the second one and added a NOTE comment so future
readers understand ``extra_kwargs`` is guaranteed non-None at the
point of use.
3. Greptile P2: misleading "already translated" docstring
The docstring claimed the translator above mapped
``output_config.format`` to ``response_format``, but Greptile
correctly traced the code and found that only the legacy top-level
``output_format`` was being translated — ``output_config.format``
was being silently dropped on the adapter path. Two-part fix:
a. Code: extended ``_translate_output_format_to_openai`` to accept
both shapes (top-level ``output_format`` AND
``output_config.format`` sub-key). Top-level still takes
precedence when both are supplied. This means callers using the
newer Anthropic Structured Outputs API now have their schema
properly forwarded to non-Anthropic backends as
``response_format``.
b. Tests: rewrote the misleading docstring to describe what
actually happens, plus added two new tests:
* ``test_output_format_top_level_still_translates`` —
regression guard for the legacy path
* ``test_output_format_takes_precedence_over_output_config_format``
— documents the precedence rule explicitly
Tests: 28/28 pass (was 26/26 before; +2 for the new translation
behavior + precedence). All run in ~0.5s, no real network calls.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Resolves the silent strip of Anthropic Structured Outputs across the
Vertex AI Claude transformation paths and the Anthropic-adapter
re-merge. Consolidates and supersedes four stalled community PRs
addressing overlapping aspects of the same root bug:
- #23475 (Vertex AI Claude blanket-strip removal)
- #23396 (Vertex AI Claude conditional passthrough)
- #23706 (Anthropic adapter exclude output_config from non-Anthropic
backends)
- #22727 (Anthropic adapter strip output_config for non-Anthropic
backends)
Closes / addresses: #23380 (Vertex AI Claude output_config drop),
related: #26423, #25079, #24549, #25971, #25957, #26163, #24856.
What was broken
---------------
* Vertex AI Claude paths called ``data.pop("output_config")`` and
``data.pop("output_format")`` unconditionally even when Vertex
accepted those fields. Callers asking for Structured Outputs got a
200 with prose and never knew the schema constraints had been
silently dropped (often masked for months by permissive fallback
parsers).
* The ``/v1/messages`` -> ``/chat/completions`` adapter
(``LiteLLMMessagesToCompletionTransformationHandler``) re-merged the
raw Anthropic-shaped ``output_config`` into ``completion_kwargs``
AFTER the translator already mapped its meaningful parts to
``response_format`` / ``reasoning_effort``. Non-Anthropic backends
(Azure OpenAI, Fireworks, Bedrock Nova, etc.) then 400'd with
"Extra inputs are not permitted".
Approach
--------
Vertex AI Claude (chat-completion + experimental_pass_through paths):
Replace the unconditional pop with a sanitizer
``_sanitize_vertex_anthropic_output_params`` that strips only the
Vertex-unsupported keys (today: ``effort``) from ``output_config``
while forwarding ``format`` and the legacy top-level
``output_format``. Defensive: non-dict ``output_config`` values are
dropped to avoid sending malformed payloads downstream.
Greptile P1 from PR #23396 addressed: when ``output_config`` carries
both ``format`` and ``effort``, the prior conditional pass-through
forwarded ``effort`` and reproduced the 400. The new helper filters
per-key.
Anthropic ``/v1/messages`` adapter:
Add ``output_config`` to a named module-level constant
``ANTHROPIC_ONLY_REQUEST_KEYS`` and wire it into ``excluded_keys`` so
the post-translation re-merge skips re-adding the raw key. This
fixes the 400 on non-Anthropic backends and avoids the conflicting
duplicate (``response_format`` + raw ``output_config``) on
Anthropic-family backends.
Greptile P2 from PR #23706 addressed: the constant gives reviewers
one grep target instead of an inline literal that silently grows.
Greptile P2 from PR #22727 addressed: ``extra_kwargs or {}`` is
replaced with explicit ``is None`` checks so empty-dict callers no
longer skip the fallback path.
Tests
-----
* tests/test_litellm/llms/vertex_ai/vertex_ai_partner_models/anthropic/
test_vertex_ai_partner_models_anthropic_transformation.py:
- 5 new/updated cases plus a direct unit test for
``_sanitize_vertex_anthropic_output_params``.
- Updated ``test_vertex_ai_claude_sonnet_4_5_structured_output_fix``
so its mock-injected ``output_format`` is asserted to FLOW THROUGH
(the original test asserted the now-buggy strip behavior).
* tests/test_litellm/llms/anthropic/experimental_pass_through/
adapters/test_handler_output_config_passthrough.py (new):
- Constant export sanity, output_config strip with ``effort`` only,
output_config strip with ``format`` only, regression guard that
unrelated extras still flow, explicit-empty-dict path, and the
``extra_kwargs=None`` no-crash path.
Test-quality fixes incorporated from Greptile review on the
superseded PRs:
* No ``inspect.getsource`` source-text assertions (PR #24114 / #23475).
* ``sys.path`` insertion is anchored to ``__file__`` (PR #23706).
* Assertion messages are positional, not tuple (PR #24114-class bug).
* No ``or {}`` masking explicit empty dicts in helper signatures
(PR #22727).
Verified locally: 26/26 pass with this commit. The new tests
fail (or fail to import) on ``main`` without it.
Out of scope
------------
* The ``max_tokens`` capping logic from PR #22727 — independent
concern, deserves its own PR with a focused test plan.
* Architectural rework of the ``excluded_keys`` mechanism (Greptile
P2 on PR #23706 noted point-fix growth). The named constant gives
maintainers a clear place to extend; a registry-based approach
would be a follow-up.
Co-Authored-By: netbrah <netbrah>
Co-Authored-By: s-zx <s-zx>
Co-Authored-By: invoicepulse <invoicepulse>
Co-Authored-By: cfdude <cfdude>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(team_endpoints): auto-add SSO team members to org for proxy admins
* test: proxy_admin vs team_admin security boundary for team→org move
* screenshots: before/after for team-org SSO fix
* fix(team_endpoints): restore staging security features dropped in SSO commit
Co-Authored-By: Ishaan Jaff <ishaan@berri.ai>
* style: black formatting for team_endpoints
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py --ignore=tests/proxy_unit_tests/test_p… (push) Has been cancelled
Use decoded managed container model_id to resolve deployment credentials for container file calls and add regressions to verify provider/model metadata decoding and api_base selection.
Made-with: Cursor
- _get_masked_values now recurses into nested dict values and covers
additional field name patterns (credentials, password, passwd)
- _row_to_submission_item applies masking before returning litellm_params
- list_guardrails_v2 filters DB and in-memory guardrails to the caller's
team memberships for non-admin users; admins still see all guardrails
- approve_guardrail_submission propagates team_id into the in-memory
guardrail dict so ownership is preserved after approval
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
MCP server CRUD endpoints (/v1/mcp/server*) were bundled with MCP
tool-call / passthrough endpoints under llm_api_routes, so setting
DISABLE_LLM_API_ENDPOINTS=true on admin-only nodes also blocked the
Admin UI from listing, adding, or attaching MCP servers.
Separate mcp_inference_routes (data-plane, gated by
DISABLE_LLM_API_ENDPOINTS) from mcp_management_routes (control-plane,
gated by DISABLE_ADMIN_ENDPOINTS). Keep mcp_routes as a union for
backward compat with allowed_routes=["mcp_routes"] virtual key configs.
Upgrade is_management_route to pattern-aware matching so
/v1/mcp/server/{path:path} resolves for concrete IDs.
Temporary MCP OAuth sessions were kept in process-local memory, so on
multi-instance/LB proxy deployments a session created on instance A could
not be found when the follow-up /server/oauth/{server_id}/... request
landed on instance B.
Persist temporary session records to Redis (encrypted with the existing
proxy encryption helpers) as a best-effort L2 cache alongside the current
in-memory L1. Convert get_cached_temporary_mcp_server to async and await
it from the authorize/token/register OAuth endpoints.
Made-with: Cursor
Vertex multi-region endpoints (e.g. us, eu) use the rep host pattern, not
{geo}-aiplatform.googleapis.com. Regional IDs still contain a hyphen.
common_utils.get_vertex_base_url centralizes the rule for SDK/API URL building.
Proxy pass-through duplicates the same branching in a local get_vertex_base_url
(with trailing slashes) to avoid importing from common_utils there; live
WebSocket passthrough uses the same multi-region host logic for wss://.
Tests cover us/eu for the common_utils helper.
Made-with: Cursor
Sibling tests were mutating litellm.proxy.proxy_server.master_key and
prisma_client with raw setattr. Values leaked across tests in the same
xdist worker, flipping the auth short-circuit in user_api_key_auth and
causing unrelated tests (e.g. test_ui_view_session_spend_logs_pagination)
to return 401 instead of 200.
Replace raw setattr with monkeypatch in the two offending files and add
an autouse conftest fixture that snapshots/restores the known-leaky
module globals for every proxy test.
PR feedback (greptile P1 / veria high): with the previous change, a team
storing mcp_tool_permissions={"my-alias": ["read_file"]} would pass the
server-access check (because the alias expanded to a concrete id in the
allowed-servers list) but the per-server tool lookup still did
dict.get(server_id) against the raw name-keyed dict — missing, returning
None, which callers treat as "no restrictions" → all tools allowed instead
of only the declared ones.
Add MCPServerManager.expand_tool_permissions() that rewrites the dict so
every key is a concrete server_id where possible (tool lists from keys
pointing at the same server are unioned). Unresolved keys pass through
unchanged so stale id-keyed restrictions still apply when the same string
is used for lookup. Wire the helper into the four dict-lookup sites:
get_allowed_tools_for_server (key + team paths), the agent tool lookup,
and the rest_endpoints.py tool filter.
Also switch expand_permission_list to pass through unresolved entries
(rather than dropping them) so existing test fixtures that use bare string
placeholders continue to work. The downstream access check denies unknown
entries when compared to the concrete request server_id, so security
posture is unchanged.
Sanitize the debug log to use %r formatting so an admin-controlled
identifier with newlines can't forge log entries (CodeQL log-injection
warning).
Two fixes to proxy-db CI:
1. test_realtime_webrtc_endpoints.py's `proxy_app` fixture mutated the
module-global `proxy_server.master_key` without restoring it, leaking
state into any test that shared the same xdist worker. Under
--dist=loadscope with 2 workers (GHA proxy-endpoints), this caused the
google_endpoints tests to fail with "No api key passed in." because
user_api_key_auth saw a set master_key and a missing API key on the
test request. The fixture now saves and restores the original value.
2. Address the Greptile note that the semantic shard design has no
catch-all, so a new test file added to tests/proxy_unit_tests/ without
a matrix entry would silently skip CI. Adds an assert-shard-coverage
job that enumerates test_*.py files and fails the workflow if any are
not referenced by a matrix entry, with a clear message telling the
author which semantic shard to place it in. All proxy-db shards now
depend on this guard.
The mocked async_increment_cache_pipeline is invoked from Router's
deployment_callback_on_success, registered as an async success callback.
Those callbacks are enqueued to GLOBAL_LOGGING_WORKER and run on a
background task, so the mock may not have been called yet when the test
asserts on it. Flush the worker before asserting.
Two independent deflakes:
1. test_ui_view_spend_logs_unauthorized (unit) was returning 400 instead
of 401/403 when earlier tests in the file left proxy-auth globals
(prisma_client, master_key, user_custom_auth, general_settings,
user_api_key_cache) in a state that let invalid tokens pass auth and
fall through to the endpoint's own start_date/end_date validation.
Add an autouse fixture that pins those globals to their import-time
defaults for every test in the file. Harden the assertion to include
response body so future flakes are diagnosable.
2. test_basic_spend_accuracy (CI job proxy_spend_accuracy_tests) depends
on the Redis transaction buffer flushing spend to Postgres. The buffer
uses a single global pod-lock key (cronjob_lock:db_spend_update_job)
and a single global buffer list key. Pointing the proxy at the shared
remote Redis means concurrent CI pipelines contend for the same lock
and can drain each other's buffer into the wrong database. Add a
start_redis reusable command that boots a per-job redis:7-alpine
container (digest-pinned), and switch proxy_spend_accuracy_tests to
REDIS_HOST=host.docker.internal:6379 so lock and buffer state are
isolated per CI run.
- Add gpt-5.5 to GPT5_MODELS parametrized list so both OpenAIGPT5Config
and AzureOpenAIGPT5Config routing tests cover the new model.
- Add test_generic_cost_per_token_gpt55 verifying the new entry's
cost-map values ($5/$0.50/$30 per 1M) and that generic_cost_per_token
returns the expected prompt/completion costs.
* feat: add gpt-5.5 to model cost map
Add gpt-5.5 entry with pricing from OpenAI flagship page:
input $5/1M, cached input $0.50/1M, output $30/1M, 272K context.
* test: add gpt-5.5 coverage for model cost map and gpt-5 routing
- Add gpt-5.5 to GPT5_MODELS parametrized list so both OpenAIGPT5Config
and AzureOpenAIGPT5Config routing tests cover the new model.
- Add test_generic_cost_per_token_gpt55 verifying the new entry's
cost-map values ($5/$0.50/$30 per 1M) and that generic_cost_per_token
returns the expected prompt/completion costs.
The periodic budget-window reset job filtered keys/teams with
`where={"budget_limits": {"not": None}}`. The prisma-client-python
library does not support null-filtering on `Json?` columns (no
DbNull/JsonNull sentinel — upstream issue #714). The client drops the
`None` value during serialization and the engine rejects the query with
`MissingRequiredValueError: where.budget_limits.not: A value is
required but not set`, so neither the key nor team reset path runs.
Switch those two `find_many` calls to `query_raw` with
`WHERE budget_limits IS NOT NULL`, selecting only the PK and the
`budget_limits` column. Writes still go through the ORM. Add unit tests
covering the expired/unexpired paths for keys and teams, string-encoded
JSON payloads, empty payloads, error isolation between the two paths,
and a regression guard asserting the query still uses `IS NOT NULL`.
team.object_permission.mcp_servers (and the per-key equivalent) previously
only accepted server_id strings. For config-loaded MCP servers, the id is
derived from a hash that includes the server URL, so the same logical
server in two regions ends up with two different ids in a shared database.
Permission lists had to enumerate every region's id.
Add a single MCPServerManager.expand_permission_list() helper that resolves
each entry against the current region's config + DB registry union: entries
that match a server_id pass through, entries that match an alias/server_name/
name expand to every matching id, and unresolved entries drop with a debug
log so stale or typo entries are diagnosable. Wire it into the four
_get_allowed_mcp_servers_for_* helpers so direct server entries and
mcp_tool_permissions dict keys are both expanded before the intersection.
Access-check outcomes are unchanged for existing id-based permissions;
name-based entries now resolve instead of being silently denied.
Align the Ruby, Node.js, and npm install path with the rest of the
config. Three separate upstream installers were being invoked via
\`curl ... | bash\` or unlocked \`npm install\`:
- RVM's \`get.rvm.io/stable\` installer (mutable upstream script).
Replace with a shallow git clone of the rvm/rvm repo at tag 1.29.12
and verify HEAD matches the published commit SHA before running the
local \`./install\` script. Same pattern already used for the
helm-unittest plugin in .github/workflows/helm_unit_test.yml.
- NodeSource's \`deb.nodesource.com/setup_18.x\` piped into sudo bash.
Replace with a direct download of the Node.js 18.20.8 linux-x64
tarball from nodejs.org, verified against the published
SHASUMS256.txt digest before extraction.
- \`npm install @google-cloud/vertexai @google/generative-ai\` and
\`--save-dev jest\` resolved fresh from the npm registry on every
run. Add \`tests/pass_through_tests/package.json\` with pinned
direct-dep versions and commit the generated package-lock.json, then
switch CI to \`npm ci\` (exact lockfile install, fails on drift).
Also scopes the Ruby+JS test runners to \`tests/pass_through_tests/\`
so they pick up the committed package.json rather than writing
node_modules at repo root.
Address codex review P1 + P2 findings:
- BYOK /token now accepts OAuth 2.1 clients that omit redirect_uri
(draft-15 §4.1.3 dropped the requirement). When the client does
submit a value, equality is still enforced vs the /authorize record.
PKCE + client_id binding cover the security role redirect_uri
played under RFC 6749.
- _user_id_from_session_cookie requires the ``exp`` claim on the UI
session JWT (PyJWT options={"require": ["exp"]}) so leaked cookies
have a bounded lifetime.
- validate_loopback_redirect_uri rejects URIs with a fragment
(RFC 6749 §3.1.2) and catches malformed-URI ValueError so
unparseable input surfaces as 400 invalid_request instead of 500.
The HTTPException arm in _run_centralized_common_checks assumed the
exception came from the team-object fetch, but asyncio.gather raises
the first exception from any of the five gathered coroutines. If
get_user_object / get_project_object / get_end_user_object raises
HTTPException on a token with team_id=None, the assert inside
_team_obj_from_token fires and the outer auth-exception handler
mishandles it.
Guard on team_id before calling _team_obj_from_token and default
team_object to None otherwise. (Greptile P1.)
Address codex review P0 + P1 findings on the discoverable OAuth proxy:
- /callback now re-validates that the decoded base_url is loopback before
302-redirecting to it. State is encrypted but pre-existing states minted
before the /authorize validation was added have no expiry and remain
valid; validating at the sink closes the open-redirect + code-theft
primitive for those stale states too. (VERIA-57 root cause B, P0.)
- /token responses now set Cache-Control: no-store + Pragma: no-cache
per RFC 6749 §5.1 (P1).
- Move TOKEN_NO_CACHE_HEADERS constant from byok_oauth_endpoints into
the shared oauth_utils module so both endpoints use the same value.
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py --ignore=tests/proxy_unit_tests/test_p… (push) Has been cancelled