* fix(guardrails): apply team-level guardrails alongside global policy guardrails
Two bugs prevented team-direct guardrails from being automatically applied
when using a team-scoped API key:
1. Auth caching: `valid_token.team_metadata` was never refreshed from the
freshly-fetched team object at the "Check 6" step in
`_user_api_key_auth_builder`. Guardrails added to a team after the key
was first cached were therefore invisible to `move_guardrails_to_metadata`.
Fix: propagate `_team_obj.metadata` → `valid_token.team_metadata` after
every "Check 6" team fetch (user_api_key_auth.py).
2. Guardrail execution: `get_guardrail_from_metadata` checked
`data["litellm_metadata"]` before `data["metadata"]`. When a request
carried a non-empty `litellm_metadata` without a "guardrails" key, the
merged guardrail list written to `data["metadata"]` by
`move_guardrails_to_metadata` was shadowed and the guardrail received an
empty requested-guardrails list (custom_guardrail.py).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix merge conflict
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Mirrors the existing api_base handling in _load_credentials_from_config():
when a vector_store_config references litellm_credential_name and the stored
credential does not itself define aws_sts_endpoint or aws_web_identity_token,
any caller-supplied value on those keys is removed. Keeps endpoint and
identity fields aligned with the credential definition.
* fix(bedrock guardrail): dedupe post-call log when only post_call is configured
When a Bedrock guardrail runs with only post_call configured, the post-call
trace section showed the same guardrail twice (one entry per parallel
INPUT/OUTPUT API call). Add skip_logging param to make_bedrock_api_request
and pass it for the INPUT scan so the OUTPUT scan stands as the single
canonical post_call log entry. INPUT exceptions still propagate, so the
input-side blocking behavior is preserved.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(bedrock guardrail): post_call only scans OUTPUT, not INPUT
post_call is the response-validation hook by definition — input scanning
belongs to pre_call / during_call. The previous code ran an extra INPUT
scan in post_call when no pre/during hook was configured, which produced
a duplicate "post-call" entry in the trace and was semantically wrong
for a "post-call" event.
Drops the should_validate_input branch and parallel asyncio.gather in both
async_post_call_success_hook and async_post_call_streaming_iterator_hook
in favor of a single OUTPUT scan. Reverts the now-unneeded skip_logging
parameter on make_bedrock_api_request introduced in the previous commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Fall back to email match when looking up the caller in
members_with_roles — email-onboarded members may have user_id=None on the
stored entry, which caused a false 404 for valid members. (P1)
- Replace 3 raw Prisma queries with get_team_object / get_team_membership /
get_user_object so the endpoint reuses the cache + retry layer the rest of
the proxy uses. (P2)
- Allow internal_user role to reach /team/{team_id}/members/me by adding the
route to LiteLLMRoutes.self_managed_routes (the handler already enforces
member-of-team access).
- Return null from the UI fetch on 404 instead of throwing, so a proxy admin
who isn't a team member sees the existing empty state rather than an error
string in the always-visible tab. (P2)
- Move the fetch out of networking.tsx into a colocated React Query hook
(useMyTeamMember) next to MyUserTab; TeamInfo now passes only teamId.
- Tooltip + empty-state copy on Model Scope: drop "(all team models)"
parenthetical and the redundant tooltip line.
- Tests: build real LiteLLM_TeamMembership / LiteLLM_BudgetTableFull
fixtures (with created_at) so the Pydantic Union resolves to the Full
variant; add an assertion that budget_reset_at survives end-to-end; add a
test for the email-only member match path.
* feat(azure): add azure/gpt-5.5 + azure/gpt-5.5-pro entries (+ dated variants)
Azure variants of OpenAI's GPT-5.5 family. Microsoft has not yet
shipped GPT-5.5 on Azure OpenAI (latest GA on the Foundry models page
is GPT-5.4 as of 2026-04-24), but adding the entries day-0 mirrors the
established precedent for azure/gpt-5.4* (which were in the cost map
before the Azure rollout) so cost tracking and capability flags work
the moment customers deploy.
Schema follows the existing azure/gpt-5.4* shape:
- Same base/long-context pricing as openai/gpt-5.5*: $5/$30 chat,
$60/$360 pro per 1M, with priority tier 2x base
- Azure variants drop the flex/batches keys (Azure has no flex tier)
but keep priority pricing, matching gpt-5.4* precedent
- mode=chat for the thinking model, mode=responses for pro
reasoning_effort capability flags mirror the OpenAI variants exactly
since Azure proxies the same API contract: minimal rejection on both
chat and pro, low/none rejection on pro. Once #26456 (which sets
supports_low_reasoning_effort + minimal=false on openai/gpt-5.5*)
lands, OpenAI and Azure flag profiles align.
Tests pin entry presence + pricing for all four Azure variants and
verify the live-API-derived reasoning_effort flags.
* test: register supports_low_reasoning_effort in cost-map JSON schema
azure/gpt-5.5-pro and azure/gpt-5.5-pro-2026-04-23 added in this branch
carry supports_low_reasoning_effort=false. The strict
'additionalProperties: false' schema in
test_aaamodel_prices_and_context_window_json_is_valid rejected the new
key. Register it alongside the other supports_*_reasoning_effort
entries.
Note: the runtime side of this flag (code that reads it) lands in
#26456. Until that PR merges the flag is inert for both Azure and
OpenAI pro entries, but having the schema accept it lets cost-map
tests pass on either merge order.
Adds a new "My User" tab on the team detail page (between Overview and
Virtual Keys) so non-admin team members can see their own spend, budget,
budget reset date, rate limits, model scope, and team role.
Backend
- New `GET /team/{team_id}/members/me` endpoint that resolves the caller
from the API key and returns only their own LiteLLM_TeamMembership row
plus minimal team context (alias, role, email). Returns 404 if the
caller is not a member of the team. Avoids exposing other members'
data, which would happen if we filtered `/team/info` client-side.
- New `TeamMemberInfoResponse` Pydantic model.
Frontend
- New `MyUserTab` component (antd) — read-only summary cards.
- New `teamMemberMeCall` helper in networking.tsx.
- Tab is visible to all team members (including non-admins).
When litellm_credential_name resolves to no values (name not registered
in litellm.credential_list), bail out before touching the resolved
config. Prevents stripping a caller-supplied api_base in the no-match
case.
Mirrors the existing vertex_credentials handling for the newer
vertex_ai_credentials field, and extends the dynamic masker's sensitive
pattern set to recognize the plural form so other plural-named credential
fields are also covered.
When a vector_store config references a stored credential by name,
overwrite caller-supplied values on the same keys with the credential
values, and drop api_base from the resolved config when the credential
does not define one. Keeps the api_key / api_base pair consistent with
the credential definition.
Bring `/server/oauth/{server_id}/authorize`, `/token`, and `/register`
in line with `fetch_mcp_server`: the helper that resolves the server now
also applies the per-caller access policy. Admin-view callers are
unrestricted; non-admins must have the server in their allowed-servers
set; servers resolved from the admin-only `/server/oauth/session`
temporary cache reject non-admins.
dimensions=1 was silently dropped before #24415 wired the param through
to Vertex. Now that it's forwarded, Vertex rejects it (valid values:
128, 256, 512, 1408). Switch the test to a valid dimension.
Add a new CI workflow that rejects pull requests from forks when they:
- Modify uv.lock (any change at all)
- Add new dependencies to any pyproject.toml file (root, litellm-proxy-extras, enterprise)
Security properties:
- Uses pull_request (not pull_request_target) so no secrets are exposed
- All action refs pinned to full SHA hashes
- persist-credentials: false on all checkouts
- permissions: {} (no GitHub token permissions)
- No user-controlled input in run: blocks (no script injection)
- Proper TOML parsing via stdlib tomllib (not regex on raw text)
- Only triggers when dependency files are actually changed (paths filter)
Internal PRs (from branches in the canonical repo) skip the job entirely.
Co-authored-by: Krrish Dholakia <krrish-berri-2@users.noreply.github.com>
Sync CI configs with upstream/litellm_internal_staging:
- Migrate from Poetry to uv (PR #25007)
- Pull in zizmor security fixes for workflows
- Add isolated unit test workflows + reusable bases
- Drop redundant matrix workflow and azure-batches workflow
- Drop .circleci/requirements.txt (replaced by uv lockfile)
After ``return_exceptions=True``, gather results are typed
``Any | BaseException``. Mypy could not narrow through
``None if isinstance(x, HTTPException) else x`` ternaries, leaving
BaseException in the type passed to ``common_checks``.
Switch the narrowing checks to ``BaseException`` (HTTPException is the
only one still possible after authz failures were re-raised) and add
explicit ``Optional[...]`` annotations so the typed objects flow into
``common_checks``. No behavior change.
The asyncio.gather in `_run_centralized_common_checks` ran with
`return_exceptions=False` and a single bare `except HTTPException`
arm, so an HTTPException from any one fetch (the realistic case is a
404 from `get_team_object` when a token references a deleted team)
zeroed out the user, end-user, project, and global-spend contexts in
addition to falling back the team object. That silently skipped the
user budget, end-user budget, and project enforcement passes inside
`common_checks` for the unrelated contexts that had actually fetched
fine.
Switch to `return_exceptions=True` and apply per-fetch fallback
(matches the pre-refactor per-fetch try/except pattern in the builder):
- ProxyException / BudgetExceededError still propagate as authz failures.
- HTTPException on the team fetch reconstructs from the token; on the
other fetches it nulls only that one context.
- Successful fetches always reach `common_checks` intact.
Adds two unit tests covering the team-404 and user-404 cases to lock
the per-fetch isolation in. Drops the inaccurate `PROXY_ADMIN tokens
short-circuit` claim from the docstring — admin tokens still flow
through `common_checks`; admin status is only honored where the
underlying check exempts it.
- test_prepare_key_update_data: replace bare MagicMock with
MagicMock(spec=LiteLLM_VerificationToken) and explicitly set
existing_key_row.metadata = {}, so reserved-field reads return real
values instead of MagicMock-returning-MagicMock. Fixes a regression
surfaced by the new reserved-metadata preservation logic.
- test_key_management_endpoints.py: black-format-only changes from
recent edits.
The previous version appended NULLS LAST to every ORDER BY, which would
silently change DESC semantics for any nullable sort column added to the
whitelist later. Today the existing sort columns (spend, total_tokens,
startTime, endTime, request_duration_ms, model) are all non-null in the
result set, so the clause is a no-op for them — but the broader form is
misleading.
Apply NULLS LAST only when sorting by ttft_ms (the only column whose
computed expression actually produces NULLs). Update the model test to
assert the clause is absent for non-nullable columns.
The routes in `global_spend_tracking_routes` (e.g. /global/spend/report,
/global/spend/teams, /global/spend/keys) return spend aggregated across
every team, customer, and api_key in the proxy. They were included in
`internal_user_routes` and `internal_user_view_only_routes`, so non-admin
roles could read proxy-wide spend.
Drop them from both non-admin route lists. PROXY_ADMIN and
PROXY_ADMIN_VIEW_ONLY access is preserved through their existing branches
in route_checks.py, and the `get_spend_routes` permission opt-in
continues to grant access for keys that need it.
Updates two pre-existing test parametrizations whose expected results
flip from True to False, and adds parametrized coverage over every
route in `global_spend_tracking_routes` for: PROXY_ADMIN_VIEW_ONLY
allowed, INTERNAL_USER blocked, INTERNAL_USER_VIEW_ONLY blocked,
INTERNAL_USER + get_spend_routes permission allowed.
Extend the /spend/logs/ui sort_by whitelist to accept "model" and
"ttft_ms", and wrap the Model and TTFT (s) column headers with the
existing SortableHeader component so users can sort by either.
TTFT has no stored column, so it is computed inline as
(completionStartTime - startTime) milliseconds. Non-streaming rows
(completionStartTime null or equal to endTime) yield NULL and the
ORDER BY uses NULLS LAST so they always sort to the bottom regardless
of direction, matching the existing "-" display in the UI.
Adds parametrized backend tests for sort_by=model and a dedicated test
covering streaming + non-streaming TTFT ordering in both directions.