The startup Redis-prerequisite warning for PKCE gated the Okta path on
OKTA_CLIENT_ID being set at process start. The runtime check in
_is_oidc_pkce_enabled does not — Okta defaults PKCE to enabled whenever
OKTA_CLIENT_USE_PKCE is unset or 'true'. When Okta credentials are
injected after startup via the SSO settings API, the runtime PKCE flow
activates but the startup warning never fires, hiding the multi-instance
Redis requirement.
Use the centralized _is_oidc_pkce_enabled helper for both generic and
Okta at startup so the warning fires whenever PKCE *could* activate at
request time.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
get_generic_sso_response had 52 statements (ruff PLR0915 limit is 50).
Move the 5 PKCE callback guards (state/code/client_id/token_endpoint/type)
into a new _validate_pkce_oauth_callback helper. Pure refactor — same
exceptions, same control flow.
- detectSSOProvider mock in EditSSOSettingsModal.test.tsx now matches the
simplified real implementation (generic_client_id always maps to 'generic').
- generic_response_convertor: extract _get_oidc_attribute_env helper that
skips the redundant GENERIC_* fallback when env_prefix is already GENERIC.
- get_generic_sso_response: when PKCE is enabled but the verifier was lost
from cache and no client_secret is configured, raise an actionable
ProxyException instead of falling through to fastapi-sso with a None
secret (which produced an opaque token-exchange error).
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Bug #7 (ui_sso.py auth_callback): env_prefix and generic_client_id were
derived solely from whether OKTA_CLIENT_ID was set, so a Google or Microsoft
login with OKTA_CLIENT_ID also configured would route through the generic
attribute-mapping branch with OKTA_* env-var lookups. Track the active
provider in the elif chain and derive env_prefix / callback client_id from
that — Google/Microsoft pass None.
Bug #8 (SSOModals.tsx): handleFormSubmit duplicated processSSOSettingsPayload
from utils.ts. The two had already diverged (legacy modal missing
use_team_mappings / team_ids_jwt_field handling and the
provider-supports-role-mappings guard). Replace inline logic with the shared
utility.
Users who configured Okta via generic OIDC env vars (GENERIC_*) previously had
their config detected as 'okta' but rendered against okta_* fields, leaving the
SSO settings page empty. Treat presence of generic_client_id as the 'generic'
provider so the existing generic_* values render correctly.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
- ui_sso.debug_sso_callback: Okta branch now captures received_response
and access_token_payload from get_generic_sso_response so the debug
page surfaces raw_claims and access_token_claims (matches generic
branch behavior).
- ui_crud_endpoints.update_sso_settings: redact secret fields in the
PATCH 200 response body so plaintext secrets (including those restored
from the **redacted** sentinel) never leak back to the caller. GET was
already redacted; PATCH now matches.
- Update test_update_sso_settings to expect redacted secrets in the
response.
- ui_sso: thread env_prefix through get_redirect_response_from_openid
and _get_user_email_and_id_from_result so OKTA_USER_ROLE_ATTRIBUTE
is honored on the Okta callback path (falls back to
GENERIC_USER_ROLE_ATTRIBUTE then 'role').
- handle_jwt: use 'is None' rather than 'or' when falling back from
litellm_jwtauth.audience/issuer to JWT_AUDIENCE/JWT_ISSUER env vars,
so explicit empty values are not silently overridden.
- proxy_setting_endpoints: use _decrypt_db_variables instead of
_decrypt_and_set_db_env_variables when restoring redacted SSO
secrets, to avoid polluting os.environ with lowercase model-field
names.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Read-only admins could previously retrieve plaintext okta_client_secret,
google_client_secret, microsoft_client_secret, and generic_client_secret
via GET /get/sso_settings. Now these fields are replaced with "**redacted**"
in the response. PATCH /update/sso_settings preserves the stored secret
when the sentinel value is echoed back unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- _types.py: narrow LiteLLM_JWTAuth.issuer from Union[str, List[str]] to
str — PyJWT performs strict string equality on the 'iss' claim and
silently rejects all tokens when passed a list value
- SSOModals.tsx: replace inline provider-detection block with the shared
detectSSOProvider() utility (same logic already used in
EditSSOSettingsModal.tsx via utils.ts) to prevent the two paths
from diverging when new providers are added
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- handle_jwt.py: initialize self.litellm_jwtauth in JWTHandler.__init__
so legacy tests that skip update_environment() don't crash with
AttributeError (fixes 6 tests in tests/proxy_unit_tests/test_jwt.py)
- ui_sso.py: add # noqa: PLR0915 to generic_response_convertor and
sso_readiness which exceed ruff's 50-statement limit after the Okta
additions (fixes Ruff PLR0915 lint errors)
- test_env_keys.py: add PENDING_DOCS_PR_VARS exclusion for the six
new OKTA_* env-var literals until BerriAI/litellm-docs#110 is merged
(fixes documentation and code-quality CI checks)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(gemini): normalize response_schema on native generateContent
The /v1beta/models/{model}:generateContent passthrough forwarded
generationConfig.response_schema verbatim, so schemas containing $defs,
$ref, anyOf-with-null, default, or title were rejected by Gemini even
though /chat/completions already handles them.
GoogleGenAIConfig.transform_generate_content_request now calls a new
_normalize_response_schema helper that mirrors the chat/completions
path: Gemini 2.0+ models get the schema promoted to responseJsonSchema
via _build_json_schema (preserving $defs/$ref natively), older models
keep responseSchema but the schema is flattened with
_build_vertex_schema. VertexAIGoogleGenAIConfig (which overrides the
transform entirely) calls the same helper before building the request.
* fix(gemini): preserve caller-supplied responseJsonSchema when responseSchema co-present
Previously, when both responseJsonSchema and responseSchema were present
on Gemini 2.0+, _normalize_response_schema processed responseJsonSchema
first (no-op normalization) then unconditionally promoted responseSchema
to responseJsonSchema, clobbering the caller-supplied value.
Now skip the promotion (and drop the redundant responseSchema) when the
caller already supplied responseJsonSchema.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* chore: strip restating comments from response-schema normalize
Drop the docstring on _normalize_response_schema and the two inline
comments that just restated what the surrounding code/asserts already
say. Function name + variable names carry the intent; PR description
covers the why-it-exists context.
* perf(gemini): drop redundant deepcopy on responseJsonSchema normalize
_build_json_schema is a no-op (returns its argument unchanged), so the
deepcopy + round-trip on the responseJsonSchema branch allocated a full
schema copy on every request with no observable effect. Forward the
caller's value as-is, and just move the popped responseSchema value when
promoting on Gemini 2.0+.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* style: remove unneeded comment
* fix(gemini): drop unsupported responseJsonSchema for older models
* test(gemini): add parity test between native and chat schema normalization
Per @Sameerlite review: lock the two Gemini schema-normalization paths
together. If either GoogleGenAIConfig._normalize_response_schema (native
generateContent) or VertexGeminiConfig.apply_response_schema_transformation
(/chat/completions) drifts, the parity test fails — forcing both to be
updated together.
* fix(google_genai): preserve key naming convention in _normalize_response_schema
When the input schema key is snake_case (response_schema), the promoted
JSON schema key should also be snake_case (response_json_schema) instead
of mixing in camelCase (responseJsonSchema). This matters for the Vertex
AI google_genai path which converts all keys to snake_case before
calling _normalize_response_schema.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
OpenAI returns 'The model dall-e-3 does not exist' for the test account,
breaking test_openai_img_gen_health_check and test_image_generation.
Switch to gpt-image-1, matching the existing TestOpenAIGPTImage1 pattern.
SERVER_ROOT_PATH is a process-startup env var. Read it once in
__init__ instead of calling get_server_root_path() + rstrip on every
request that arrives before all lazy features have loaded.
LazyFeatureMiddleware compared the raw scope path against registered
prefixes (e.g. /policies), so requests under a server root path like
/api/v1/policies/... never matched, the feature never loaded, and the
endpoint returned 404. Strip the configured root path before matching,
normalizing trailing slashes and enforcing a component boundary so
/api does not falsely match /apiv2.
1. Missing litellm_request child span when proxy parent in metadata:
_get_span_context now returns (ctx, None) for the metadata-injected
proxy parent so the primary span is always emitted as a child of ctx.
Proxy span lifecycle managed by new _end_proxy_span_from_kwargs.
2. open_telemetry_logger overwrite by later handlers:
_init_otel_logger_on_litellm_proxy now uses first-registered-wins —
only assigns proxy_server.open_telemetry_logger when currently None.
3. Duplicate litellm_request success spans in streaming paths:
Added _mark_success_span_once with per-handler dedupe key stored in
kwargs metadata, suppressing the second span when both sync and async
success callbacks fire for the same request.
Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`apply_client_tag_policy_pre_auth` overwrote string-typed metadata
with `{}` before merging header tags, dropping any tags inside. A
caller could send `metadata='{"tags":["over-budget"]}'` plus
`x-litellm-tags: within-budget` and bypass `_tag_max_budget_check`
on the body tag. Parse the string via `safe_json_loads` first so
existing tags survive the merge.
Also drop the empty `tests/test_litellm/proxy/credential_endpoints/`
directory — the cascade-rename tests it held imported a function
that was never implemented (out of scope for this PR).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
``model_dump(exclude_unset=True)`` in ``prepare_key_update_data``
includes any field the caller explicitly set, even when the value is
``None``. The previous guard short-circuited on ``getattr(data,
'user_id', None) is None``, which conflated "field omitted" (safe)
with "field explicitly set to null" (writes NULL to the token row,
detaching the key from its user and bypassing user-row role
checks).
Switch the omitted-vs-set distinction to ``data.model_fields_set``;
treat explicit-null and explicit-empty-string identically as a
removal attempt, both 403-rejected for non-admin callers.
Parametrized regression adds ``explicit_null_blocked`` alongside the
existing ``rebind_blocked`` / ``empty_blocked`` / ``same_user_id_allowed``
cases.
Removes the allow_client_tags metadata check from apply_client_tag_policy_pre_auth so
x-litellm-tags headers are always merged into request metadata, matching the post-auth
behavior in add_litellm_data_to_request. Updates pre-call tests accordingly and adds a
new test suite covering cascading credential renames into model rows.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Previous wording read "User=<new_owner> is not allowed to update the
key to belong to user=<current_owner>" — easy to misread as "caller
wants to keep the key on its current owner". Reframe as
"Non-admin caller is not allowed to rebind the key from
user=<existing> to user=<incoming>" so the direction of the failed
operation is unambiguous.
Same shape preserved (HTTPException 403); only the ``detail`` string
changes. Regression test substring updated.
A non-admin caller could rebind their own key's ``user_id`` via
``/key/regenerate``. ``_execute_virtual_key_regeneration`` had org/team
guards but no ``user_id`` guard, and ``prepare_key_update_data`` did not
strip the field — it survived ``model_dump(exclude_unset=True)`` into
the Prisma update. On the next request,
``_return_user_api_key_auth_obj`` resolved the rebound ``user_id``
against ``litellm_usertable`` and returned ``PROXY_ADMIN`` whenever
the target row's ``user_role`` was admin (e.g. the default
``user_id="default_user_id"`` created on first password-UI login).
``/key/update`` had the equivalent guard inline at
``_validate_update_key_data``; extract it to a shared helper
``_validate_caller_can_change_key_ownership`` and call from both
``/key/update`` and ``_execute_virtual_key_regeneration``. Future
regenerate-style endpoints inherit the guard for free.
Also tighten the premium gate that allowed the master-key rotation
branch to skip the enterprise check. The previous predicate was
``data.new_master_key is not None`` — a field-presence test, not an
identity check. Any non-premium caller could send any value in that
field and the premium check would no-op. Verify the caller actually
holds the master key via ``_is_master_key`` before allowing the
non-premium path.
Tests:
- ``test_regenerate_user_id_rebind_guard`` — parametrized table over
cross-user rebind (blocked), empty-string removal (blocked), and
same-user no-op rebind (allowed).
- ``test_regenerate_premium_gate_requires_actual_master_key`` /
``test_regenerate_premium_gate_allows_actual_master_key_holder`` —
ensure the premium check requires the caller actually present the
master key, and that legitimate master-key rotation still works.
* fix(proxy): resolve cache handling issues in _lookup_deprecated_key
- Updated the in-memory cache for deprecated key lookups to store a 3-tuple (active_token_id, cache_expires_at_ts, revoke_at_ts) instead of a 2-tuple, ensuring proper unpacking and backward compatibility.
- Removed duplicate cache reads and added logic to handle legacy cache entries gracefully.
- Enhanced unit tests to cover scenarios for cache hits, DB misses, and respect for revoke_at timestamps, ensuring robust handling of the grace-period key-rotation feature.
* refactor(proxy): streamline cache handling in _lookup_deprecated_key
- Simplified the cache retrieval logic by directly unpacking the 3-tuple cache entries, removing the need for backward compatibility checks for 2-tuple entries.
- Updated unit tests to ensure that pre-warmed 3-tuple cache entries are served correctly without unnecessary database lookups.
* chore(ci): add new unit test for deprecated key grace period
- Included `test_deprecated_key_grace_period.py` in the CI workflow to enhance coverage for deprecated key handling scenarios.
* fix(proxy): remove unnecessary check for revoke_at in _lookup_deprecated_key
- Eliminated the redundant check for None on revoke_at, streamlining the logic for handling deprecated keys in the cache. This change enhances the efficiency of the key lookup process.
* test(proxy): add end-to-end tests for deprecated key lookup behavior
- Introduced a new test class `TestDeprecatedKeyLookupDbE2E` to validate the behavior of deprecated key lookups against a real Prisma-backed database.
- The test ensures that old key hashes resolve correctly and that repeated lookups utilize the in-memory cache without errors.
- Cleaned up the `_lookup_deprecated_key` function by removing an unnecessary check for `revoke_at`, enhancing the efficiency of the key lookup process.
* feat(ui): add Vertex AI Search as vector store provider
Adds a "Vertex AI Search" entry to the provider dropdown
(custom_llm_provider=vertex_ai/search_api) with fields for project,
location (global/us/eu select), and optional collection ID. Extends
VectorStoreFieldConfig with `options` so select fields can be
data-driven instead of falling through to the embedding-model list.
* fix(ui): clarify vertex_collection_id placeholder copy
Placeholder previously displayed "default_collection" — the literal
fallback value — which invited users to type it instead of leaving the
field blank. Switch to an example placeholder and tighten the tooltip.
The tag-strip block was removed in the parent commit but two surrounding
comments still referenced "tags without opt-in" and "runs AFTER the
strip". Update them to describe the remaining user_api_key_* and
_pipeline_managed_guardrails strip that the snapshot/merge ordering
actually protects against.
Caller-supplied tags (`x-litellm-tags` header, body `tags`, `metadata.tags`)
were silently dropped unless the key/team had
`metadata.allow_client_tags: true` set. Restore the documented behavior:
tags from the request always flow into `metadata.tags` and union with any
admin-configured static tags from key/team/project metadata.
Removes the `allow_client_tags` opt-in flag from the pre-call pipeline.
The flag was only ever read here; it has no schema or endpoint footprint,
so leftover values in existing key metadata are inert.
Test cleanup mirrors the simplification: drop the three tests that
verified the strip-when-not-opted-in path, drop the `allow_client_tags`
fixture lines from the merge/union tests.
Second wave of failures from the 2026-05-12 DALL-E shutdown:
- tests/image_gen_tests/test_image_edits.py::TestOpenAIImageEditDallE2
and tests/image_gen_tests/test_image_generation.py::TestOpenAIDalle3
are explicitly named for the deprecated models and can't pass; remove.
gpt-image-1 coverage already exists in sibling classes.
- tests/local_testing/test_router.py image gen tests use dall-e-3 only
as a routing example; swap to gpt-image-1.
- tests/local_testing/test_custom_callback_input.py image_generation
success/failure paths swapped to gpt-image-1.
DALL-E 2 and DALL-E 3 were removed from the OpenAI API on 2026-05-12,
causing e2e image-generation tests to fail with "model does not exist".
Swap all live-API DALL-E references in proxy-backed tests to gpt-image-1
and update the dall-e-2 alias in proxy_server_config.yaml to point at
openai/gpt-image-1 (preserves any historical dall-e-2 callers).
* chore: reject bare str at file-input sinks to prevent local-file read (#27667)
Squash-merged by litellm-agent from stuxf's PR.
* fix: use os.PathLike in ocr sink and check truthy reasoningSummary for bridge
- ocr/main.py: widen Path check to os.PathLike for consistency with other sinks
- main.py: bridge condition checks truthiness of reasoning_summary, not just None
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix: remove unused pathlib.Path import in ocr/main.py
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Co-authored-by: stuxf <70670632+stuxf@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Second wave of failures from the 2026-05-12 DALL-E shutdown:
- tests/image_gen_tests/test_image_edits.py::TestOpenAIImageEditDallE2
and tests/image_gen_tests/test_image_generation.py::TestOpenAIDalle3
are explicitly named for the deprecated models and can't pass; remove.
gpt-image-1 coverage already exists in sibling classes.
- tests/local_testing/test_router.py image gen tests use dall-e-3 only
as a routing example; swap to gpt-image-1.
- tests/local_testing/test_custom_callback_input.py image_generation
success/failure paths swapped to gpt-image-1.
DALL-E 2 and DALL-E 3 were removed from the OpenAI API on 2026-05-12,
causing e2e image-generation tests to fail with "model does not exist".
Swap all live-API DALL-E references in proxy-backed tests to gpt-image-1
and update the dall-e-2 alias in proxy_server_config.yaml to point at
openai/gpt-image-1 (preserves any historical dall-e-2 callers).
The tag-strip block was removed in the parent commit but two surrounding
comments still referenced "tags without opt-in" and "runs AFTER the
strip". Update them to describe the remaining user_api_key_* and
_pipeline_managed_guardrails strip that the snapshot/merge ordering
actually protects against.
Caller-supplied tags (`x-litellm-tags` header, body `tags`, `metadata.tags`)
were silently dropped unless the key/team had
`metadata.allow_client_tags: true` set. Restore the documented behavior:
tags from the request always flow into `metadata.tags` and union with any
admin-configured static tags from key/team/project metadata.
Removes the `allow_client_tags` opt-in flag from the pre-call pipeline.
The flag was only ever read here; it has no schema or endpoint footprint,
so leftover values in existing key metadata are inert.
Test cleanup mirrors the simplification: drop the three tests that
verified the strip-when-not-opted-in path, drop the `allow_client_tags`
fixture lines from the merge/union tests.
The enterprise package is installed as `litellm_enterprise` (per
enterprise/pyproject.toml), but several tests imported it as
`enterprise.litellm_enterprise.*` — a path that only resolves
because the repo root happens to sit on sys.path, letting Python's
implicit namespace package machinery discover `enterprise/` as a
directory.
This breaks any test runner that relocates source (e.g. the
mutation-testing workflow, which copies tests under `mutants/`) and
also caused two `patch()` strings to target a module path that does
not match what production code imports — meaning those mocks were
never actually patching the production module's attribute.
Replace `from enterprise.litellm_enterprise.` with the canonical
`from litellm_enterprise.` across 6 test files, and fix two
`patch()` target strings (and one `sys.modules` patch key in the
SSO test) to match.
* feat(ui): search teams by team ID alongside name
The Teams page search box only matched team_alias, so pasting a team UUID
returned zero results. Detect when the input is a full UUID and route it
to the team_id filter instead; otherwise keep the existing alias substring
search. Placeholder now reads "Search teams by name or ID...".
Resolves LIT-2648
* refactor: search teams via backend OR clause, drop client-side UUID detection
Adds a `search` query param to /v2/team/list that ORs across team_id
(exact) and team_alias (case-insensitive contains), so the search box
sends one param regardless of input format. Removes the isLikelyTeamId
helper and the client-side branching it fed.
* fix(ui): resolve created_by to human-readable name for team keys
The team's Virtual Keys table rendered a raw UUID under Created By
for keys created on behalf of a real user, and the key details page
header showed "-" because it was reading the wrong field
(user_email/user_id instead of created_by_user).
- TeamVirtualKeysTable Created By column now prefers
created_by_user.user_alias > user_email > UUID, with a Popover
on hover that surfaces all three fields with copy icons
(matches the existing AllKeys table pattern in VirtualKeysTable.tsx)
- key_info_view passes the resolved created_by_user value to the
details-page header so it renders the readable name
Resolves LIT-2517
* fix(ui): add created_by to KeyResponse type
The backend returns created_by on every key row, but the frontend
type omitted it. The Created By cell already reads the field via
info.row.original, which trips the production typecheck.
* feat(ui): add Expires to key Overview header; merge User into one field
The key details page header omitted Expires (only Settings tab had it)
and showed User Email + User ID as two separate rows. This PR:
- adds Expires below Created At, reusing the Settings tab's
formatTimestamp(...) ?? "Never" formatting so both views agree
- merges User Email / User ID into a single "User" field that
displays alias / email / user_id (in that fallback order) with a
Popover on hover exposing all three with copy icons — mirrors the
Created By pattern in TeamVirtualKeysTable.tsx and VirtualKeysTable.tsx
Refs LIT-2517
* chore(ui): User field icon + alias-primary test
Address review feedback on #27696:
- Swap MailOutlined → UserOutlined on the merged User cell; the
envelope icon implied "this is an email" but the cell can render
alias / email / user_id depending on what's available.
- Add a test asserting userAlias displays as primary and overrides
userEmail — closes the gap where a fallback-order regression
would have silently passed.
* fix(ui): truncate long User values and reshuffle header layout
- Swap column groupings so Created By stays paired with Created At
(matches the pre-merge layout); User and Expires now share col 1.
- Ellipsis-truncate the visible User cell at maxWidth 200 so a raw
UUID fallback doesn't sprawl. Full identity still revealed via the
hover Popover.
- Cap each Popover row at maxWidth 220 so a UUID truncates inside the
panel too; antd's built-in ellipsis tooltip surfaces the full value.