litellm/tests/test_litellm/proxy
Ishaan Jaff 7697b1c397
fix(sso): direct PKCE token exchange + Redis wiring for multi-instance SSO (#22923)
* fix(sso): add direct PKCE token exchange and Redis cache wiring for multi-instance SSO

When PKCE is enabled, bypass fastapi-sso and perform direct token exchange so
code_verifier is correctly included. Store PKCE verifiers as dict in cache
for proper JSON serialization in Redis. Wire user_api_key_cache to Redis when
available so PKCE verifiers are shared across ECS tasks/pods.

Also adds clearer error messages when PKCE is required but not configured.

* refactor(sso): extract PKCE token exchange into SSOAuthenticationHandler methods

- Move import httpx/jwt to module level (top of file, not inside function)
- Extract inline PKCE token exchange + userinfo logic into two static methods:
  _pkce_token_exchange() and _get_pkce_userinfo()
- get_generic_sso_response PKCE path is now a single method call
- Fix double-logging in except block for non-PKCE errors
- Use %-style log formatting (no f-strings in log calls)

* fix: address greptile review feedback

- Fix access_token missing in PKCE path: read from combined_response directly
  instead of generic_sso.access_token (which is only set by verify_and_process)
- Fix PKCE error hint firing when PKCE is already enabled: only show
  'set GENERIC_CLIENT_USE_PKCE=true' advice when code_verifier was absent
- Fix unguarded KeyError on access_token: check for error field in HTTP 200
  responses before accessing token_response['access_token']
- Fix silent empty userinfo: raise ProxyException when both userinfo endpoint
  and id_token fallback produce no user data
- Fix backward-incompatible Redis wiring: only attach Redis to user_api_key_cache
  when GENERIC_CLIENT_USE_PKCE=true, preserving existing in-memory behaviour

* fix: address second round of greptile review feedback

- Fix PKCE error hint: check env var directly (not code_verifier presence) to
  distinguish 'PKCE not configured' from 'PKCE enabled but cache miss'
- Fix misleading Redis TTL comment in proxy_server.py

* fix: address third round of greptile review feedback

- Fix CRITICAL log firing on every non-PKCE callback: only log when PKCE is enabled
- Remove unused pkce_env_value intermediate variable
- Prefer reusing redis_usage_cache over creating separate RedisCache instance
  (avoids losing advanced connection options like SSL, timeouts, db)

* fix: address fourth round of greptile review feedback

- Strip OAuth token credentials from response_convertor input to prevent
  access_token/id_token appearing in restricted-group error messages
- Reuse single httpx.AsyncClient for both token exchange and userinfo requests
  to avoid a second TCP/TLS handshake per SSO callback
- Revert Redis wiring to user_api_key_cache: PKCE code already uses
  redis_usage_cache directly; wiring would route all API-key lookups through
  Redis unnecessarily. Add startup warning instead when PKCE+Redis mismatch.
- Move _OAUTH_TOKEN_FIELDS to module level

* fix remaining PKCE test assertion for dict-format verifier storage

* sanitize PKCE cache log to not expose verifier content

* address greptile review feedback (greploop iteration 3)

* address greptile review feedback (greploop iteration 4)

* address greptile review feedback (greploop iteration 5)

* simplify _get_pkce_userinfo: remove shared-client complexity, use async with directly

* address greptile review feedback (greploop iteration 6)

* address greptile review feedback (greploop iteration 7)

* address greptile review feedback (greploop iteration 8)

* address greptile review feedback (greploop iteration 9)

* address greptile review feedback (greploop iteration 10)

* address greptile review feedback (greploop iteration 11)

* address greptile review feedback (greploop iteration 12)

* fix misleading comment on user_api_key_cache TTL line

* address greptile review feedback (greploop iteration 13)

* address greptile review feedback (greploop iteration 14)

* address greptile review feedback (greploop iteration 15)

* address greptile review feedback (greploop iteration 16)

* address greptile review feedback (greploop iteration 17)

* address greptile review feedback (greploop iteration 18)

* address greptile review feedback (greploop iteration 19)

* address greptile review feedback (greploop iteration 20)

* address greptile review feedback (greploop iteration 21)

* address greptile review feedback (greploop iteration 22)

* address greptile review feedback (greploop iteration 23)

* address greptile review feedback (greploop iteration 24)

* address greptile review feedback (greploop iteration 25)

* address greptile review feedback (greploop iteration 26)

* address greptile review feedback (greploop iteration 27)

* address greptile review feedback (greploop iteration 28)

* address greptile review feedback (greploop iteration 29)

* address greptile review feedback (greploop iteration 30)

* address greptile review feedback (greploop iteration 31)

* address greptile review feedback (greploop iteration 32)

* address greptile review feedback (greploop iteration 33)

* address greptile review feedback (greploop iteration 34)

* address greptile review feedback (greploop iteration 35)

* address greptile review feedback (greploop iteration 37)

- read GENERIC_CLIENT_USE_PKCE env var once in prepare_token_exchange_parameters
- include actual decode error in jwt.decode failure exception message
- add GENERIC_CLIENT_USE_PKCE=true to no-state regression test

* defer PKCE verifier deletion until after all downstream processing

Move _delete_pkce_verifier to after response_convertor and
process_sso_jwt_access_token complete. If JWT processing raises,
the verifier stays in cache so the user can retry without restarting
the full OAuth flow.

* address greptile review feedback (greploop iteration 38)

- fix strict-mode cache miss error message to differentiate
  cross-instance routing failures (Redis configured) from single-instance
  issues (TTL expiry, pod restart) when only in-memory cache is available
- add comment above _get_pkce_userinfo call explaining that bearer
  credentials are always sourced from token_response in the merge step

* fix null JSON response body in _pkce_token_exchange

- Guard against HTTP 200 with body null: response.json() returns None
  for JSON null, and calling .get() on None raises AttributeError.
  Now raises a clean ProxyException with a clear error message.
- Fix misleading userinfo warning: was always saying "empty dict" but
  also fires for JSON null responses; updated to say "empty or null".
- Add HTTP status code assertion to cache miss test.

* address greptile review feedback (greploop iteration 39)

- fix credential leakage: directly assign received_response from
  combined_response instead of relying on nonlocal mutation; Pyright
  was flagging the old guard as unreachable, meaning credential stripping
  might not execute — now it always runs unconditionally
- add test for legacy plain-string cache format backward compat branch
- add test for HTTP 200 with no error field and no access_token (else branch)
- add test for HTTP 200 with JSON null body (new AttributeError guard)

* fix _OAUTH_TOKEN_FIELDS merge loop to preserve userinfo values on absent fields

When the token endpoint omits a bearer-credential field entirely (field
absent from token_response), the previous code deleted it from merged even
if userinfo provided a valid value. Now:
- non-null in token_response → restore authoritative token endpoint value
- explicit null in token_response → remove key from merged (clean absence)
- field absent from token_response → leave userinfo value unchanged

* use HTTP 401 for PKCE missing config errors

GENERIC_CLIENT_ID and GENERIC_TOKEN_ENDPOINT missing when PKCE is
enabled are auth-flow failures, not server errors. Use 401 instead
of 500 to avoid triggering false-positive server error alerts in
monitoring systems.

* address greptile review feedback (greploop iteration 40)

- fix duplicate error logging: demote first format-error log to DEBUG
  so the detailed ERROR in strict-mode branch is not duplicated
- add HTTP status code assertions to all PKCE ProxyException tests
  for better regression protection against accidental code changes

* add credential absence assertions to test_pkce_token_exchange_basic_auth

Verify that client_id and client_secret are NOT double-sent in the POST
body when Basic Auth is used (include_client_id=False with client_secret).
Catches regressions where credentials leak into both Auth header and body.

* address greptile review feedback (greploop iteration 41)

- add Bearer token header assertion to test_pkce_token_exchange_credentials_in_body
- add cache query assertions to both non-strict mode tests to confirm
  the cache was accessed before the warning path triggers

* address greptile review feedback (greploop iteration 42)

- assert null id_token is absent from merged result in basic auth test
- add test for HTTP 200 empty/null userinfo body with no id_token fallback

* use caplog to verify warning logs in non-strict cache miss tests

The two non-strict mode tests now use pytest's caplog fixture to assert
that a warning is actually emitted, not just that the code continues
without raising. This catches regressions where the warning silently
disappears.

* remove dead-code response=None guard in _pkce_token_exchange

* clean up stale pkce verifier cache entries in non-strict mode

* fix test: configure async_delete_cache as AsyncMock and assert cleanup called

* add sentinel guard so pkce-no-redis warning only fires once across hot-reloads

* fix misleading comments: code_verifier init and bearer-credential merge docs

* add best-effort cleanup in strict-mode for corrupt/empty cache entries

* add redirect_uri assertion, userinfo body in non-200 log, sentinel comment
2026-03-12 12:41:39 -07:00
..
_experimental/mcp_server merge: resolve conflicts with upstream staging (bedrock + mcp tests) 2026-03-12 13:40:16 -03:00
agent_endpoints fix(tests): restore litellm_params=None on mock agent in a2a invoke test (#23125) 2026-03-09 07:16:02 -07:00
anthropic_endpoints [Fix] 404 Not Found on /api/event_logging/batch endpoint (#20504) 2026-02-05 10:58:08 -08:00
auth merge: resolve conflicts with upstream staging (bedrock + mcp tests) 2026-03-12 13:40:16 -03:00
client add a new feature fix to expose the team alias when authenticating th… (#17725) 2025-12-10 10:10:28 -08:00
common_utils Merge pull request #20688 from BerriAI/litellm_budget_tier_enforcement_for_keys 2026-03-06 20:44:58 -08:00
db Revert "feat(proxy): add Prisma DB pool and engine health metrics to Promethe…" 2026-03-09 14:55:11 -07:00
discovery_endpoints Support auto_redirect_ui_login_to_sso in config.yaml general_settings 2026-03-11 11:34:17 -07:00
experimental/mcp_server
google_endpoints fix: Metadata / Trace ID Missing in S3 Streaming Callbacks 2026-02-25 14:16:42 +05:30
guardrails merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
health_endpoints Prisma DB Failure Detection and Self-Healing (#21059) 2026-03-05 13:44:49 -08:00
hooks fix(proxy): make async_post_call_response_headers_hook consistent across all endpoints (#22985) 2026-03-12 08:51:00 -07:00
image_endpoints fixing core proxy tests 2026-02-12 17:54:32 -08:00
management_endpoints fix(sso): direct PKCE token exchange + Redis wiring for multi-instance SSO (#22923) 2026-03-12 12:41:39 -07:00
management_helpers [Docs] Fix "Page Not Found" link for Anthropic endpoint (#23349) 2026-03-11 20:17:41 +05:30
middleware feat: add in_flight_requests metric to /health/backlog + prometheus (#22319) 2026-02-27 18:00:50 -08:00
openai_files_endpoint CircleCI test stability (#23055) 2026-03-07 15:19:39 -08:00
pass_through_endpoints Merge branch 'main' into litellm_oss_staging_03_11_2026 2026-03-12 16:21:28 -03:00
policy_engine Guardrail Policy Versioning (#21862) 2026-02-21 20:14:31 -08:00
prompts fix(prompts): fix prompt info lookup and delete using correct IDs (#19358) 2026-01-20 12:28:34 -08:00
public_endpoints [Feature] Add /public/endpoints endpoint for provider endpoint support 2026-02-26 18:17:37 -08:00
rag_endpoints tests and route permissions (#21508) 2026-02-18 16:58:38 -08:00
realtime_endpoints address greptile review feedback (greploop iteration 2) 2026-03-12 18:53:22 +05:30
response_api_endpoints Fix x-litellm-key-spend update 2025-12-12 11:44:51 +05:30
spend_tracking merge: resolve conflicts with upstream staging (bedrock + mcp tests) 2026-03-12 13:40:16 -03:00
test_configs
ui_crud_endpoints [Fix] PATCH /update/ui_settings now merges with existing record instead of overwriting 2026-03-05 15:49:40 -08:00
vector_store_endpoints test_delete_vector_store_checks_access 2026-01-31 12:05:09 -08:00
__init__.py test fix 2025-10-17 10:46:42 -07:00
conftest.py Add health endpoint tests to CI with database and Redis support (#17877) 2025-12-12 07:35:50 -08:00
test_aiohttp_cleanup_closed.py fix(aiohttp): only set enable_cleanup_closed when required (#21897) 2026-02-23 21:06:29 -08:00
test_api_key_masking_in_errors.py fix: mask API keys in error responses for invalid/malformed keys (#20289) 2026-02-12 19:58:05 +05:30
test_audio_speech_prometheus_hooks.py fix req changes 2026-02-28 21:32:57 +05:30
test_batch_expiry.py fix(proxy): improve team expiry enforcement validation 2026-03-03 17:29:39 -08:00
test_batch_metadata_none_fix.py Fix issue #13995: Handle None metadata in batch requests (#13996) 2025-08-27 14:51:09 -07:00
test_caching_routes.py [Bug Fix] Ensure /redis/info works on GCP Redis (#11732) 2025-06-14 15:35:09 -07:00
test_chat_completion_metadata.py fix: propagate JWT auth metadata to OTEL spans (#19627) 2026-01-23 21:21:23 -08:00
test_common_request_processing.py Fix azure model router 2026-03-12 12:40:37 +05:30
test_custom_proxy.py fix(ui/): fix routing for custom server root path (#15701) 2025-10-23 13:59:29 -07:00
test_empty_model_list.py [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
test_enforce_user_param.py Enforce support of enforce_user_param to openai post endpoints 2025-12-03 12:19:21 +05:30
test_fallback_management_endpoints.py Add fallback endpoints support 2026-01-16 10:51:33 +05:30
test_fastapi_offline_routes.py [Bug Fix] - Get Routes (#13466) 2025-08-09 12:52:23 -07:00
test_health_check_functions.py Fix_mapped tests part 2 2026-02-26 12:43:39 +05:30
test_health_check_max_tokens.py add docs and formatting 2026-02-28 14:08:09 +05:30
test_litellm_pre_call_utils.py Bug Fix: auto-inject prompt caching support for Gemini models (#21881) 2026-03-03 20:25:35 -08:00
test_model_dump_with_preserved_fields.py Fix_mapped tests part 2 2026-02-26 12:43:39 +05:30
test_model_id_header_propagation.py (fix) propagate x-litellm-model-id in responses (#16986) 2025-11-24 20:40:43 -08:00
test_openapi_schema_validation.py Reapply "feat: add model_cost aliases expansion support" 2026-03-12 13:36:57 -03:00
test_prometheus_cleanup.py Add Prometheus child_exit cleanup for gunicorn workers 2026-02-27 16:11:15 -08:00
test_proxy_cli.py merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
test_proxy_server.py Fix inflight mypy 2026-03-02 19:41:32 +05:30
test_proxy_types.py fix: Add PROXY_ADMIN role to system user for key rotation (#21896) 2026-02-27 19:11:29 -08:00
test_proxy_utils.py [Feat[ extends OAuth2 M2M authentication support to info routes (/key/info, /team/info, /user/info, /model/info) (#22713) 2026-03-06 17:29:25 -08:00
test_pyroscope.py Fix CI/CD pyroscope test failure (#21219) 2026-02-14 12:07:20 -08:00
test_response_model_sanitization.py feat(azure_ai): show actual model used in Azure Model Router response 2026-03-12 11:41:19 +05:30
test_route_a2a_models.py Fix test_route_a2a_model_bypasses_router 2026-02-05 09:47:05 +05:30
test_route_llm_request.py Override router settings 2026-01-31 16:04:52 -08:00
test_shared_health_check.py Fix_mapped tests part 2 2026-02-26 12:43:39 +05:30
test_spend_log_cleanup.py Fix spend log cleanup: lock tracking, integer retention, skip log level 2026-03-03 10:12:08 -08:00
test_swagger_chat_completions.py [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
test_team_member_update.py fix mapped tests (#12320) 2025-07-04 10:04:43 -07:00
test_tools_allowlist_enforcement.py Bug Fix: auto-inject prompt caching support for Gemini models (#21881) 2026-03-03 20:25:35 -08:00
test_update_llm_router_resilience.py fix(proxy): isolate get_config failures from model loading in sync loop 2026-02-26 17:49:44 -03:00