Commit graph

34478 commits

Author SHA1 Message Date
Ishaan Jaffer
dfac66ed51 fix misleading comment on user_api_key_cache TTL line 2026-03-05 16:01:48 -08:00
Ishaan Jaffer
bf27a497c3 address greptile review feedback (greploop iteration 12) 2026-03-05 15:48:51 -08:00
Ishaan Jaffer
f953e171e1 address greptile review feedback (greploop iteration 11) 2026-03-05 15:29:14 -08:00
Ishaan Jaffer
1f503fdcbd address greptile review feedback (greploop iteration 10) 2026-03-05 15:15:47 -08:00
Ishaan Jaffer
de15b54c00 address greptile review feedback (greploop iteration 9) 2026-03-05 15:02:23 -08:00
Ishaan Jaffer
d8906d33f7 address greptile review feedback (greploop iteration 8) 2026-03-05 14:50:32 -08:00
Ishaan Jaffer
27f766424d address greptile review feedback (greploop iteration 7) 2026-03-05 14:39:23 -08:00
Ishaan Jaffer
6001a0b30c address greptile review feedback (greploop iteration 6) 2026-03-05 14:29:08 -08:00
Ishaan Jaffer
03a39c50d2 simplify _get_pkce_userinfo: remove shared-client complexity, use async with directly 2026-03-05 14:23:03 -08:00
Ishaan Jaffer
530627ab78 address greptile review feedback (greploop iteration 5) 2026-03-05 14:14:10 -08:00
Ishaan Jaffer
8a667f096d address greptile review feedback (greploop iteration 4) 2026-03-05 14:02:55 -08:00
Ishaan Jaffer
9161253d6a address greptile review feedback (greploop iteration 3) 2026-03-05 13:55:01 -08:00
Ishaan Jaffer
cc1f8c23f2 sanitize PKCE cache log to not expose verifier content 2026-03-05 13:45:39 -08:00
Ishaan Jaffer
c8dcf45fc9 fix remaining PKCE test assertion for dict-format verifier storage 2026-03-05 13:45:28 -08:00
Ishaan Jaffer
1a16dcf598 fix: address fourth round of greptile review feedback
- Strip OAuth token credentials from response_convertor input to prevent
  access_token/id_token appearing in restricted-group error messages
- Reuse single httpx.AsyncClient for both token exchange and userinfo requests
  to avoid a second TCP/TLS handshake per SSO callback
- Revert Redis wiring to user_api_key_cache: PKCE code already uses
  redis_usage_cache directly; wiring would route all API-key lookups through
  Redis unnecessarily. Add startup warning instead when PKCE+Redis mismatch.
- Move _OAUTH_TOKEN_FIELDS to module level
2026-03-05 12:24:44 -08:00
Ishaan Jaffer
c6f2446fa5 fix: address third round of greptile review feedback
- Fix CRITICAL log firing on every non-PKCE callback: only log when PKCE is enabled
- Remove unused pkce_env_value intermediate variable
- Prefer reusing redis_usage_cache over creating separate RedisCache instance
  (avoids losing advanced connection options like SSL, timeouts, db)
2026-03-05 12:11:05 -08:00
Ishaan Jaffer
50631f8959 fix: address second round of greptile review feedback
- Fix PKCE error hint: check env var directly (not code_verifier presence) to
  distinguish 'PKCE not configured' from 'PKCE enabled but cache miss'
- Fix misleading Redis TTL comment in proxy_server.py
2026-03-05 12:04:01 -08:00
Ishaan Jaffer
c94886d2b3 fix: address greptile review feedback
- Fix access_token missing in PKCE path: read from combined_response directly
  instead of generic_sso.access_token (which is only set by verify_and_process)
- Fix PKCE error hint firing when PKCE is already enabled: only show
  'set GENERIC_CLIENT_USE_PKCE=true' advice when code_verifier was absent
- Fix unguarded KeyError on access_token: check for error field in HTTP 200
  responses before accessing token_response['access_token']
- Fix silent empty userinfo: raise ProxyException when both userinfo endpoint
  and id_token fallback produce no user data
- Fix backward-incompatible Redis wiring: only attach Redis to user_api_key_cache
  when GENERIC_CLIENT_USE_PKCE=true, preserving existing in-memory behaviour
2026-03-05 11:56:19 -08:00
Ishaan Jaffer
da82ba4dc3 refactor(sso): extract PKCE token exchange into SSOAuthenticationHandler methods
- Move import httpx/jwt to module level (top of file, not inside function)
- Extract inline PKCE token exchange + userinfo logic into two static methods:
  _pkce_token_exchange() and _get_pkce_userinfo()
- get_generic_sso_response PKCE path is now a single method call
- Fix double-logging in except block for non-PKCE errors
- Use %-style log formatting (no f-strings in log calls)
2026-03-05 11:37:00 -08:00
Ishaan Jaffer
08e249b6c7 fix(sso): add direct PKCE token exchange and Redis cache wiring for multi-instance SSO
When PKCE is enabled, bypass fastapi-sso and perform direct token exchange so
code_verifier is correctly included. Store PKCE verifiers as dict in cache
for proper JSON serialization in Redis. Wire user_api_key_cache to Redis when
available so PKCE verifiers are shared across ECS tasks/pods.

Also adds clearer error messages when PKCE is required but not configured.
2026-03-05 11:31:48 -08:00
Sameer Kankute
bf9c96b912
Merge pull request #22679 from giulio-leone/fix/websearch-thinking-constraint
fix: WebSearch interception fails with thinking enabled + SpendLog dedup
2026-03-06 00:49:17 +05:30
Sameer Kankute
728e5b13f7
Merge pull request #22922 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:43:28 +05:30
Sameer Kankute
f06e9e6368 Fix doc 2026-03-06 00:42:45 +05:30
Sameer Kankute
7aff1dc0d3
Merge pull request #22919 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:26:34 +05:30
Sameer Kankute
04f38332de Fix doc 2026-03-06 00:25:31 +05:30
Spencer Burridge
c919031ff0
feat(proxy): include user_email in jwt upsert user creation (#22915)
* Include user_email in new user creation within get_user_object

Enhance the get_user_object function to include user_email in the parameters when creating a new user. This change is accompanied by a new test to verify that user_email is correctly included during the upsert process.

* Improve error handling in test_get_user_object by logging exceptions

Updated the test_get_user_object_upsert_includes_user_email function to log exceptions when they occur, enhancing the visibility of potential issues during testing. This change helps in diagnosing failures related to the mock LiteLLM_UserTable.
2026-03-05 10:55:11 -08:00
Sameer Kankute
46fa9a33da
Merge pull request #22918 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:53:34 +05:30
Sameer Kankute
cae1f5fbae Fix doc 2026-03-05 23:52:56 +05:30
Sameer Kankute
9df686e044
Merge pull request #22917 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:51:00 +05:30
Sameer Kankute
cf376d2c0e Fix doc 2026-03-05 23:50:20 +05:30
Ishaan Jaff
a42132f329
fix(passthrough): propagate Azure 429/5xx errors in async streaming instead of silent HTTP 200 (#22913)
* fix(passthrough): raise_for_status in _async_streaming to propagate Azure 429s

* address greptile review feedback (greploop iteration 1)

Guard data/json args when content is provided to avoid httpx ValueError

* address greptile review feedback (greploop iteration 2)

Use bare raise to preserve original traceback in _async_streaming exception handler

* address greptile review feedback (greploop iteration 3)

Close httpx streaming response on error to prevent connection pool exhaustion

* address greptile review feedback (greploop iteration 4)

Guard aclose() call to prevent masking original exception; add explicit test for content param forwarding

* address greptile review feedback (greploop iteration 5)

Pass content to sign_request so AWS body-hash signing is correct when content is the sole body source

* revert sign_request content change - request_data expects dict, not bytes

Bedrock's sign_request calls json.dumps(request_data) — passing content bytes
would TypeError. sign_request should only receive data/json (dict), not raw bytes.
2026-03-05 10:12:43 -08:00
Sameer Kankute
8dca085640
Merge pull request #22916 from BerriAI/litellm_gpt-5.4_day_0
Add day 0 support for gpt-5.4
2026-03-05 23:41:35 +05:30
Sameer Kankute
3b457b5d8e Add day 0 support for gpt-5.4 2026-03-05 23:40:24 +05:30
giulio-leone
d6310ff36e fix: downgrade WebSearch logs from info to debug to reduce production noise
All operational/diagnostic messages in WebSearchInterceptionLogger are now
debug-level to avoid flooding production logs while still remaining available
when verbose logging is enabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:58:40 +01:00
Sameer Kankute
b9a8d42882 Add day 0 support for gpt-5.4 2026-03-05 23:26:24 +05:30
giulio-leone
7b0ed0ff91 fix: replace sk-fake with safe test key to avoid secret scanner
Replace 'sk-fake' with 'fake-key-for-testing' in websearch interception
tests to prevent false-positive secret scanner triggers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:29:28 +01:00
Julio Quinteros Pro
6db3f2f668
Merge pull request #22892 from BerriAI/fix/greptile-type-safety-improvements
fix(types): address type-safety issues from mypy PR review
2026-03-05 14:06:17 -03:00
Julio Quinteros
023794ba62 fix(merge): resolve conflict with main in cost_tracking_settings
PR #22890 used cast(str, ...) / cast(Optional[str], ...) for the return
statements; this PR's approach uses str() for explicit runtime coercion
(addressing Greptile's concern). Keep the str() version and drop the
now-unused cast import.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 14:01:07 -03:00
Giulio Leone
6b7d767637
feat(anthropic): support top-level cache_control for automatic prompt caching (#22442) 2026-03-05 08:34:56 -08:00
giulio-leone
660de94493 fix: change all verbose_logger.warning → info in websearch handler
Per Sameerlite's review: warning-level logs trigger Slack alerts.
All 6 remaining .warning() calls were operational/fallback messages,
not actual errors. Changed to .info() to match the first fix at L510.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 17:14:30 +01:00
giulio-leone
b04ba60e6e fix(websearch): downgrade max_tokens adjustment log level
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 14:19:15 +01:00
Sameer Kankute
5183a6e850
Merge pull request #22866 from mubashir1osmani/feat/bedrock-mantle-provider-clean
feat: bedrock mantle provider
2026-03-05 18:24:00 +05:30
Sameer Kankute
a282bf9726
Merge pull request #22893 from BerriAI/litellm_messages-to-responses-mapping-docs
docs(anthropic): add v1/messages → /responses parameter mapping reference
2026-03-05 18:21:37 +05:30
Sameer Kankute
0620f99fa4
Merge pull request #22867 from BerriAI/litellm_bedrock-azure-cache-control-scope
fix(bedrock,azure_ai): strip scope from cache_control for Anthropic messages
2026-03-05 18:20:59 +05:30
Sameer Kankute
c04c120df2
Merge pull request #22884 from BerriAI/litellm_vertex-output-config-drop
fix(vertex_ai): drop unsupported output_config parameter from all requests
2026-03-05 18:20:47 +05:30
Sameer Kankute
4fda3e8351
Merge pull request #22896 from BerriAI/litellm_mistral-document-ai-2512-cost-map
feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
2026-03-05 16:14:27 +05:30
Sameer Kankute
bb1297fe1b feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
Made-with: Cursor
2026-03-05 16:07:46 +05:30
Sameer Kankute
6a8adf8bdf docs(anthropic): add v1/messages → /responses parameter mapping reference
Documents exactly how every request and response field gets translated
when LiteLLM routes an Anthropic /v1/messages call through the OpenAI
Responses API path (for OpenAI/Azure targets). Covers messages content
block mapping, tools, tool_choice, thinking→reasoning, context_management,
and the reverse response translation. Wired into the /v1/messages sidebar.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 15:51:19 +05:30
Julio Quinteros Pro
b52bd43740
Merge pull request #22822 from BerriAI/fix/plr0915-too-many-statements
fix(lint): resolve PLR0915 too-many-statements in 4 files
2026-03-05 07:13:10 -03:00
Julio Quinteros
c1076de5bd fix(types): address type-safety issues from mypy PR review
- CreateBatchRequest.output_expires_after: drop Optional since total=False
  already makes the key absent-or-present; Optional[T] incorrectly allowed
  the key to exist with value None, which is incompatible with the OpenAI
  SDK's OutputExpiresAfter | NotGiven expectation on batches.create()
- cost_tracking_settings._resolve_model_for_cost_lookup: replace implicit
  object-to-str returns with explicit str() calls so the function is safe
  even if the surrounding truthiness guards are later weakened

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 07:11:31 -03:00