Commit graph

34463 commits

Author SHA1 Message Date
Yuneng Jiang
e33fc6167d
chore: fixes 2026-04-04 23:13:14 -07:00
Cursor Agent
b21140775d feat(spend-logs): add truncation note when error logs are truncated for DB storage
When the messages or response JSON fields in spend logs are truncated
before being written to the database, the truncation marker now includes
a note explaining:
- This is a DB storage safeguard
- Full, untruncated data is still sent to logging callbacks (OTEL, Datadog, etc.)
- The MAX_STRING_LENGTH_PROMPT_IN_DB env var can be used to increase the limit

Also emits a verbose_proxy_logger.info message when truncation occurs in
the request body or response spend log paths.

Adds 3 new tests:
- test_truncation_includes_db_safeguard_note
- test_response_truncation_logs_info_message
- test_request_body_truncation_logs_info_message

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-03-05 23:43:44 +00:00
weiguang li
3d027c0f7a
fix(bedrock): filter out custom field from tools to prevent 400 errors (#22861)
Claude Code v2.1.69+ sends `custom: {defer_loading: true}` on tool
definitions. Anthropic's API accepts this field, but Bedrock rejects it
with "Extra inputs are not permitted", causing ~90% of requests to fail.

Strip the `custom` field from each tool in the request body before
sending to Bedrock, in both the Messages API and Chat API invoke paths.

Fixes #22847

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
2026-03-05 13:54:23 -08:00
Curtis
725c0c158f
Prisma DB Failure Detection and Self-Healing (#21059)
* fix(proxy): readiness check returns 200 when database is unreachable

_db_health_readiness_check() catches health_check() exceptions but
never updates db_health_cache to "disconnected" and never re-raises.
The caller health_readiness() always returns 200 with "db": "connected"
hardcoded, regardless of actual DB state.

In Kubernetes, this means pods with dead database connections stay in
the Service endpoints and continue receiving traffic they cannot serve.

Changes:
- Set db_health_cache to "disconnected" and re-raise the exception on
  health_check failure so health_readiness() returns 503
- Use actual db_health_status["status"] in the response instead of
  hardcoding "db": "connected"
- Reduce cache TTL from 2 minutes to 15 seconds. The 2-minute window
  is too wide for readiness probes (typically 10-15s intervals) and
  means a pod can report healthy for up to 2 minutes after the DB dies
- Only serve cached results when status is "connected". The previous
  condition (status != "unknown") would also cache "disconnected" for
  2 minutes, delaying recovery detection after a DB comes back

* fix(proxy): add DB connection self-healing to readiness check

When the Prisma query engine's internal TCP connection pool holds dead
connections (caused by network blips, Cloud SQL proxy restarts, or
node-level issues), health_check() fails with httpx.ConnectError.
The engine never recovers on its own because nothing triggers a
disconnect/connect cycle to restart the subprocess with fresh
connections.

This leaves pods permanently failing readiness checks until they are
manually restarted, even after the underlying DB becomes reachable
again.

Add a reconnect attempt to _db_health_readiness_check() when
health_check() fails:
1. disconnect() - kills the query engine subprocess and closes all
   connections (has built-in backoff retry: 3 tries, 10s max)
2. connect() - starts a new engine with fresh TCP connections (has
   built-in backoff retry: 3 tries, 10s max)
3. health_check() - verifies the new connection works (has built-in
   backoff retry: 3 tries, 10s max)

If reconnect succeeds, the pod immediately returns to service (200).
If it fails, the original exception is re-raised (503). Reconnect
attempts are rate-limited by probe frequency (~10-15s), so a
permanently unreachable DB gets one attempt per cycle with no retry
loops.

This uses the same disconnect/connect mechanism that
PrismaWrapper.recreate_prisma_client() uses for IAM token refresh,
and aligns with the community-documented pattern for Prisma connection
recovery in long-running processes (prisma/prisma#24718, #27024).

* Add poetry lock and modify test_health_endpoints

* Address allow_requests_on_db_unavailable regression

* Address comments

* resolve greptile issue

* Restore accidentally deleted UI HTML files

These were removed in an earlier commit but still exist on main.
Restoring to keep the PR diff clean.

* Guard reconnect with is_database_transport_error

Only attempt disconnect/connect/health_check cycle for transport-level
failures (unreachable DB, dropped connection). Data-layer errors like
UniqueViolationError indicate the DB is reachable, so reconnecting
would be pointless churn.

* Address greptile's comments

* Fix module alias after rebase and add adversarial test coverage

- Unify module alias to _health_endpoints_module after rebase conflict
- Add test for non-transport error with flag on (exercises is_database_transport_error guard)
- Add test for disconnect() failure during reconnect cycle
- Split non-transport error test into flag-off (re-raises) and flag-on (skips reconnect) variants

* Remove stale UI HTML files reintroduced during rebase
2026-03-05 13:44:49 -08:00
Ishaan Jaff
503eb2fd4c
fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22925)
* fix: don't close HTTP/SDK clients on LLMClientCache eviction

Removing the _remove_key override that eagerly called aclose()/close()
on evicted clients. Evicted clients may still be held by in-flight
streaming requests; closing them causes:

  RuntimeError: Cannot send a request, as the client has been closed.

This is a regression from commit fb72979432. Clients that are no longer
referenced will be garbage-collected naturally. Explicit shutdown cleanup
happens via close_litellm_async_clients().

Fixes production crashes after the 1-hour cache TTL expires.

* test: update LLMClientCache unit tests for no-close-on-eviction behavior

Flip the assertions: evicted clients must NOT be closed. Replace
test_remove_key_closes_async_client → test_remove_key_does_not_close_async_client
and equivalents for sync/eviction paths.

Add test_remove_key_removes_plain_values for non-client cache entries.
Remove test_background_tasks_cleaned_up_after_completion (no more _background_tasks).
Remove test_remove_key_no_event_loop variant that depended on old behavior.

* test: add e2e tests for OpenAI SDK client surviving cache eviction

Add two new e2e tests using real AsyncOpenAI clients:
- test_evicted_openai_sdk_client_stays_usable: verifies size-based eviction
  doesn't close the client
- test_ttl_expired_openai_sdk_client_stays_usable: verifies TTL expiry
  eviction doesn't close the client

Both tests sleep after eviction so any create_task()-based close would
have time to run, making the regression detectable.

Also expand the module docstring to explain why the sleep is required.

* docs(AGENTS.md): add rule — never close HTTP/SDK clients on cache eviction

* docs(CLAUDE.md): add HTTP client cache safety guideline
2026-03-05 12:00:38 -08:00
Sameer Kankute
bf9c96b912
Merge pull request #22679 from giulio-leone/fix/websearch-thinking-constraint
fix: WebSearch interception fails with thinking enabled + SpendLog dedup
2026-03-06 00:49:17 +05:30
Sameer Kankute
728e5b13f7
Merge pull request #22922 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:43:28 +05:30
Sameer Kankute
f06e9e6368 Fix doc 2026-03-06 00:42:45 +05:30
Sameer Kankute
7aff1dc0d3
Merge pull request #22919 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:26:34 +05:30
Sameer Kankute
04f38332de Fix doc 2026-03-06 00:25:31 +05:30
Spencer Burridge
c919031ff0
feat(proxy): include user_email in jwt upsert user creation (#22915)
* Include user_email in new user creation within get_user_object

Enhance the get_user_object function to include user_email in the parameters when creating a new user. This change is accompanied by a new test to verify that user_email is correctly included during the upsert process.

* Improve error handling in test_get_user_object by logging exceptions

Updated the test_get_user_object_upsert_includes_user_email function to log exceptions when they occur, enhancing the visibility of potential issues during testing. This change helps in diagnosing failures related to the mock LiteLLM_UserTable.
2026-03-05 10:55:11 -08:00
Sameer Kankute
46fa9a33da
Merge pull request #22918 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:53:34 +05:30
Sameer Kankute
cae1f5fbae Fix doc 2026-03-05 23:52:56 +05:30
Sameer Kankute
9df686e044
Merge pull request #22917 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:51:00 +05:30
Sameer Kankute
cf376d2c0e Fix doc 2026-03-05 23:50:20 +05:30
Ishaan Jaff
a42132f329
fix(passthrough): propagate Azure 429/5xx errors in async streaming instead of silent HTTP 200 (#22913)
* fix(passthrough): raise_for_status in _async_streaming to propagate Azure 429s

* address greptile review feedback (greploop iteration 1)

Guard data/json args when content is provided to avoid httpx ValueError

* address greptile review feedback (greploop iteration 2)

Use bare raise to preserve original traceback in _async_streaming exception handler

* address greptile review feedback (greploop iteration 3)

Close httpx streaming response on error to prevent connection pool exhaustion

* address greptile review feedback (greploop iteration 4)

Guard aclose() call to prevent masking original exception; add explicit test for content param forwarding

* address greptile review feedback (greploop iteration 5)

Pass content to sign_request so AWS body-hash signing is correct when content is the sole body source

* revert sign_request content change - request_data expects dict, not bytes

Bedrock's sign_request calls json.dumps(request_data) — passing content bytes
would TypeError. sign_request should only receive data/json (dict), not raw bytes.
2026-03-05 10:12:43 -08:00
Sameer Kankute
8dca085640
Merge pull request #22916 from BerriAI/litellm_gpt-5.4_day_0
Add day 0 support for gpt-5.4
2026-03-05 23:41:35 +05:30
Sameer Kankute
3b457b5d8e Add day 0 support for gpt-5.4 2026-03-05 23:40:24 +05:30
giulio-leone
d6310ff36e fix: downgrade WebSearch logs from info to debug to reduce production noise
All operational/diagnostic messages in WebSearchInterceptionLogger are now
debug-level to avoid flooding production logs while still remaining available
when verbose logging is enabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:58:40 +01:00
Sameer Kankute
b9a8d42882 Add day 0 support for gpt-5.4 2026-03-05 23:26:24 +05:30
giulio-leone
7b0ed0ff91 fix: replace sk-fake with safe test key to avoid secret scanner
Replace 'sk-fake' with 'fake-key-for-testing' in websearch interception
tests to prevent false-positive secret scanner triggers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:29:28 +01:00
Julio Quinteros Pro
6db3f2f668
Merge pull request #22892 from BerriAI/fix/greptile-type-safety-improvements
fix(types): address type-safety issues from mypy PR review
2026-03-05 14:06:17 -03:00
Julio Quinteros
023794ba62 fix(merge): resolve conflict with main in cost_tracking_settings
PR #22890 used cast(str, ...) / cast(Optional[str], ...) for the return
statements; this PR's approach uses str() for explicit runtime coercion
(addressing Greptile's concern). Keep the str() version and drop the
now-unused cast import.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 14:01:07 -03:00
Giulio Leone
6b7d767637
feat(anthropic): support top-level cache_control for automatic prompt caching (#22442) 2026-03-05 08:34:56 -08:00
giulio-leone
660de94493 fix: change all verbose_logger.warning → info in websearch handler
Per Sameerlite's review: warning-level logs trigger Slack alerts.
All 6 remaining .warning() calls were operational/fallback messages,
not actual errors. Changed to .info() to match the first fix at L510.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 17:14:30 +01:00
giulio-leone
b04ba60e6e fix(websearch): downgrade max_tokens adjustment log level
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 14:19:15 +01:00
Sameer Kankute
5183a6e850
Merge pull request #22866 from mubashir1osmani/feat/bedrock-mantle-provider-clean
feat: bedrock mantle provider
2026-03-05 18:24:00 +05:30
Sameer Kankute
a282bf9726
Merge pull request #22893 from BerriAI/litellm_messages-to-responses-mapping-docs
docs(anthropic): add v1/messages → /responses parameter mapping reference
2026-03-05 18:21:37 +05:30
Sameer Kankute
0620f99fa4
Merge pull request #22867 from BerriAI/litellm_bedrock-azure-cache-control-scope
fix(bedrock,azure_ai): strip scope from cache_control for Anthropic messages
2026-03-05 18:20:59 +05:30
Sameer Kankute
c04c120df2
Merge pull request #22884 from BerriAI/litellm_vertex-output-config-drop
fix(vertex_ai): drop unsupported output_config parameter from all requests
2026-03-05 18:20:47 +05:30
Sameer Kankute
4fda3e8351
Merge pull request #22896 from BerriAI/litellm_mistral-document-ai-2512-cost-map
feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
2026-03-05 16:14:27 +05:30
Sameer Kankute
bb1297fe1b feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
Made-with: Cursor
2026-03-05 16:07:46 +05:30
Sameer Kankute
6a8adf8bdf docs(anthropic): add v1/messages → /responses parameter mapping reference
Documents exactly how every request and response field gets translated
when LiteLLM routes an Anthropic /v1/messages call through the OpenAI
Responses API path (for OpenAI/Azure targets). Covers messages content
block mapping, tools, tool_choice, thinking→reasoning, context_management,
and the reverse response translation. Wired into the /v1/messages sidebar.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 15:51:19 +05:30
Julio Quinteros Pro
b52bd43740
Merge pull request #22822 from BerriAI/fix/plr0915-too-many-statements
fix(lint): resolve PLR0915 too-many-statements in 4 files
2026-03-05 07:13:10 -03:00
Julio Quinteros
c1076de5bd fix(types): address type-safety issues from mypy PR review
- CreateBatchRequest.output_expires_after: drop Optional since total=False
  already makes the key absent-or-present; Optional[T] incorrectly allowed
  the key to exist with value None, which is incompatible with the OpenAI
  SDK's OutputExpiresAfter | NotGiven expectation on batches.create()
- cost_tracking_settings._resolve_model_for_cost_lookup: replace implicit
  object-to-str returns with explicit str() calls so the function is safe
  even if the surrounding truthiness guards are later weakened

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 07:11:31 -03:00
Julio Quinteros Pro
ad2969badb
Merge pull request #22890 from BerriAI/fix/mypy-type-errors
fix(mypy): resolve type errors across 9 files
2026-03-05 07:05:23 -03:00
Julio Quinteros Pro
44498da62a
Merge pull request #22887 from BerriAI/fix/schema-add-realtime-mode
fix(test): add 'realtime' to model mode enum in schema validation
2026-03-05 07:03:27 -03:00
Julio Quinteros Pro
de18b47f83
Merge pull request #22891 from BerriAI/fix/prisma-schema-duplicate-spec-path
fix(schema): remove duplicate spec_path field in LiteLLM_MCPServerTable
2026-03-05 07:02:33 -03:00
Julio Quinteros
16f415ad74 fix(schema): remove duplicate spec_path field in LiteLLM_MCPServerTable
PR #22850 (BYOK MCP servers) accidentally re-declared spec_path which was
already added by PR #22820, causing Prisma schema validation to fail with
error P1012 "Field is already defined".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 07:00:11 -03:00
Julio Quinteros
f45a9df52d fix(mypy): resolve type errors across 9 files
- batches/main.py: import FileExpiresAfter, cast output_expires_after on assignment
- openai/openai.py, azure/batches/handler.py: add # type: ignore[arg-type] on
  batches.create / batches.retrieve TypedDict unpacking calls
- searchapi/transformation.py: cast optional_params["country"] to str before .lower()
- openrouter/image_edit/transformation.py: cast iterated value to str for size/quality params
- spend_log_cleanup.py: narrow bool | None to bool with `or False`
- cost_tracking_settings.py: cast base_model/resolved_model to str and
  custom_llm_provider to Optional[str] in return statements
- text_moderation.py: suppress misc TypedDict ** expansion error; use cast for response
- prompt_shield.py: use cast instead of TypedDict(**response_json) construction

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:58:20 -03:00
Julio Quinteros Pro
1c7f93f90c
Update litellm/fine_tuning/main.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-05 06:45:31 -03:00
Julio Quinteros
db8e909ef2 fix(test): add 'realtime' to model mode enum in schema validation
gemini/gemini-live-2.5-flash-preview-native-audio-09-2025 uses mode='realtime'
but the schema in test_aaamodel_prices_and_context_window_json_is_valid did
not include 'realtime' as a valid enum value, causing a ValidationError.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:41:51 -03:00
Julio Quinteros
0133d11eb2 fix: replace assert with RuntimeError and fix return type annotation
- a2a_protocol/main.py: replace bare assert with descriptive RuntimeError
  in _execute_a2a_send_with_retry so retry exhaustion gives a clear message
- fine_tuning/main.py: fix _resolve_fine_tuning_timeout return type from
  float to Union[float, httpx.Timeout] to accurately reflect the passthrough path

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:38:38 -03:00
Sameer Kankute
9a13c76e2f
Merge pull request #22553 from dsteeley/fix/streaming-multi-tool-call-premature-finish
fix(streaming): output_item.done for function_call must not emit finish_reason
2026-03-05 15:05:43 +05:30
Julio Quinteros
b3bbcd3955 fix(lint): remove unreachable None check in _resolve_fine_tuning_timeout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:29:57 -03:00
Sameer Kankute
a2c11d431a fix(vertex_ai): drop unsupported output_config parameter from all requests
Vertex AI does not support the output_config parameter in its API.
This parameter is being added by Anthropic/Gemini transformations but needs
to be removed before sending requests to Vertex AI endpoints.

This fix addresses the "Extra inputs are not permitted" error (issue #22312)
when using Claude models with structured outputs on Vertex AI.

Changes:
- Drop output_config in Gemini model transformation
- Drop output_config in Anthropic partner model transformation
- Drop output_config in Anthropic experimental pass-through transformation
- Add comprehensive tests to verify output_config is dropped

Fixes: #22312
Made-with: Cursor
2026-03-05 13:02:17 +05:30
Sameer Kankute
cdf2d67fc8
Merge pull request #22503 from giulio-leone/fix/graceful-tool-args-repair
fix(tools): gracefully repair truncated JSON in tool call arguments
2026-03-05 13:00:07 +05:30
Sameer Kankute
f7d5ff9e2a
Merge pull request #22692 from giulio-leone/fix/vertex-ai-streaming-truncation
fix(streaming): prevent Vertex AI Claude content truncation when finish_reason races content
2026-03-05 12:50:49 +05:30
mubashir1osmani
b3f3918e98 fix(provider): register bedrock_mantle in model_list and models_by_provider
Adds bedrock_mantle_models to the model_list union and models_by_provider
dict so models are discoverable via litellm.model_list and
litellm.models_by_provider["bedrock_mantle"].

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 00:34:05 -05:00
mubashir1osmani
1bf0a3adc4
Update ui/litellm-dashboard/src/components/provider_info_helpers.tsx
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-05 00:20:20 -05:00