Commit graph

18103 commits

Author SHA1 Message Date
mateo-berri
fcc7efa4db fix(responses): forward the routed input and report routing rejections on the websocket 2026-09-18 18:10:26 -07:00
mateo-berri
f213855558 refactor: drop the docstrings from the websocket relay and its tests 2026-09-18 17:44:32 -07:00
mateo-berri
5030214671 Merge origin/main into litellm_fix_responses_ws_encrypted_content_affinity
Resolves the conflicts with the WebSocket request defaults from main (PR #41881):
the relay keeps both custom_llm_provider and request_defaults, and a masked
response.create frame is re-serialized when the defaults changed it.

Keeps the first-frame routing hints (input, previous_response_id) out of the
deployment request defaults so they never get injected into later frames on the
same connection, with a regression test.
2026-09-18 17:23:59 -07:00
mateo-berri
febe9aec65 fix(responses): book a rejected WebSocket connection as a failed request 2026-09-18 17:10:01 -07:00
yucheng-berri
500e880a40
Merge pull request #41895 from BerriAI/litellm_openai_moderations_model_default
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 17:06:41 -07:00
ryan-crabbe-berri
4bb1ae115b
Merge pull request #41347 from BerriAI/litellm_team_member_budget_link_default
fix(team): apply team_member_budget updates to members still on the team default
2026-09-18 17:06:25 -07:00
Mateo Wang
ff7dc86947
Merge pull request #41892 from BerriAI/litellm_gemini_contentless_candidate_finish_reason
Some checks failed
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Issue label sync / sync-issue-labels-tests (push) Has been cancelled
Issue label sync / sync-issue-labels (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
VS Code Extension / vscode-extension (push) Has been cancelled
fix(gemini): preserve candidates with finishReason and no content (#40477)
2026-09-18 16:34:46 -07:00
yujonglee
018f640b30
Merge pull request #41885 from BerriAI/litellm_rust_callback_contract
refactor(rust): formalize legacy callback contract
2026-09-18 16:16:02 -07:00
Mateo Wang
ec05cd0128
Merge pull request #41887 from BerriAI/litellm_gemma_4_26b_maas_context_window
fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144
2026-09-18 16:14:56 -07:00
ryan-crabbe-berri
4ab23d7343
Merge pull request #41894 from BerriAI/litellm_remove_agent_shin
ci: remove the dead Agent Shin triage workflows and scripts
2026-09-18 16:10:35 -07:00
yucheng-berri
8e93031c19
Merge pull request #41786 from BerriAI/litellm_passthrough_xpass_trace
Pass-through requests inject the proxy span into upstream headers since #40669, which
replaced an explicit x-pass-traceparent with an unrelated trace and dropped its
x-pass-tracestate. Keep the caller's context when the carrier already names a
different trace, and keep the proxy child span for same-trace or missing headers.

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:07:28 -07:00
yucheng
1704aeebb4 fix(enterprise): resolve openai_moderations model at call time and default to omni-moderation-latest
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 23:06:47 +00:00
Mateo Wang
3490754e65
Merge pull request #41871 from BerriAI/litellm_bedrock_eager_input_streaming
feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools
2026-09-18 16:05:10 -07:00
ryan
b4c5f6fa44 Merge remote-tracking branch 'origin/main' into litellm_team_member_budget_link_default
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/management_helpers/utils.py
#	tests/test_litellm/proxy/management_helpers/test_management_helpers_utils.py
2026-09-18 23:05:03 +00:00
Mateo Wang
ec9435cbf4
Merge pull request #41881 from BerriAI/litellm_responses_ws_deployment_defaults
fix(responses): merge deployment litellm_params into native websocket response.create frames
2026-09-18 16:04:39 -07:00
mateo-berri
a4624b6c6c fix(gemini): derive the finish reason key set from the Candidates type
Candidates.finishReason listed eleven values while the mapping key set
carried twenty-one, so typed fixtures could not spell the reasons this
PR handles. GeminiFinishReason is now the one list, the key set derives
from it, and a test checks every documented reason has an explicit
mapping instead of falling through to "stop"
2026-09-18 16:03:45 -07:00
yucheng-berri
711a1924d4
Merge pull request #41583 from BerriAI/litellm_applied_guardrails_blocker
* fix(proxy): name the blocking guardrail in x-litellm-applied-guardrails

When a guardrail hook raises, the common ProxyLogging dispatch (sequential and parallel pre_call, pipeline block, during_call and post_call metrics wrapper, streaming iterator wrapper) now records that guardrail in applied_guardrails before re-raising, and pre_call_hook folds request-declared guardrails in on its raising path. Buffered streams rebuild their response headers after the first chunk so a post_call block reached while buffering carries the blocker too

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): attribute only the raising layer in stream and pipeline blocks

The streaming wrapper caught every exception crossing its boundary and named its own
callback, so a block by an inner guardrail or a provider stream failure also named every
outer guardrail. The wrapper now runs the hook over an upstream boundary that remembers
the exception it raised, and skips attribution when the same exception passes through

Pipeline blocks converted from SensitiveDataRouteException or ModifyResponseException into
a generic guardrail_pipeline_error now still record the blocking step's guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop explanatory docstrings from the stream attribution helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:00:27 -07:00
ryan
139445179a ci: remove the dead Agent Shin triage workflows and scripts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 22:59:32 +00:00
Yujong Lee
ba6b22cf56 test(rust): isolate callback registries per hypothesis example
Replace the module-level LATEST_EDITS list with per-example callback
registry isolation, and import litellm names with from-imports in the
legacy callback shim so the module uses one import style.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 15:57:50 -07:00
Yujong Lee
1a52bae779 add streaming message 2026-09-18 15:45:08 -07:00
mateo-berri
1c15d9f291 fix(responses): restore encrypted_content and apply affinity on the native WebSocket relay 2026-09-18 15:44:06 -07:00
Yujong Lee
ab27ee7efc test(rust): fix request count and header case in new OCR callback tests 2026-09-18 15:43:08 -07:00
Yujong Lee
a737e3430a test(rust): property and parametrized tests for legacy callback contracts
The payload boundary of callbacks-legacy gets a model-based proptest: for any
JSON body, caller keywords and callback edit, keywords the route sends unchanged
reach pre_call as the caller's own objects, and the wire is the body pre_call
received as the callback left it. A parametrized test pins that a keyword the
bridge never reads keeps its identity through setup, the deployment hook,
check_limits and prepare.

Behaviour owned by the real Logging object is pinned end to end in the OCR
tests: a hypothesis version of the body property over HTTP, sync hooks seeing no
running event loop, retained payloads staying intact after the call, success
callbacks sharing one standard logging payload, and state stashed before a
blocking deployment hook raises reaching both failure callback families
2026-09-18 15:43:08 -07:00
Yujong Lee
b4bfd92a2a refactor(rust): route-neutral callback contract
Every legacy callback call from callbacks-legacy now goes through one typed
Python shim, litellm.rust_bridge.legacy_callbacks, the only Python module
the crate reaches. Before, the crate called Logging methods, litellm.utils
hooks, the logging worker, the executor and several litellm globals
directly, and its tests retyped those signatures by hand, so an outdated
fake could accept a call the real code rejects. python_contract.json lists
each shim function's parameters: a Python test pins it to the real
signatures and a Rust test pins it to the Rust enum.

The lifecycle contract changes to match the Python @client wrapper:
- the driver emits CallEvent::Started before begin, so every host sees one
  start time
- RequestContext carries the route-resolved api_key, so legacy pre_call and
  post_call receive it, and post_call's additional_args match the Python OCR
  path
- Passthrough and its re-aliasing are gone
- async deployment hooks always run, and the "no callbacks" shortcut that
  skipped the logging payload is removed, as in the Python path

The OCR api_key is a SecretValue from the wire request onward, so Debug
output upstream of the callback contract cannot leak it.

host-python's RouteHost now classifies native failures once through
classify, and host ops return HostOpError. The OCR route host keeps main's
public errors by sending both through the existing Python map_failure.
2026-09-18 15:43:08 -07:00
mateo-berri
44034c1d5e test: source the gemma context window limits and isolate the cost map cache 2026-09-18 15:43:07 -07:00
mateo-berri
1adbfbfbb1 fix: strip eager_input_streaming for non-Claude providers next to input_examples 2026-09-18 15:28:45 -07:00
Yassin Kortam
b0887b63a5
Merge pull request #41838 from BerriAI/litellm_fix_tpm_window_reset_sibling_counters
fix(proxy): reset sibling tpm/rpm counters when the shared rate limit window rolls over
2026-09-18 15:23:05 -07:00
Yassin Kortam
5f83d97669
Merge pull request #41483 from BerriAI/litellm_v1_models_alias_metadata
fix(proxy): resolve model_group_alias to its target for /v1/models metadata
2026-09-18 15:22:02 -07:00
mateo-berri
417a88daed fix(responses): carry dict-valued reasoning_effort and keep the frame type on websocket defaults
A deployment whose reasoning_effort is an object is copied through as
reasoning the way the HTTP mapper does it instead of being dropped, and
the relay re-asserts the response.create frame type after merging
extra_body so a type key inside it can never replace it. The lazy
OpenAPI snapshot goes back to main: the earlier regeneration came from a
Python 3.14 interpreter dedenting docstrings, which CI on 3.12 rejects
2026-09-18 15:13:07 -07:00
Mateo Wang
7283293d83
Merge pull request #41138 from BerriAI/litellm_bedrock_files_s3_endpoint_url
fix(bedrock): carry s3_endpoint_url and s3_region_name into file content downloads
2026-09-18 15:04:01 -07:00
yucheng-berri
84df4c0d1b
Merge pull request #41783 from BerriAI/litellm_rate_limit_fallback_guardrails
fix(proxy): keep requested model guardrails and key disable_fallbacks on rate-limit fallback
2026-09-18 14:59:16 -07:00
mateo-berri
55249c7128 fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144 2026-09-18 14:58:58 -07:00
mateo-berri
cf05466a27 fix(gemini): map every documented finishReason and reset per-candidate state
A content-less candidate is now kept as a choice whenever it carries a
finishReason, with the raw value on the choice's provider_specific_fields.
NO_IMAGE, IMAGE_RECITATION, IMAGE_OTHER and ESCALATION map to content_filter;
UNEXPECTED_TOOL_CALL and MISSING_THOUGHT_SIGNATURE map to stop. The
/v1/responses bridge reports content_filter and refusal as incomplete with
incomplete_details, and tool calls and reasoning no longer leak from one
candidate into the next.
2026-09-18 14:54:18 -07:00
Yassin Kortam
87694c26ef
Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table
feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard
2026-09-18 14:53:08 -07:00
Yassin Kortam
47209d37f2
Merge pull request #41882 from BerriAI/litellm_azure_speech_api_base_prefix
fix(proxy): classify Azure Speech short audio behind a prefixed api base
2026-09-18 14:50:28 -07:00
yujonglee
59604b2b19
Merge pull request #41884 from BerriAI/litellm_ocr_test_matrix
test(ocr): declarative provider x auth x input matrix for tests/ocr_tests
2026-09-18 14:50:05 -07:00
Yassin Kortam
52d6aab421
Merge pull request #41554 from BerriAI/litellm_deepgram_listen_websocket_passthrough
feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
2026-09-18 14:48:37 -07:00
Mateo Wang
d45e04a9fd
Merge pull request #41062 from BerriAI/litellm_mistral_codex_reasoning_effort_client_metadata
fix(mistral): accept reasoning_effort on all models and drop client_metadata for Codex compatibility
2026-09-18 14:43:58 -07:00
Yujong Lee
f72b7155ac fix(ocr): map Rust upstream 401/403 to the public auth exceptions
The httpx.Response built for a Rust upstream failure had no request attached,
so constructing openai.AuthenticationError raised RuntimeError inside the
exception mapper and every bad-key OCR call surfaced as APIConnectionError 500
instead of AuthenticationError 401 (the Python path already returned 401)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:40:13 +00:00
Devin AI
16d63eaf88 Merge remote-tracking branch 'origin/main' into litellm_fix_tpm_window_reset_sibling_counters 2026-09-18 21:38:26 +00:00
Mateo Wang
a6bd779bd1
Merge pull request #39424 from emerzon/litellm_azure_ai_flux_2_flex
feat(azure_ai): support FLUX.2 flex images
2026-09-18 14:30:25 -07:00
yassin
0b5b69ea3a fix(deepgram): forward only the first model and language values to /listen
Authorization and pricing read the first model and language query value, but the raw query was forwarded, so Deepgram (which honours the last repeated value) could be sent a model the key was never allowed. Later duplicates of those two keys are now dropped before the upstream URL is built

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:29:29 +00:00
Yujong Lee
9767878425 test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:27:03 +00:00
ryan-crabbe-berri
a43a4924a6
Merge pull request #40878 from BerriAI/litellm_null_cost_unpriced_deployments
fix(router): report null cost for unpriced deployments instead of 0
2026-09-18 14:20:52 -07:00
mateo-berri
4a951847bb fix(responses): merge deployment litellm_params into native websocket response.create frames 2026-09-18 14:15:39 -07:00
yassin
3449ae9d0d fix(proxy): advance the daily global spend marker in one conditional upsert so overlapping runs cannot rewind it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:15:32 +00:00
yassin
f1b9642c41 fix(proxy): classify Azure Speech short audio behind a prefixed api base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:15:01 +00:00
Mateo Wang
59c24abcbe
Merge pull request #33101 from BerriAI/litellm_fix_responses_ws_litellm_params_leak
fix(responses): stop managed Responses WebSocket from leaking litellm_params into provider request body
2026-09-18 14:11:19 -07:00
yassin
0897b663c5 Merge remote-tracking branch 'origin/litellm_deepgram_listen_websocket_passthrough' into litellm_deepgram_listen_websocket_passthrough 2026-09-18 21:08:45 +00:00
yassin
93d61abfa5 fix(deepgram): refuse /listen sessions that have no streaming price
A caller could pick a model with only a pre-recorded registry row, or no row at all, and the session would be billed at the pre-recorded rate or logged at zero cost, so budgets did not apply. The route now closes the WebSocket with 1008 before dialing Deepgram unless deepgram/streaming/<model> (or the -multilingual row for language=multi) is an exact registry hit, and the logging handler applies the same check so a registry change under a live session records the duration with no cost instead of a substitute rate

Regression tests cover the route refusal, an operator-supplied streaming row for another model being accepted, and the handler never substituting the pre-recorded rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:08:14 +00:00