Commit graph

18080 commits

Author SHA1 Message Date
Mateo Wang
ec05cd0128
Merge pull request #41887 from BerriAI/litellm_gemma_4_26b_maas_context_window
fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144
2026-09-18 16:14:56 -07:00
ryan-crabbe-berri
4ab23d7343
Merge pull request #41894 from BerriAI/litellm_remove_agent_shin
ci: remove the dead Agent Shin triage workflows and scripts
2026-09-18 16:10:35 -07:00
yucheng-berri
8e93031c19
Merge pull request #41786 from BerriAI/litellm_passthrough_xpass_trace
Pass-through requests inject the proxy span into upstream headers since #40669, which
replaced an explicit x-pass-traceparent with an unrelated trace and dropped its
x-pass-tracestate. Keep the caller's context when the carrier already names a
different trace, and keep the proxy child span for same-trace or missing headers.

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:07:28 -07:00
Mateo Wang
3490754e65
Merge pull request #41871 from BerriAI/litellm_bedrock_eager_input_streaming
feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools
2026-09-18 16:05:10 -07:00
Mateo Wang
ec9435cbf4
Merge pull request #41881 from BerriAI/litellm_responses_ws_deployment_defaults
fix(responses): merge deployment litellm_params into native websocket response.create frames
2026-09-18 16:04:39 -07:00
yucheng-berri
711a1924d4
Merge pull request #41583 from BerriAI/litellm_applied_guardrails_blocker
* fix(proxy): name the blocking guardrail in x-litellm-applied-guardrails

When a guardrail hook raises, the common ProxyLogging dispatch (sequential and parallel pre_call, pipeline block, during_call and post_call metrics wrapper, streaming iterator wrapper) now records that guardrail in applied_guardrails before re-raising, and pre_call_hook folds request-declared guardrails in on its raising path. Buffered streams rebuild their response headers after the first chunk so a post_call block reached while buffering carries the blocker too

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): attribute only the raising layer in stream and pipeline blocks

The streaming wrapper caught every exception crossing its boundary and named its own
callback, so a block by an inner guardrail or a provider stream failure also named every
outer guardrail. The wrapper now runs the hook over an upstream boundary that remembers
the exception it raised, and skips attribution when the same exception passes through

Pipeline blocks converted from SensitiveDataRouteException or ModifyResponseException into
a generic guardrail_pipeline_error now still record the blocking step's guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop explanatory docstrings from the stream attribution helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:00:27 -07:00
ryan
139445179a ci: remove the dead Agent Shin triage workflows and scripts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 22:59:32 +00:00
mateo-berri
44034c1d5e test: source the gemma context window limits and isolate the cost map cache 2026-09-18 15:43:07 -07:00
mateo-berri
1adbfbfbb1 fix: strip eager_input_streaming for non-Claude providers next to input_examples 2026-09-18 15:28:45 -07:00
Yassin Kortam
b0887b63a5
Merge pull request #41838 from BerriAI/litellm_fix_tpm_window_reset_sibling_counters
fix(proxy): reset sibling tpm/rpm counters when the shared rate limit window rolls over
2026-09-18 15:23:05 -07:00
Yassin Kortam
5f83d97669
Merge pull request #41483 from BerriAI/litellm_v1_models_alias_metadata
fix(proxy): resolve model_group_alias to its target for /v1/models metadata
2026-09-18 15:22:02 -07:00
mateo-berri
417a88daed fix(responses): carry dict-valued reasoning_effort and keep the frame type on websocket defaults
A deployment whose reasoning_effort is an object is copied through as
reasoning the way the HTTP mapper does it instead of being dropped, and
the relay re-asserts the response.create frame type after merging
extra_body so a type key inside it can never replace it. The lazy
OpenAPI snapshot goes back to main: the earlier regeneration came from a
Python 3.14 interpreter dedenting docstrings, which CI on 3.12 rejects
2026-09-18 15:13:07 -07:00
Mateo Wang
7283293d83
Merge pull request #41138 from BerriAI/litellm_bedrock_files_s3_endpoint_url
fix(bedrock): carry s3_endpoint_url and s3_region_name into file content downloads
2026-09-18 15:04:01 -07:00
yucheng-berri
84df4c0d1b
Merge pull request #41783 from BerriAI/litellm_rate_limit_fallback_guardrails
fix(proxy): keep requested model guardrails and key disable_fallbacks on rate-limit fallback
2026-09-18 14:59:16 -07:00
mateo-berri
55249c7128 fix: set vertex gemma-4-26b-a4b-it-maas context window to 262144 2026-09-18 14:58:58 -07:00
Yassin Kortam
87694c26ef
Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table
feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard
2026-09-18 14:53:08 -07:00
Yassin Kortam
47209d37f2
Merge pull request #41882 from BerriAI/litellm_azure_speech_api_base_prefix
fix(proxy): classify Azure Speech short audio behind a prefixed api base
2026-09-18 14:50:28 -07:00
yujonglee
59604b2b19
Merge pull request #41884 from BerriAI/litellm_ocr_test_matrix
test(ocr): declarative provider x auth x input matrix for tests/ocr_tests
2026-09-18 14:50:05 -07:00
Yassin Kortam
52d6aab421
Merge pull request #41554 from BerriAI/litellm_deepgram_listen_websocket_passthrough
feat(passthrough): deepgram streaming /v1/listen WebSocket passthrough with duration-based cost tracking
2026-09-18 14:48:37 -07:00
Mateo Wang
d45e04a9fd
Merge pull request #41062 from BerriAI/litellm_mistral_codex_reasoning_effort_client_metadata
fix(mistral): accept reasoning_effort on all models and drop client_metadata for Codex compatibility
2026-09-18 14:43:58 -07:00
Yujong Lee
f72b7155ac fix(ocr): map Rust upstream 401/403 to the public auth exceptions
The httpx.Response built for a Rust upstream failure had no request attached,
so constructing openai.AuthenticationError raised RuntimeError inside the
exception mapper and every bad-key OCR call surfaced as APIConnectionError 500
instead of AuthenticationError 401 (the Python path already returned 401)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:40:13 +00:00
Devin AI
16d63eaf88 Merge remote-tracking branch 'origin/main' into litellm_fix_tpm_window_reset_sibling_counters 2026-09-18 21:38:26 +00:00
Mateo Wang
a6bd779bd1
Merge pull request #39424 from emerzon/litellm_azure_ai_flux_2_flex
feat(azure_ai): support FLUX.2 flex images
2026-09-18 14:30:25 -07:00
yassin
0b5b69ea3a fix(deepgram): forward only the first model and language values to /listen
Authorization and pricing read the first model and language query value, but the raw query was forwarded, so Deepgram (which honours the last repeated value) could be sent a model the key was never allowed. Later duplicates of those two keys are now dropped before the upstream URL is built

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:29:29 +00:00
Yujong Lee
9767878425 test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:27:03 +00:00
ryan-crabbe-berri
a43a4924a6
Merge pull request #40878 from BerriAI/litellm_null_cost_unpriced_deployments
fix(router): report null cost for unpriced deployments instead of 0
2026-09-18 14:20:52 -07:00
mateo-berri
4a951847bb fix(responses): merge deployment litellm_params into native websocket response.create frames 2026-09-18 14:15:39 -07:00
yassin
3449ae9d0d fix(proxy): advance the daily global spend marker in one conditional upsert so overlapping runs cannot rewind it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:15:32 +00:00
yassin
f1b9642c41 fix(proxy): classify Azure Speech short audio behind a prefixed api base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:15:01 +00:00
Mateo Wang
59c24abcbe
Merge pull request #33101 from BerriAI/litellm_fix_responses_ws_litellm_params_leak
fix(responses): stop managed Responses WebSocket from leaking litellm_params into provider request body
2026-09-18 14:11:19 -07:00
yassin
0897b663c5 Merge remote-tracking branch 'origin/litellm_deepgram_listen_websocket_passthrough' into litellm_deepgram_listen_websocket_passthrough 2026-09-18 21:08:45 +00:00
yassin
93d61abfa5 fix(deepgram): refuse /listen sessions that have no streaming price
A caller could pick a model with only a pre-recorded registry row, or no row at all, and the session would be billed at the pre-recorded rate or logged at zero cost, so budgets did not apply. The route now closes the WebSocket with 1008 before dialing Deepgram unless deepgram/streaming/<model> (or the -multilingual row for language=multi) is an exact registry hit, and the logging handler applies the same check so a registry change under a live session records the duration with no cost instead of a substitute rate

Regression tests cover the route refusal, an operator-supplied streaming row for another model being accepted, and the handler never substituting the pre-recorded rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:08:14 +00:00
Yassin Kortam
6759f28e73
Merge pull request #41557 from BerriAI/litellm_azure_speech_passthrough
feat(proxy): add Azure AI Speech pass-through route
2026-09-18 14:05:55 -07:00
ryan-crabbe-berri
006080ea6d
Merge pull request #41349 from BerriAI/litellm_team_member_spend_without_budget
fix(proxy): track team member spend when the member has no budget
2026-09-18 14:04:39 -07:00
Mateo Wang
c553bc92bd
Merge pull request #41875 from BerriAI/litellm_passthrough_stream_timeout
fix(router): honor stream_timeout on the SDK-native passthrough route (/v1/messages, /converse)
2026-09-18 13:59:02 -07:00
Mateo Wang
20e1e6f2a9
Merge pull request #41384 from BerriAI/litellm_fix_azure_vector_store_search_url
fix(azure): keep api-version query after vector store search path
2026-09-18 13:56:05 -07:00
yassin
abf530fbeb fix(proxy): never rewind the daily global spend marker from an overlapping reconcile run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:55:11 +00:00
ryan
61d4c5b9b5 fix(proxy): skip members already on the team before resolving a per-member budget
A mixed /team/member_add list that names an existing member used to run
add_new_member for them, which created or cloned a budget that the empty
upsert update branch never linked to their membership row. Filter the
requested members against the freshly locked roster first so budgets and
membership rows are only written for members who are actually new

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:17 -07:00
ryan
b6f4ad190e refactor(proxy): freeze the member spend arrays and budget link to stay within the type discipline budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:17 -07:00
ryan
2d61fa66b1 fix(proxy): lock teams in sorted team id order and keep existing member budgets on re-add
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:17 -07:00
ryan
499c334fce fix(proxy): lock each team in sorted order before the member spend upsert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:17 -07:00
ryan
60077e90aa fix(proxy): take the team advisory lock before the member spend flush
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:17 -07:00
ryan
efef6ab684 fix(proxy): write team member spend as one roster checked upsert statement
Replaces the per team advisory lock and Pydantic roster parse in the spend flush
with a single INSERT ... ON CONFLICT statement that checks the stored roster in
SQL, so malformed roster JSON cannot fail the whole flush and large batches no
longer issue two queries per team inside the fixed transaction deadline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:16 -07:00
ryan
6dce85c728 fix(proxy): skip recreating membership rows for members removed before a spend flush
Take the team advisory lock in the spend flush transaction and read the roster
through it, so a delayed flush after /team/member_delete cannot recreate the
deleted LiteLLM_TeamMembership row. TEAM_ADVISORY_LOCK_SQL moves to
team_repository so the spend writer can import it without a circular import

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:16 -07:00
ryan
d9b48ac941 test(proxy): mock team membership upsert in team admin member add test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:16 -07:00
ryan
de4b520153 fix(proxy): track team member spend when the member has no budget
add_new_member only wrote a LiteLLM_TeamMembership row when a budget id
resolved, and the spend writer used update_many so a missing row failed
silently. Members without a budget therefore never accrued per-member spend.

The membership row is now always upserted (budget_id NULL when no budget
applies) and the spend write is an upsert so members added before this fix
start accruing on their next request.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:54:16 -07:00
yassin
b7d808e416 Merge remote-tracking branch 'origin/main' into litellm_deepgram_listen_websocket_passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:52:42 +00:00
kerry-berri
6fa34a299b
Merge pull request #41754 from BerriAI/litellm_qwen3_8_omni_flash
feat(models): add qwen3.8 flash rows, fix Cohere embed v3 context, Bedrock Mantle and OpenRouter pricing
2026-09-18 13:49:35 -07:00
mateo-berri
e5fd2bec83 fix: forward eager_input_streaming through the Responses API tool bridge 2026-09-18 13:34:29 -07:00
mateo-berri
e75ad61fce fix(router): rank litellm_settings.request_timeout on the passthrough route like the completion route 2026-09-18 13:32:43 -07:00