Commit graph

50986 commits

Author SHA1 Message Date
kerry-berri
0f5bc0ffa9
Merge pull request #41627 from BerriAI/litellm_fireworks_minimax_m3_vision_tests
test(fireworks_ai): stop pinning vision support on minimax-m3
2026-09-17 12:28:23 -07:00
Yuneng Jiang
9207c9a3d8
Merge remote-tracking branch 'origin/main' into litellm_add_spendlogs_api_key_startTime_index
# Conflicts:
#	litellm-proxy-extras/litellm_proxy_extras/schema.prisma
#	litellm/proxy/schema.prisma
#	schema.prisma
2026-09-17 12:20:26 -07:00
yujonglee
dc81cf57f7
Merge pull request #41550 from BerriAI/new-ocr-mapping2
refactor(ocr): mirror Python provider layout and preserve tests
2026-09-17 12:18:02 -07:00
yucheng
0f3556b1bf refactor(proxy): drop explanatory docstrings from the stream attribution helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:17:49 +00:00
Devin AI
14e4b9f906 fix(gemini): gemini-3.5-flash-lite priority cache read is $0.054/M
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:14:04 +00:00
yassin
4212e1ff6f Merge remote-tracking branch 'origin/main' into litellm_transcribe_passthrough 2026-09-17 19:11:49 +00:00
yassin
784fe5bfd8 fix(proxy): price Amazon Transcribe jobs at completion so budgets apply
StartTranscriptionJob was logged with response_cost 0.0, so key, team and proxy
budgets never stopped repeated jobs on the proxy's AWS credentials. The success
handler now polls GetTranscriptionJob to completion, reads the audio duration
from the transcript artifact and charges whole seconds at the cost map rate,
charging the longest media AWS accepts when the duration cannot be read. The
route refuses job classes and surcharge features the cost map does not price

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:11:45 +00:00
Yujong Lee
cd4d78a26a fix(ocr): narrow public error attribute writes and cover callback failure mapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:06:28 +00:00
yassin
2fea3f53b7 perf(anthropic-bridge): reorder mid-conversation system runs in a single pass
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:06:06 +00:00
Devin AI
d877253a7a Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_17 2026-09-17 19:04:40 +00:00
joshua-berri
1e7b03a6ed
Merge pull request #41619 from BerriAI/litellm_fix_mcp_guardrail_context_4889
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 19:04:04 +00:00
joshua-berri
f075417643
Merge pull request #41609 from BerriAI/litellm_fix_mcp_health_permissions_4504
fix(mcp): restrict health discovery to virtual key grants
2026-09-17 19:03:50 +00:00
yassin
db560ca652 fix(passthrough): bill every Deepgram channel, not just wall-clock duration
Deepgram charges for the total processed audio across channels, so a stereo /listen session with multichannel=true costs twice its duration. The handler now multiplies the session duration by a validated channel count taken from Metadata.channels, then the widest Results channel_index, then the channels query parameter, defaulting to one. Booleans, floats, strings, zero and negative values are ignored

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:02:44 +00:00
yucheng
73dea4c567 fix(proxy): attribute only the raising layer in stream and pipeline blocks
The streaming wrapper caught every exception crossing its boundary and named its own
callback, so a block by an inner guardrail or a provider stream failure also named every
outer guardrail. The wrapper now runs the hook over an upstream boundary that remembers
the exception it raised, and skips attribution when the same exception passes through

Pipeline blocks converted from SensitiveDataRouteException or ModifyResponseException into
a generic guardrail_pipeline_error now still record the blocking step's guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:58:21 +00:00
kerry
6d20e68706 test(fireworks_ai): stop pinning vision support on minimax-m3
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:57:50 +00:00
yassin
701c880922 feat(ui): temporary budget increase controls for team members
Adds temp_budget_increase and temp_budget_expiry to the team member edit form with pair validation,
seeds stored values into edit mode, sends both through /team/member_update, and adds cached-key auth
and reservation regression tests for active and expired increases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:56:30 +00:00
kerry-berri
ab03850666
Merge pull request #41623 from BerriAI/litellm_lit_8010_mock_response_provider_custom_pricing
fix(mock_completion): keep the resolved provider so router custom pricing resolves for azure_ai deployments
2026-09-17 11:49:33 -07:00
Yujong Lee
1f0c10147d merge: port OCR request validation and upstream error mapping onto main's dispatch layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:33:20 +00:00
yassin
849859001f fix(deepgram): refuse callback delivery on the /listen passthrough so sessions cannot go unbilled
With callback or callback_method in the query, Deepgram sends every Results and Metadata frame to the caller's URL and only a request id down this socket, so the proxy would meter zero seconds of audio while its own Deepgram credential paid for the transcription. The route now closes such connections with 1008 before contacting Deepgram, naming the offending parameters in the close reason. Adds helper and route tests for both parameters and a nine mutation sweep, all killed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:18:40 +00:00
yujonglee
d5b8400aa9
Merge pull request #41479 from BerriAI/litellm_rust_bridge_declarative_route_catalog
refactor(rust_bridge): declarative route catalog and shared runtime selection
2026-09-17 11:18:36 -07:00
Yujong Lee
15f0d83305 merge: take main's e2e team allow-list settle helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:16:46 +00:00
yassin
8d972eefc7 feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use
Replace the per-deployment asyncio.Semaphore with MaxParallelRequestsLimit, which admits a call synchronously or raises the router's RateLimitError (429) right away. Nothing waits for a slot any more, so the max_parallel_requests_queue_size and default_max_parallel_requests_queue_size settings from the earlier commits are dropped along with their proxy validation, dashboard control and generated schema entries. The rpm/tpm derivation of the cap is unchanged. Every router endpoint family now enters the slot through one _deployment_slot context, and the provider coroutine is only created once the slot is held

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:14:09 +00:00
yassin
5f64dfd8dd fix(proxy): price Azure Speech fast transcription and limit unpriced batch writes to admins
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:14:04 +00:00
Devin AI
f972fddafc test(proxy): include temp budget fields in customer budget table fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:10:00 +00:00
Yujong Lee
766f45e0eb merge: resolve conflicts with main for anthropic layout rename
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:09:27 +00:00
Yujong Lee
b26935416a Merge remote-tracking branch 'github/main' into litellm_rust_bridge_declarative_route_catalog
# Conflicts:
#	tests/e2e/access_control/test_model_access_group_e2e.py
2026-09-17 11:08:08 -07:00
kerry
3cf42f6565 test(mock_completion): cover the provider inference fallback for direct calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:07:58 +00:00
yassin
7c6fd3090f Merge remote-tracking branch 'origin/main' into litellm_bridge_mid_conversation_system_turns 2026-09-17 18:07:06 +00:00
yassin
99667ad633 fix(anthropic-bridge): keep mid-conversation system turns when the target declares supports_mid_conversation_system
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:06:55 +00:00
Yujong Lee
b170d61b8d route stuff through dispatch no direct main 2026-09-17 11:06:46 -07:00
yuneng-jiang
acf75a525c
Merge pull request #41616 from BerriAI/litellm_/buildkite-241-triage-4a48cf
fix(e2e): clear the three standing errors in the scheduled Buildkite suite
2026-09-17 11:05:34 -07:00
mateo
b2ef8daee8 fix(proxy): stop duplicating query params on the TypeSafe passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:03:08 +00:00
berriai-litellm-provider-info-sync[bot]
730195c603
chore(prices): sync Together AI prices: 6 models, 6 deprecated [sync failed: Google Gemini]
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4-Pro-0813: deprecation_date
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-17 18:01:11 +00:00
yujonglee
e038a4feb2
Merge pull request #41531 from BerriAI/litellm_anthropic_stream_types
feat(rust): map Anthropic Messages transformations
2026-09-17 11:00:56 -07:00
Devin AI
b94cd21707 test(proxy): suppress TQ008 on member temp budget patches
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:57:42 +00:00
yassin
3b2b9bf7b6 Merge branch 'litellm_usage_key_free_aggregate_split' into litellm_daily_global_spend_table
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-17 17:55:47 +00:00
yassin
cf07ee1dec Merge remote-tracking branch 'origin/main' into litellm_usage_key_free_aggregate_split
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:55:10 +00:00
Yujong Lee
56ba988b62 fix wrong assertion 2026-09-17 10:54:58 -07:00
Joshua Valluru
f4918e69f4 test(mcp): use the shared guardrail exception in regression 2026-09-17 10:54:28 -07:00
yassin
542cfb218e Merge remote-tracking branch 'origin/main' into litellm_vertex_gcs_file_content_streaming
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-17 17:53:23 +00:00
Devin AI
7c1eb197bf chore(ui): regenerate dashboard API types for team member temp budget fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:52:20 +00:00
Devin AI
32d1dd0cde fix(proxy): apply temp budget increase at member spend admission and reservation checks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:48:33 +00:00
Devin AI
e43f19fc7c docs(proxy): document temp budget fields on organization endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:42:46 +00:00
Yujong Lee
4dcbef0558 refactor(ocr): drop mutable collection builds flagged by LIT002 gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:42:44 +00:00
Joshua Valluru
5b91195406 fix(mcp): retain selected guardrails for virtual REST calls 2026-09-17 10:40:33 -07:00
Yuneng Jiang
fa01e2d5b7
Merge branch 'main' into litellm_/buildkite-241-triage-4a48cf 2026-09-17 10:39:37 -07:00
Yuneng Jiang
dd6ef9e1bc
fix(e2e): delete raw cloud-storage batch files with the master key
DELETE /v1/files/{id} only lets a proxy admin key delete a raw s3:// or
gs:// file id, because such ids skip the managed-file owner check. The
batch lifecycle cleanup deleted the vertex_ai raw ids with the test's
own virtual key and got a 403 at teardown on every build since #194

Raw cloud-storage ids now go through the master key; managed and
provider-native ids keep using the creating key
2026-09-17 10:37:08 -07:00
Yuneng Jiang
1d71564063
fix(e2e): settle the team allow-list through /team/info
The team access-group fixture polled a 403 until its message enumerated
the team's allow-list, because registering a team-scoped deployment
appends that deployment to the list and the fixture has to wait for the
reset to land. #41310 replaced that message with a fixed client-facing
one, so the poll never matched and both tests errored at setup

The allow-list is now read back from /team/info until it holds exactly
the access group
2026-09-17 10:37:07 -07:00
Yuneng Jiang
c3048dcd30
test(http): move the outbound HTTP/2 check into a new integration sdk suite
The check spins up a hypercorn TLS peer and drives the SDK's own httpx
handlers at it, so it needs litellm importable, hypercorn installed and a
loopback socket. It lived under tests/e2e, whose Buildkite runner image
installs neither litellm nor hypercorn by design (the suite drives a
remote proxy over HTTP), so every scheduled e2e build since #230 failed
to import the module and pytest reported it as a collection error. The
unit tree bans sockets, so it does not belong there either

tests/integration is the CircleCI tier built for real TCP against local
protocol peers. This adds an sdk shard to it for cases that exercise the
SDK's clients with no gateway in the path, registers the two HTTP/2
nodes in the contracts manifest, and adds the shard to the CircleCI
matrix. The test now flips the feature through LITELLM_HTTP2 (the user
surface) instead of patching module attributes, and asserts the version
the peer observed on the wire next to the one the client reports
2026-09-17 10:37:07 -07:00
kerry
c2dd7bd98a fix(mock_completion): stamp the resolved provider on mock responses so router custom pricing resolves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:36:11 +00:00