Commit graph

19740 commits

Author SHA1 Message Date
Yassin Kortam
356b8d4074
Merge pull request #41444 from zachbernstein-sdx/fix/scim-pagination-count-clamp
fix(scim): align pagination `count` validation with RFC 7644
2026-09-17 14:59:22 -07:00
mateo-berri
4669041ae8 fix(a2a): copy registered agent headers so one caller's bearer never reaches the next
The chat route handed the registry's stored headers dict straight to validate_environment, which wrote the caller's bearer into it, so the next caller of the same agent with no key of their own sent the previous caller's token. The registry lookup now copies the stored headers and validate_environment returns a new dict instead of mutating its input. A regression test drives two completions through one registered agent and asserts the second carries no Authorization and the stored agent is unchanged.
2026-09-17 14:58:50 -07:00
yassin
4885594a1e fix(proxy): use path-style S3 URLs for dotted Transcribe media buckets
Virtual-hosted URLs for bucket names containing dots fail TLS verification, so the
media duration fetch failed and completed jobs were charged the eight hour maximum

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:55:32 +00:00
Mateo Wang
80b0a875dc
Merge pull request #41673 from BerriAI/litellm_deprecate_litellm_proxy_entrypoint
feat(cli): deprecate the litellm-proxy entrypoint in favour of lite
2026-09-17 14:54:47 -07:00
kerry
6858663fd2 test: share the video cost expectation helper and trim the gate workflow triggers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:51:06 +00:00
kerry
bdd9335116 docs(e2e): list the govcloud bedrock test as a coverage matrix row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:51:02 +00:00
mateo-berri
255190ca35 Merge remote-tracking branch 'origin/main' into litellm_fix_image_edits_bracketed_alias 2026-09-17 14:49:54 -07:00
kerry
c988a5002b ci: gate changed tests against a mutated cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:49:16 +00:00
kerry
ccff1fa95f test: derive the remaining cost-map pins from the catalog entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:49:16 +00:00
kerry
4b45fd5f44 docs(e2e): drop govcloud keys from the contributing starter env
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:47:21 +00:00
yassin
fb99e3dde3 Merge remote-tracking branch 'origin/main' into litellm_mcp_client_allowlist 2026-09-17 21:46:21 +00:00
yassin
c0f335d8ca fix(mcp): cap the body the client allowlist inspects at 64 KiB
With mcp_allowed_clients set the gateway used to read the whole POST body to
find clientInfo.name, so an authenticated client could make the proxy buffer an
arbitrarily large payload. Inspection is now capped at MCP_ALLOWLIST_PEEK_MAX_BYTES
and a sessionless POST that exceeds the cap is rejected with 403 before routing,
while posts on an admitted session stream through unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:46:15 +00:00
jesus
83b5f68801 chore: merge main into litellm_org_alias_from_team
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:45:41 +00:00
yassin
16500bdf07 fix(proxy): cap Transcribe pricing media downloads by size and concurrency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:41:33 +00:00
yucheng
1f3b58a528 fix(proxy): dispatch llm_api_check moderation through during_call_hook
ProxyLogging.during_call_hook only ran async_moderation_hook for CustomGuardrail callbacks, so a
CustomLogger such as the prompt injection detector with llm_api_check enabled never called the
configured moderation model. Dispatch any CustomLogger that overrides async_moderation_hook and hand
the proxy router to every registered prompt injection detector at startup so that call can route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:36:36 +00:00
mateo-berri
36b471ff24 Merge commit '1dd4c13815' into litellm_bedrock_openai_no_cachepoint 2026-09-17 14:36:33 -07:00
yucheng
54bf3e9bd9 Merge remote-tracking branch 'origin/main' into litellm_prompt_injection_async_llm_check 2026-09-17 21:36:27 +00:00
yucheng
b4f71df2d2 refactor(proxy): move llm_api_check moderation dispatch to its own PR
Keeps this branch scoped to running the prompt injection heuristics off the event loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:36:13 +00:00
kerry
5aec6d7bb6 test(e2e): assert govcloud file content round-trips the uploaded record
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:31:33 +00:00
kerry
733a8a482e ci(auto-merge): stop requiring Greptile and Bugbot on price sync pull requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:28:25 +00:00
Yassin Kortam
bc9f4fec5b
Merge pull request #41542 from BerriAI/litellm_bedrock_realtime_sdk_0_11
fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime
2026-09-17 14:26:39 -07:00
Yassin Kortam
1b4739c415
Merge pull request #41493 from BerriAI/litellm_bridge_mid_conversation_system_turns
fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions
2026-09-17 14:25:36 -07:00
jesus
cfd8c18616 fix(auth): inherit org budget, tpm and rpm limits for JWT and team-linked keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:24:36 +00:00
yassin
6e7c3f68a1 test(deepgram): pin litellm.max_budget to zero in the model authorization route test
Under xdist the per-test litellm reload is skipped, so a leaked max_budget from another proxy test sent the real
key auth path into the global spend lookup, which the MagicMock prisma client cannot await

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:06:07 +00:00
mateo-berri
ec0e6dd98a feat(cli): deprecate the litellm-proxy entrypoint in favour of lite 2026-09-17 14:05:31 -07:00
mateo-berri
a440d6d452 refactor(cli): rename lite autoroute up/down to start/stop 2026-09-17 14:05:20 -07:00
yassin
d240a5b6bb test(mcp): import json at module level in the MCP allowlist tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:04:56 +00:00
yassin
664688f337 fix(proxy): do not carry a zero team default cap onto a new member budget row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:00:31 +00:00
yassin
6e84ff0cb2 fix(bedrock): keep raw SDK import failure out of the realtime client error
Log the underlying ImportError server side and send the client only the installed
version, the supported range and the install hint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:57:28 +00:00
yassin
14639bbb5a fix(mcp): read the whole initialize body under allowlist enforcement and surface a stored empty allowlist in the UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:55:09 +00:00
kerry
f5940bcee4 Merge remote-tracking branch 'origin/main' into litellm_remove_brittle_price_pinning_tests 2026-09-17 20:52:32 +00:00
yassin
2801614878 fix(proxy): bill Transcribe jobs by media length and refuse media LiteLLM cannot measure
Amazon Transcribe bills every second of the media file, silence included, while the
transcript's last end_time stops at the last word, so pricing from the transcript
undercharged. After a job completes, download Media.MediaFileUri from S3 with the
proxy's credentials and read its length with libsndfile. Formats libsndfile cannot
read (mp4, m4a, webm, amr) and custom language models under LanguageIdSettings are
refused before signing. The S3 signature is only sent to hosts in the AWS partition

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:52:10 +00:00
yassin
3084d2af31 test(utils): allow /v1/listen in the registry supported_endpoints schema
The deepgram/streaming/* rows added for the Deepgram WebSocket passthrough declare /v1/listen as their endpoint, so the registry validation test needs it in the enum, the same way /vertex_ai/live was added for that passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:51:50 +00:00
yassin
1206fa802b refactor(agents): resolve attached access groups from the agent registry instead of the DB on the request path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:51:04 +00:00
yassin
e73b8d49da fix(proxy): seed a new member budget row from the team default cap
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:49:58 +00:00
yassin
11ec157d71 fix(deepgram): authorize the effective model and price /listen sessions at streaming rates
Key auth on the Deepgram WebSocket route now sees the same model the upstream target will carry, so a key restricted to other models can no longer reach nova-3 by leaving model out of the query. user_api_key_auth_websocket keeps its signature and delegates to user_api_key_auth_websocket_for_model, which the Deepgram route calls with deepgram_listen_requested_model

Sessions are priced from new deepgram/streaming/* registry rows (nova-3, nova-3-multilingual for language=multi) plus per-minute add-on rows for redact, keyterm, detect_entities and diarize, all read from Deepgram's pricing page on 2026-09-17. Models without a streaming row fall back to their pre-recorded row as before

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:36:03 +00:00
yuneng-jiang
efce8b0485
Merge pull request #41373 from BerriAI/litellm_fix_integration_conftest_import
fix(tests): resolve the integration support package without run.py's PYTHONPATH
2026-09-17 13:34:05 -07:00
yassin
a56390ed09 Merge remote-tracking branch 'origin/main' into litellm_bedrock_realtime_sdk_0_11
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	uv.lock
2026-09-17 20:32:59 +00:00
ryan-crabbe-berri
5396810bb6 fix(management_v1): authorize bulk member budget writes off the writer and reject unschedulable reset windows
The roster the authorization check reads came from the routed reader, so a
replica lagging behind a team-admin demotion could still grant that caller
member-budget writes. Pin that read to the writer, as the model reconcile does.

A budget_duration the reset job can never schedule from, a non-positive one
that leaves the row permanently due or an unparseable one that blew up mid
batch as a 500, is now a 422 naming the row it came from, with nothing written.
The check is the same one /team/member_update and /budget/new already run,
lifted out of validate_budget_duration so both surfaces share it.
2026-09-17 13:31:32 -07:00
yassin
95a2d5088a fix(proxy): keep custom-auth end-user caps under a key default budget
Custom auth callables that already capped an end user keep their cap; the key
default fills only unset limits. The proxy-wide default still reaches an
uncapped custom-auth token, and the missing-budget log strips line breaks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:31:26 +00:00
yassin
3a86567c9d fix(agents): evict the cached agent access groups on every agent write and cap the model listing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:31:03 +00:00
yassin
97c50acc19 fix(proxy): persist a temp budget pair for members without a private budget row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:30:29 +00:00
yassin
25846d1341 fix(proxy): limit the whole Azure Speech batch API to proxy admin keys
Ordinary keys could read, patch and delete batch transcription jobs that other keys created with the proxy's shared Azure subscription, so every /speechtotext/v3.2 method is now admin only while fast transcription stays open. Also clears SERVER_ROOT_PATH in the real-auth test helper because test_custom_proxy leaves it set at import time and the shared app then 404s pass-through routes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:29:51 +00:00
ryan-crabbe-berri
fbbddb922e
Merge pull request #41488 from BerriAI/litellm_bound_enduser_reset_invalidation
fix(budgets): page end-user cache invalidation after a budget reset
2026-09-17 13:21:23 -07:00
mateo
dcbb77326d chore(proxy): merge origin/main into TypeSafe passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:20:16 +00:00
kerry
435ab2ccc0 Merge remote-tracking branch 'origin/main' into litellm_aws_govcloud_partition_gate 2026-09-17 20:16:28 +00:00
mateo
6a76ca0c72 refactor(vertex_ai): remove constant-False is_using_v1beta1_features stub and its dead call sites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:13 +00:00
kerry
5a5b18550c test(e2e): cover bedrock batch files in govcloud
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:09 +00:00
kerry
fc35b78eb4 ci: drop aws partition hardcode gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:07 +00:00
mateo
07d593d831 refactor(interactions): remove expired use_legacy_interactions_schema shim
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:14:07 +00:00