Commit graph

53363 commits

Author SHA1 Message Date
yassin
c0f335d8ca fix(mcp): cap the body the client allowlist inspects at 64 KiB
With mcp_allowed_clients set the gateway used to read the whole POST body to
find clientInfo.name, so an authenticated client could make the proxy buffer an
arbitrarily large payload. Inspection is now capped at MCP_ALLOWLIST_PEEK_MAX_BYTES
and a sessionless POST that exceeds the cap is rejected with 403 before routing,
while posts on an admitted session stream through unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:46:15 +00:00
jesus
83b5f68801 chore: merge main into litellm_org_alias_from_team
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:45:41 +00:00
yassin
16500bdf07 fix(proxy): cap Transcribe pricing media downloads by size and concurrency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:41:33 +00:00
yucheng
1f3b58a528 fix(proxy): dispatch llm_api_check moderation through during_call_hook
ProxyLogging.during_call_hook only ran async_moderation_hook for CustomGuardrail callbacks, so a
CustomLogger such as the prompt injection detector with llm_api_check enabled never called the
configured moderation model. Dispatch any CustomLogger that overrides async_moderation_hook and hand
the proxy router to every registered prompt injection detector at startup so that call can route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:36:36 +00:00
mateo-berri
36b471ff24 Merge commit '1dd4c13815' into litellm_bedrock_openai_no_cachepoint 2026-09-17 14:36:33 -07:00
yucheng
54bf3e9bd9 Merge remote-tracking branch 'origin/main' into litellm_prompt_injection_async_llm_check 2026-09-17 21:36:27 +00:00
yucheng
b4f71df2d2 refactor(proxy): move llm_api_check moderation dispatch to its own PR
Keeps this branch scoped to running the prompt injection heuristics off the event loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:36:13 +00:00
kerry
5aec6d7bb6 test(e2e): assert govcloud file content round-trips the uploaded record
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:31:33 +00:00
kerry
733a8a482e ci(auto-merge): stop requiring Greptile and Bugbot on price sync pull requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:28:25 +00:00
Yassin Kortam
bc9f4fec5b
Merge pull request #41542 from BerriAI/litellm_bedrock_realtime_sdk_0_11
fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime
2026-09-17 14:26:39 -07:00
Yassin Kortam
1b4739c415
Merge pull request #41493 from BerriAI/litellm_bridge_mid_conversation_system_turns
fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions
2026-09-17 14:25:36 -07:00
jesus
cfd8c18616 fix(auth): inherit org budget, tpm and rpm limits for JWT and team-linked keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:24:36 +00:00
yassin
25729c521f refactor(agents): combine key and team agent grants without a fall-through match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:22:39 +00:00
yassin
6e7c3f68a1 test(deepgram): pin litellm.max_budget to zero in the model authorization route test
Under xdist the per-test litellm reload is skipped, so a leaked max_budget from another proxy test sent the real
key auth path into the global spend lookup, which the MagicMock prisma client cannot await

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:06:07 +00:00
mateo-berri
ec0e6dd98a feat(cli): deprecate the litellm-proxy entrypoint in favour of lite 2026-09-17 14:05:31 -07:00
mateo-berri
a440d6d452 refactor(cli): rename lite autoroute up/down to start/stop 2026-09-17 14:05:20 -07:00
yassin
d240a5b6bb test(mcp): import json at module level in the MCP allowlist tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:04:56 +00:00
yassin
7229fc1952 fix(agents): return on every branch of the agent access ceiling so CodeQL sees no fall-through
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:00:56 +00:00
yassin
664688f337 fix(proxy): do not carry a zero team default cap onto a new member budget row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:00:31 +00:00
yassin
6e84ff0cb2 fix(bedrock): keep raw SDK import failure out of the realtime client error
Log the underlying ImportError server side and send the client only the installed
version, the supported range and the install hint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:57:28 +00:00
yassin
14639bbb5a fix(mcp): read the whole initialize body under allowlist enforcement and surface a stored empty allowlist in the UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:55:09 +00:00
kerry
f5940bcee4 Merge remote-tracking branch 'origin/main' into litellm_remove_brittle_price_pinning_tests 2026-09-17 20:52:32 +00:00
yassin
2801614878 fix(proxy): bill Transcribe jobs by media length and refuse media LiteLLM cannot measure
Amazon Transcribe bills every second of the media file, silence included, while the
transcript's last end_time stops at the last word, so pricing from the transcript
undercharged. After a job completes, download Media.MediaFileUri from S3 with the
proxy's credentials and read its length with libsndfile. Formats libsndfile cannot
read (mp4, m4a, webm, amr) and custom language models under LanguageIdSettings are
refused before signing. The S3 signature is only sent to hosts in the AWS partition

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:52:10 +00:00
yassin
3084d2af31 test(utils): allow /v1/listen in the registry supported_endpoints schema
The deepgram/streaming/* rows added for the Deepgram WebSocket passthrough declare /v1/listen as their endpoint, so the registry validation test needs it in the enum, the same way /vertex_ai/live was added for that passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:51:50 +00:00
yassin
1206fa802b refactor(agents): resolve attached access groups from the agent registry instead of the DB on the request path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:51:04 +00:00
yassin
e73b8d49da fix(proxy): seed a new member budget row from the team default cap
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:49:58 +00:00
yassin
c085619f29 fix(mcp): pass the original receive to the SSE handler when no body was consumed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:44:24 +00:00
yuneng-jiang
1dd4c13815
Merge pull request #41374 from BerriAI/litellm_fix_guardrail_lifecycle_untimed_entries
fix(ui): keep untimed guardrail entries on the request lifecycle
2026-09-17 13:44:18 -07:00
yassin
18c31e6fc6 fix(agents): import assert_never from typing_extensions for Python 3.10
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:40:06 +00:00
yassin
11ec157d71 fix(deepgram): authorize the effective model and price /listen sessions at streaming rates
Key auth on the Deepgram WebSocket route now sees the same model the upstream target will carry, so a key restricted to other models can no longer reach nova-3 by leaving model out of the query. user_api_key_auth_websocket keeps its signature and delegates to user_api_key_auth_websocket_for_model, which the Deepgram route calls with deepgram_listen_requested_model

Sessions are priced from new deepgram/streaming/* registry rows (nova-3, nova-3-multilingual for language=multi) plus per-minute add-on rows for redact, keyterm, detect_entities and diarize, all read from Deepgram's pricing page on 2026-09-17. Models without a streaming row fall back to their pre-recorded row as before

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:36:03 +00:00
mateo
8d760679f6 refactor(ui): drop JSDoc example from DocsMenu
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:35:05 +00:00
yassin
89330cdac6 refactor(agents): exhaust the agent access match and drop routine comments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:34:06 +00:00
yuneng-jiang
efce8b0485
Merge pull request #41373 from BerriAI/litellm_fix_integration_conftest_import
fix(tests): resolve the integration support package without run.py's PYTHONPATH
2026-09-17 13:34:05 -07:00
yassin
a56390ed09 Merge remote-tracking branch 'origin/main' into litellm_bedrock_realtime_sdk_0_11
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	uv.lock
2026-09-17 20:32:59 +00:00
ryan-crabbe-berri
5396810bb6 fix(management_v1): authorize bulk member budget writes off the writer and reject unschedulable reset windows
The roster the authorization check reads came from the routed reader, so a
replica lagging behind a team-admin demotion could still grant that caller
member-budget writes. Pin that read to the writer, as the model reconcile does.

A budget_duration the reset job can never schedule from, a non-positive one
that leaves the row permanently due or an unparseable one that blew up mid
batch as a 500, is now a 422 naming the row it came from, with nothing written.
The check is the same one /team/member_update and /budget/new already run,
lifted out of validate_budget_duration so both surfaces share it.
2026-09-17 13:31:32 -07:00
yassin
95a2d5088a fix(proxy): keep custom-auth end-user caps under a key default budget
Custom auth callables that already capped an end user keep their cap; the key
default fills only unset limits. The proxy-wide default still reaches an
uncapped custom-auth token, and the missing-budget log strips line breaks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:31:26 +00:00
yassin
3a86567c9d fix(agents): evict the cached agent access groups on every agent write and cap the model listing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:31:03 +00:00
yassin
5b04560997 refactor(agents): keep the model listing cap within the type-discipline budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:30:37 +00:00
yassin
97c50acc19 fix(proxy): persist a temp budget pair for members without a private budget row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:30:29 +00:00
yuneng-jiang
7fb3e73276
Merge pull request #41659 from BerriAI/litellm_/release-version-bump-229f45
chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99
2026-09-17 13:30:19 -07:00
yassin
25846d1341 fix(proxy): limit the whole Azure Speech batch API to proxy admin keys
Ordinary keys could read, patch and delete batch transcription jobs that other keys created with the proxy's shared Azure subscription, so every /speechtotext/v3.2 method is now admin only while fast transcription stays open. Also clears SERVER_ROOT_PATH in the real-auth test helper because test_custom_proxy leaves it set at import time and the shared app then 404s pass-through routes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:29:51 +00:00
Yuneng Jiang
6544671a31
fix(ui): order each lifecycle phase on its own clock 2026-09-17 13:29:45 -07:00
yassin
fa1c1f27d9 docs(mcp): list client_allowlist.py in the mcp_server package map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:27:03 +00:00
mateo
ce1a8340c7 refactor(ui): rename HelpLink module to DocsMenu
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:23:11 +00:00
ryan-crabbe-berri
fbbddb922e
Merge pull request #41488 from BerriAI/litellm_bound_enduser_reset_invalidation
fix(budgets): page end-user cache invalidation after a budget reset
2026-09-17 13:21:23 -07:00
mateo
dcbb77326d chore(proxy): merge origin/main into TypeSafe passthrough
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:20:16 +00:00
kerry
435ab2ccc0 Merge remote-tracking branch 'origin/main' into litellm_aws_govcloud_partition_gate 2026-09-17 20:16:28 +00:00
mateo
6a76ca0c72 refactor(vertex_ai): remove constant-False is_using_v1beta1_features stub and its dead call sites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:13 +00:00
kerry
5a5b18550c test(e2e): cover bedrock batch files in govcloud
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:09 +00:00
kerry
fc35b78eb4 ci: drop aws partition hardcode gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:07 +00:00