Commit graph

52881 commits

Author SHA1 Message Date
yassin
d5acbbde6f chore(ui): regenerate schema.d.ts for end_user_budget_id docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:57:44 +00:00
yassin
e932451312 refactor(proxy): move effective member budget onto the budget model and reject negative temp increases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:56:35 +00:00
yassin
21cafc8780 refactor(mcp): validate initialize body and general_settings with pydantic in the client allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:53:33 +00:00
kerry
86c5627d66 Merge remote-tracking branch 'origin/main' into litellm-providers/price-sync 2026-09-17 19:51:01 +00:00
kerry-berri
decbb96382
Merge pull request #41635 from BerriAI/litellm_together_successor_test_drop_deprecation_pin
test(together_ai): stop pinning successor deprecation status
2026-09-17 12:50:50 -07:00
yassin
d9ddc4b901 test(proxy): assert temp budget increase stops at the exact expiry instant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:50:50 +00:00
yassin
b9d0008d97 fix(proxy): keep temp budget fields out of organization metadata on create
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:50:50 +00:00
kerry-berri
c5b0d6218d
Merge pull request #41633 from BerriAI/litellm_non_string_model_spend_tracking
fix(proxy): reject non-string model with 400 and log its spend as unknown-model
2026-09-17 12:49:23 -07:00
yassin
43089640a2 docs(proxy): document end_user_budget_id on key generate and update endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:48:37 +00:00
mubashir1osmani
4e4008cea9 lint(cost): justify blind except in _lookup_model_info_or_none
get_model_info raises a bare Exception for unmapped models, so BLE001 cannot
be narrowed; mark it noqa with the reason to stay within the strict budget.
2026-09-17 15:43:46 -04:00
yassin
322262db01 style(ui): format add_agent_form test with prettier
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:42:56 +00:00
kerry
427d08470f test(together_ai): stop pinning successor deprecation status
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:40:00 +00:00
mubashir1osmani
d2af0577d5 test(files): assert key_model_access_denied error type instead of message text
main (15f2e25e8a) replaced the configurable model-access-denied message
with a fixed client message, so match on the stable error type.
2026-09-17 15:39:07 -04:00
yassin
e91cd877fb Merge remote-tracking branch 'origin/main' into litellm_team_member_temp_budget_increase 2026-09-17 19:37:25 +00:00
kerry
27dd1a02aa fix(proxy): reject non-string model with 400 and log its spend as unknown-model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:37:14 +00:00
ryan-crabbe-berri
4f6dfb0480 fix(management_v1): report a zero team default as no cap
Enforcement treats max_budget 0 on the team default as "no cap" and only
honors 0 as an explicit disable on a member's own row, so reporting an
inheriting member as capped at 0 said the opposite of what happens on their
next request.
2026-09-17 12:35:46 -07:00
yuneng-jiang
dbc6c1cfaa
Merge pull request #37983 from etiennechabert/litellm_add_spendlogs_api_key_startTime_index
perf(spend_tracking): index LiteLLM_SpendLogs by (api_key, startTime)
2026-09-17 12:35:28 -07:00
yassin
99eeb813c4 Merge remote-tracking branch 'origin/main' into litellm_mcp_client_allowlist 2026-09-17 19:34:42 +00:00
yassin
5e7888f44d fix(ui): save MCP allowed clients independently of private IP ranges
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:34:42 +00:00
mubashir1osmani
9066a32363 Merge remote-tracking branch 'berri/main' into litellm_mistral_ocr_batches
# Conflicts:
#	litellm/batches/batch_utils.py
#	tests/test_litellm/llms/mistral/ocr/test_mistral_ocr_cost.py
2026-09-17 15:34:20 -04:00
yassin
177e6a0a97 test(anthropic-bridge): bound role reads instead of wall-clock time in the long system run test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:34:16 +00:00
yassin
433c804a6c refactor(agents): mark Callable loader aliases for the type discipline gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:32:39 +00:00
ryan-crabbe-berri
648373a260 feat(management_v1): bulk update team member budgets
Adds POST /management/v1/teams/{team_id}/members/bulk_update, a merge patch
over per-member limits (max_budget_in_team, tpm_limit, rpm_limit,
budget_duration, allowed_models) for up to 500 members in one transaction.

Editing a team's default member budget has never reached members who already
have a budget row, because /team/member_add clones the default per member.
This gives admins one call to roll a new cap out across the roster, and each
result carries max_budget_source so a caller can see whether a member is on
their own cap or on the team default.

Reads run on the writer inside the batch transaction, and any budget row more
than one membership points at is cloned before it is written, so raising one
member's cap never moves another's.
2026-09-17 12:30:46 -07:00
kerry
5e5e5086f0 Merge remote-tracking branch 'origin/main' into litellm-providers/price-sync 2026-09-17 19:29:18 +00:00
yassin
cc11653152 feat(proxy): per-key default budget for dynamically created customers
A service-account key can now carry end_user_budget_id in its metadata. When a request through that key names a customer that does not exist yet, the key's budget is applied to the new customer from the first request and wins over the proxy-wide max_end_user_budget_id. A customer with an explicitly assigned budget keeps it. The Admin UI exposes the setting on service-account key creation and key edit, and only proxy admins may set or clear it.

Resolves LIT-7996

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:28:53 +00:00
kerry-berri
0f5bc0ffa9
Merge pull request #41627 from BerriAI/litellm_fireworks_minimax_m3_vision_tests
test(fireworks_ai): stop pinning vision support on minimax-m3
2026-09-17 12:28:23 -07:00
yassin
cf9c8fe21b Merge remote-tracking branch 'origin/main' into litellm_agent_access_groups 2026-09-17 19:24:22 +00:00
yassin
b84f8b6a77 feat(agents): attach access groups to agents and enforce them for models, MCP servers and agent calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:24:14 +00:00
Yuneng Jiang
9207c9a3d8
Merge remote-tracking branch 'origin/main' into litellm_add_spendlogs_api_key_startTime_index
# Conflicts:
#	litellm-proxy-extras/litellm_proxy_extras/schema.prisma
#	litellm/proxy/schema.prisma
#	schema.prisma
2026-09-17 12:20:26 -07:00
yujonglee
dc81cf57f7
Merge pull request #41550 from BerriAI/new-ocr-mapping2
refactor(ocr): mirror Python provider layout and preserve tests
2026-09-17 12:18:02 -07:00
yucheng
0f3556b1bf refactor(proxy): drop explanatory docstrings from the stream attribution helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:17:49 +00:00
Devin AI
14e4b9f906 fix(gemini): gemini-3.5-flash-lite priority cache read is $0.054/M
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:14:04 +00:00
yassin
4212e1ff6f Merge remote-tracking branch 'origin/main' into litellm_transcribe_passthrough 2026-09-17 19:11:49 +00:00
yassin
784fe5bfd8 fix(proxy): price Amazon Transcribe jobs at completion so budgets apply
StartTranscriptionJob was logged with response_cost 0.0, so key, team and proxy
budgets never stopped repeated jobs on the proxy's AWS credentials. The success
handler now polls GetTranscriptionJob to completion, reads the audio duration
from the transcript artifact and charges whole seconds at the cost map rate,
charging the longest media AWS accepts when the duration cannot be read. The
route refuses job classes and surcharge features the cost map does not price

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:11:45 +00:00
Yujong Lee
cd4d78a26a fix(ocr): narrow public error attribute writes and cover callback failure mapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:06:28 +00:00
yassin
2fea3f53b7 perf(anthropic-bridge): reorder mid-conversation system runs in a single pass
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:06:06 +00:00
Devin AI
d877253a7a Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_17 2026-09-17 19:04:40 +00:00
joshua-berri
1e7b03a6ed
Merge pull request #41619 from BerriAI/litellm_fix_mcp_guardrail_context_4889
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 19:04:04 +00:00
joshua-berri
f075417643
Merge pull request #41609 from BerriAI/litellm_fix_mcp_health_permissions_4504
fix(mcp): restrict health discovery to virtual key grants
2026-09-17 19:03:50 +00:00
yassin
db560ca652 fix(passthrough): bill every Deepgram channel, not just wall-clock duration
Deepgram charges for the total processed audio across channels, so a stereo /listen session with multichannel=true costs twice its duration. The handler now multiplies the session duration by a validated channel count taken from Metadata.channels, then the widest Results channel_index, then the channels query parameter, defaulting to one. Booleans, floats, strings, zero and negative values are ignored

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:02:44 +00:00
yucheng
73dea4c567 fix(proxy): attribute only the raising layer in stream and pipeline blocks
The streaming wrapper caught every exception crossing its boundary and named its own
callback, so a block by an inner guardrail or a provider stream failure also named every
outer guardrail. The wrapper now runs the hook over an upstream boundary that remembers
the exception it raised, and skips attribution when the same exception passes through

Pipeline blocks converted from SensitiveDataRouteException or ModifyResponseException into
a generic guardrail_pipeline_error now still record the blocking step's guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:58:21 +00:00
kerry
6d20e68706 test(fireworks_ai): stop pinning vision support on minimax-m3
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:57:50 +00:00
yassin
701c880922 feat(ui): temporary budget increase controls for team members
Adds temp_budget_increase and temp_budget_expiry to the team member edit form with pair validation,
seeds stored values into edit mode, sends both through /team/member_update, and adds cached-key auth
and reservation regression tests for active and expired increases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:56:30 +00:00
yassin
a705e0396e refactor(mcp): type the allowlist 403 body and replay consumed messages immutably
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:55:47 +00:00
kerry-berri
ab03850666
Merge pull request #41623 from BerriAI/litellm_lit_8010_mock_response_provider_custom_pricing
fix(mock_completion): keep the resolved provider so router custom pricing resolves for azure_ai deployments
2026-09-17 11:49:33 -07:00
yassin
47be6c8aeb feat(mcp): allowlist client applications for MCP gateway access
Adds the mcp_allowed_clients general setting, enforced against the
clientInfo.name each MCP client sends in its initialize request. A client
not on the list, or one that does not identify itself, is rejected with
403 before any stateful session is created. The setting is configurable
from config.yaml and from the Admin UI MCP network settings page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:34:36 +00:00
Yujong Lee
1f0c10147d merge: port OCR request validation and upstream error mapping onto main's dispatch layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:33:20 +00:00
yassin
849859001f fix(deepgram): refuse callback delivery on the /listen passthrough so sessions cannot go unbilled
With callback or callback_method in the query, Deepgram sends every Results and Metadata frame to the caller's URL and only a request id down this socket, so the proxy would meter zero seconds of audio while its own Deepgram credential paid for the transcription. The route now closes such connections with 1008 before contacting Deepgram, naming the offending parameters in the close reason. Adds helper and route tests for both parameters and a nine mutation sweep, all killed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:18:40 +00:00
yujonglee
d5b8400aa9
Merge pull request #41479 from BerriAI/litellm_rust_bridge_declarative_route_catalog
refactor(rust_bridge): declarative route catalog and shared runtime selection
2026-09-17 11:18:36 -07:00
Yujong Lee
15f0d83305 merge: take main's e2e team allow-list settle helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:16:46 +00:00