Commit graph

49282 commits

Author SHA1 Message Date
mateo
a40804b7dc fix(proxy): block POST /user/new during the master key lockout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:06:57 +00:00
mateo
cbe7857af2 chore(ui): regen schema types after main merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:23:34 +00:00
mateo
2195c0ea1c Merge remote-tracking branch 'origin/main' into litellm_refuse_example_master_key 2026-09-15 17:23:13 +00:00
mateo
7f5278eab3 fix(proxy): cover credentials by_model route param in lockout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:04:08 +00:00
tin-berri
3ad9a7f336
Merge pull request #41174 from BerriAI/litellm_tier_model_affinity
fix(router): preserve session model choice within each complexity tier
2026-09-15 09:54:53 -07:00
mateo
5c2fc1c99d fix(proxy): extract request method helper to satisfy typing gates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 16:14:11 +00:00
mateo
c504ca954e fix(proxy): annotate db user row with the prisma model type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 15:56:00 +00:00
mateo
cedee4800a fix(proxy): read request method defensively for mock and scopeless requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 15:36:05 +00:00
mateo
ab95b8066e fix(proxy): keep plain-str user_role handling when dropping cast
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 15:15:24 +00:00
mateo
fd7f186929 fix: satisfy type-discipline and prettier gates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 15:06:02 +00:00
mateo
8c2e910d8a fix(proxy): unblock CI - drop banned casts, gate missing request method, fix ui lint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 15:01:22 +00:00
Devin AI
2bc38b0b65 Merge remote-tracking branch 'origin/main' into litellm_refuse_example_master_key
# Conflicts:
#	tests/test_litellm/proxy/auth/test_user_api_key_auth.py
#	tests/test_litellm/proxy/test__types.py
#	tests/test_litellm/test_utils.py
2026-09-15 14:24:56 +00:00
mateo
eaf1c15569 chore(ui): drop generated files swept into the previous commit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 14:11:02 +00:00
mateo
ed886c518b revert(ui): keep the simplified key generate error mapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 14:10:47 +00:00
Devin AI
5b7690cd2f fix(ui): surface the master key lockout message in credential, model and key toasts 2026-09-15 14:10:15 +00:00
Devin AI
a9e8e196c0 style(proxy): format lockout condition 2026-09-15 13:54:46 +00:00
Devin AI
29971b3011 refactor(proxy): drop unused lockout action parameter 2026-09-15 13:54:27 +00:00
Devin AI
74c0e1a339 feat(proxy): lock out credential storage and virtual key management while the master key is insecure 2026-09-15 13:52:30 +00:00
mateo
b9b0f0e7ea Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_refuse_example_master_key 2026-09-15 13:36:33 +00:00
Mateo Wang
c8114ba41f
Merge pull request #40921 from BerriAI/litellm_unified_key_policy_hook
feat(proxy): unified custom_key_policy hook for key generate, update and regenerate
2026-09-15 06:34:28 -07:00
Mateo Wang
9496f16f12
Merge pull request #41173 from BerriAI/litellm_realtime_health_check_credential_name
fix(health): resolve litellm_credential_name in realtime health checks
2026-09-15 05:24:29 -07:00
mateo-berri
1f2d050386 Merge remote-tracking branch 'origin/main' into litellm_unified_key_policy_hook 2026-09-15 05:11:29 -07:00
mateo-berri
87b630429d fix(realtime): send openai and xai health check keys as bearer tokens 2026-09-15 02:20:09 -07:00
Mateo Wang
3ed6c19b8d
Merge pull request #40915 from BerriAI/litellm_internal_copy_37075
fix(vertex-live): bill Gemini Live sessions end to end (internal copy of #37075)
2026-09-15 02:06:49 -07:00
Mateo Wang
c274fd8781
Merge pull request #41168 from BerriAI/litellm_bedrock_wif_session_policy_coverage
fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy
2026-09-15 01:49:41 -07:00
mateo-berri
5969fbb052 refactor(realtime): freeze the empty credential mapping in the health check 2026-09-15 01:44:46 -07:00
mateo-berri
0770f663c1 fix(vertex-live): report repeated search queries once per grounded turn
The session usage collapsed duplicate query strings across turns while the
price was per turn, so two turns asking the same question paid two fees yet
reported web_search_requests 1. Sum each turn's grounding requests so the
counter matches the bill; duplicates within one turn still collapse.
2026-09-15 01:35:36 -07:00
mateo-berri
1771255b32 Merge remote-tracking branch 'origin/main' into litellm_realtime_health_check_credential_name 2026-09-15 01:17:31 -07:00
mateo-berri
d5938ff886 Merge remote-tracking branch 'origin/main' into litellm_bedrock_wif_session_policy_coverage 2026-09-15 01:16:45 -07:00
Mateo Wang
80b9ed4f2c
Merge pull request #41191 from BerriAI/litellm_router_test_cap_resets_per_fallback_hop
fix(router): count num_retries_per_request across fallback hops
2026-09-15 01:13:41 -07:00
Mateo Wang
e5cb8b7534
Merge pull request #40984 from BerriAI/litellm_anthropic_guardrail_system_and_tool_use
fix(guardrails): scan the Anthropic top-level system prompt and tool_use arguments
2026-09-15 00:55:42 -07:00
mateo-berri
1b040af414 test(router): type the retry-cap tests this PR adds or touches 2026-09-15 00:34:38 -07:00
mateo-berri
4f27573424 merge: origin/main into litellm_internal_copy_37075 2026-09-15 00:34:18 -07:00
yuneng-jiang
81863c1b17
Merge pull request #41188 from BerriAI/litellm_spend_reconciliation
test(spend): reconcile concurrent requests and daily activity
2026-09-15 00:32:20 -07:00
tin-berri
feab83aae1
Merge pull request #41186 from BerriAI/litellm_statusline_router_cost_label
fix(cli): label savings cost bars with the auto-router name
2026-09-15 00:32:08 -07:00
Tin Chi Lo
81340439fc fix(router): preserve session model choice within each complexity tier 2026-09-15 00:09:17 -07:00
mateo-berri
92714cac0c fix(guardrails): validate tool_use rewrites before writing text rewrites back
A guardrail that rewrites text and hands back tool_use arguments that are
not a JSON object used to leave the text rewrite applied when the request
was rejected, so failure logging saw a half-rewritten request. Every
rejection now happens before any write to system or messages.
2026-09-15 00:07:45 -07:00
mateo-berri
f80cb5cb46 fix(router): ignore planted request_retry_count seeds and cover the rust OCR cap path
The router clamps a negative request_retry_count found in request metadata before counting a failure, and the proxy strips a client-supplied request_retry_count with the other router-reserved metadata fields. The rust OCR lifecycle test that trips the per-request cap now plants request_retry_count instead of attempted_retries, which the cap no longer reads since the previous commit
2026-09-15 00:04:02 -07:00
yuneng-jiang
7e3ca1421e
Merge pull request #41194 from BerriAI/litellm_stream_tool_contract
test(e2e): verify streamed answers and tool continuation
2026-09-15 00:01:49 -07:00
Tin Chi Lo
d0fdf1c237 fix(cli): show only the routed model in the footer header 2026-09-14 23:54:50 -07:00
Tin Chi Lo
4fca818f34 fix(cli): align wide and combining Unicode cost labels 2026-09-14 23:45:48 -07:00
tin-berri
4805c6d51f
Merge pull request #41072 from BerriAI/litellm_lit5201_provider_split
fix(router): honor team and key provider weights
2026-09-14 23:41:45 -07:00
Mateo Wang
1ce66e98a2
Merge pull request #40989 from BerriAI/litellm_responses_bridge_hoist_additional_tools
fix(responses): hoist Codex additional_tools input items into the chat bridge tools
2026-09-14 23:33:38 -07:00
Tin Chi Lo
398300c4e7 fix(router): honor team and key provider weights 2026-09-14 23:31:52 -07:00
Mateo Wang
b52de1675a
Merge pull request #40994 from BerriAI/litellm_sdk_exception_body_headers
fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm_proxy 400
2026-09-14 23:27:29 -07:00
mateo-berri
2bf44ed354 fix(guardrails): reject tool_use rewrites that are not JSON objects 2026-09-14 23:18:47 -07:00
mateo-berri
aaf924693a fix(router): count num_retries_per_request across fallback hops
num_retries_per_request has always capped the retries of one request with its fallback hops included. #40930 started reading the per-hop attempted_retries counter instead, and every fallback hop restarts that counter at zero, so a request could spend a fresh retry budget on each hop and the legacy fallback cap test started seeing the hop run.

Router.log_retry now also keeps request_retry_count on the request metadata, incremented on every retry and fallback hop and never truncated the way previous_models is, and max_retries_per_request_hit reads that count. The flat retry records, the litellm_metadata coverage and caps above four from #40930 stay as they are, and the legacy test goes back to its previous_models == 0 assertion.
2026-09-14 23:13:50 -07:00
Mateo Wang
c93708b2a5
Merge pull request #40228 from AaronHowell/litellm_fix_responses_credentials_affinity
fix(responses): preserve provider affinity
2026-09-14 23:09:55 -07:00
mateo-berri
9fb94ea761 fix(exceptions): keep repeated litellm_proxy response headers on the rebuilt response
httpx.Headers.items() comma-joins repeated header names, so the rebuilt
response iterates multi_items() and keeps every value, matching what the
raw openai client exposes on e.response.headers
2026-09-14 22:59:40 -07:00
mateo-berri
e3152c011d fix(responses): classify streamed tool calls on the chat name and strip guardrail edits around the grammar block
The streaming bridge restored the namespace before deciding whether a tool call was a custom tool, so a namespaced function sharing a short name with a nested custom tool streamed back as a custom_tool_call. Classify on the raw chat tool name first, the way the non-streaming path already does.

The guardrail merge only stripped the namespace prefix and grammar suffix from the ends of the edited description, so a guardrail appending text after the grammar block left the block in the member description and the chat conversion appended it a second time. Strip the first occurrence of each instead.
2026-09-14 22:58:25 -07:00