Yassin Kortam
95abc9fb0b
Merge pull request #41330 from BerriAI/litellm_team_model_max_budget_v2
...
feat(team): team-level model_max_budget with key-level overrides
2026-09-16 14:48:29 -07:00
Yassin Kortam
cd08c65002
Merge pull request #41425 from BerriAI/litellm_streaming_buffer_release_on_scan
...
feat(guardrails): release buffered stream chunks after each passing scan
2026-09-16 14:37:56 -07:00
Mateo Wang
6edb549dbd
Merge pull request #41112 from BerriAI/litellm_registry_audit_2026_09_14
...
fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching
2026-09-16 14:31:50 -07:00
kerry-berri
930ec9643a
Merge pull request #41446 from BerriAI/litellm_fix_anthropic_stream_served_model
...
fix(anthropic): carry the served model from message_start onto stream chunks
2026-09-16 14:14:54 -07:00
Mateo Wang
8e524370e1
Merge pull request #41343 from BerriAI/litellm_lit5030_bedrock_invoke_nova_prompt_caching
...
fix(bedrock): make prompt caching work on the Nova InvokeModel route
2026-09-16 13:42:21 -07:00
Yassin Kortam
c4a9341ef6
Merge pull request #34829 from max-sixty/bugfix/http-handler-del-closes-streaming-client
...
fix(http_handler): keep a handler alive while a response it issued is still reading
2026-09-16 13:39:04 -07:00
yassin
857301444e
Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan
2026-09-16 20:34:36 +00:00
yassin
8acd2477a6
fix(guardrails): hold legacy function_call stream windows until the end-of-stream scan
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 20:34:20 +00:00
mateo-berri
6015437d67
test(bedrock): assert the Nova cache-read rate as a discount instead of pinning the vendor ratio
2026-09-16 13:19:58 -07:00
Mateo Wang
737929f338
Merge pull request #41335 from BerriAI/litellm_fireworks_dict_reasoning_effort
...
fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string
2026-09-16 13:12:37 -07:00
mateo-berri
884087f01c
test(bedrock): type the vector store search test helper
2026-09-16 12:55:29 -07:00
mateo-berri
0dff64ce1a
fix(bedrock): move a Nova invoke cache point behind an image or tool result to the last text block
2026-09-16 12:42:15 -07:00
yassin
73a0c3bb8a
Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan
2026-09-16 19:09:19 +00:00
Devin AI
03931180d7
merge main into litellm_registry_audit_2026_09_14
2026-09-16 19:03:36 +00:00
yassin
ad25fe1886
fix(guardrails): hold unscannable Responses windows and key terminal envelopes by output items
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:56:56 +00:00
mateo-berri
033aa8ba6d
fix(bedrock): forward userContext in Knowledge Base Retrieve requests
...
The Bedrock vector store search only lifted retrievalConfiguration out of extra_body, so the caller's userContext (the Retrieve API's ACL identity) never reached Bedrock and ACL-enabled data sources answered with zero results. The transform now forwards userContext, taken from extra_body first and then from the top-level params where the OpenAI SDK's extra_body merge lands, as the caller sent it.
2026-09-16 11:50:36 -07:00
yassin
0143fe5583
fix(guardrails): hold tool-call windows until the final scan and expose Bedrock streaming flags to the UI
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:30:12 +00:00
kerry
686556fffc
test(anthropic): build served-model stream chunks immutably
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:49:15 +00:00
kerry
07d4936428
fix(anthropic): satisfy strict lint and update tests pinned to the dropped served model
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:44:59 +00:00
kerry-berri
4e996400e2
Merge pull request #41154 from BerriAI/litellm-providers/price-sync
...
chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated
2026-09-16 10:37:21 -07:00
kerry
847172d311
fix(anthropic): carry the served model from message_start onto stream chunks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:31:04 +00:00
Yassin Kortam
8491d01668
Merge pull request #41268 from BerriAI/litellm_outbound_http2_opt_in
...
feat(http): opt-in outbound HTTP/2 for httpx clients
2026-09-16 10:05:37 -07:00
kerry
cab6732928
test: drop price-pinning tests that break on catalog updates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:32:37 +00:00
yassin
04aad317c7
Merge remote-tracking branch 'origin/main' into pr34829
2026-09-16 16:02:08 +00:00
Devin AI
ba820783ac
fix(tests): import Bedrock usage types
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:25:14 +00:00
Devin AI
4fd69b25e1
Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14
2026-09-16 13:05:28 +00:00
kerry-berri
4b84fa9230
Merge pull request #41336 from BerriAI/litellm_fix_anthropic_stream_absent_usage
...
fix(anthropic): tolerate message_delta events without usage when streaming
2026-09-15 20:23:01 -07:00
Yassin Kortam
d108cdc431
Merge pull request #38254 from BerriAI/litellm_fix_xai_responses_instructions
2026-09-15 19:55:52 -07:00
kerry-berri
399420bb25
Merge pull request #41339 from BerriAI/litellm_fix_fireworks_cost_components
...
fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator
2026-09-15 19:16:24 -07:00
kerry
133f1e8ef5
test(fireworks-ai): drop the explanatory comment on the cache-read constant
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 02:06:42 +00:00
kerry-berri
54e1b9e3de
Merge pull request #41338 from BerriAI/litellm_fix_gemini_model_version
...
fix(gemini): propagate the provider's modelVersion to the response model
2026-09-15 19:04:50 -07:00
kerry
091a38cce6
test(anthropic): import Final for the annotated locals
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:43:08 +00:00
kerry
c4c96180e3
fix(gemini): strip version suffix from modelVersion and keep it on blocked streams
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:42:36 +00:00
kerry
8352045f32
style(fireworks-ai): apply repository conventions to the cost component change
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:39:24 +00:00
Devin AI
8d625d0400
fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web_search bridging
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:35:13 +00:00
kerry
10aef224e7
style(anthropic): apply repository conventions to the missing-usage change
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:34:28 +00:00
kerry
d1aa37b0a8
merge main into litellm_fix_fireworks_cost_components
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:28:29 +00:00
kerry
44a6d19889
test(gemini): drop review narration from the wrapper test docstring
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:14:37 +00:00
kerry
801f6a92b4
test(gemini): assert the served modelVersion reaches the assembled stream through CustomStreamWrapper
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
mateo-berri
3c15f64fd4
fix(bedrock): make prompt caching work on the Nova InvokeModel route
...
Nova InvokeModel rejects the standalone cachePoint blocks the shared Converse transform emits, so each one is folded into the block it caches and tool_config injection points are dropped before the transform runs, since this route has no tool caching to credit. Usage reads Bedrock's Count-suffixed cache keys and adds cached tokens into prompt_tokens, streaming routes every wrapped InvokeModel event through the Converse chunk parser and tolerates the missing totalTokens, and the Nova 1 cost-map entries gain cache_read_input_token_cost at a quarter of the input rate
2026-09-15 18:01:21 -07:00
kerry
5eab1feb20
fix(fireworks-ai): bill cache-write, reasoning, and audio tokens via the shared cost calculator
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:42:13 +00:00
yassin
260ff5f491
feat(team): team-level model_max_budget with key-level overrides
...
A team can now carry a per-model budget map that every key on the team
inherits. A key's own model_max_budget entry for the same model takes
precedence, so it is gated on and billed to the key alone.
Backend: NewTeamRequest/UpdateTeamRequest accept model_max_budget (validated
like the key-level field, enterprise gated); the value is hydrated onto
UserAPIKeyAuth via the token view, TeamGrants and the carried budget state;
_check_team_model_budget enforces it in the centralized common checks; the
limiter meters spend under team_model_spend:<team>:<model>:<duration> and
skips the team counter when the key overrides; /team/update lets only a
proxy admin raise, re-window or drop a cap; /team/info exposes usage.
The Anthropic context-management compaction summary subrequest runs the
same team gate. Both fallback token-view SQL definitions project the column.
UI: team create and edit forms reuse the key-level ModelMaxBudgetEditor,
premium gated, sending {} to clear and omitting unchanged fields.
A key entry overrides the team cap only when it spend-gates the model
(non-negative max_budget); a row that only carries tpm/rpm limits or a
negative cap leaves the team cap in force.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:40:58 +00:00
Devin AI
a7a61db78d
fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:37:55 +00:00
kerry
d5a36eb2ca
fix(gemini): propagate provider modelVersion onto model responses
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:36:07 +00:00
kerry
ba6bcd747f
chore: merge origin/main into test cleanup
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:33:25 +00:00
Yassin Kortam
7eeba69016
Merge pull request #41316 from BerriAI/litellm_nvidia_nim_infer_passthrough
...
feat(proxy): add /nvidia_nim passthrough route for NIM object detection and OCR /v1/infer
2026-09-15 17:28:25 -07:00
kerry
4657d43fb9
fix(anthropic): tolerate message_delta chunks without usage in streams
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:27:09 +00:00
Yassin Kortam
636313e9fc
Merge pull request #41322 from BerriAI/litellm_gcp_video_usage
...
fix(vertex_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough
2026-09-15 17:21:15 -07:00
yassin
771b2509d1
fix(proxy): reject mixed NIM model groups and strip the deployment model before the group in /nvidia_nim URLs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:34:22 +00:00
kerry
e64371062c
test: remove test files left empty by the cleanup
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:04:40 +00:00