berriai-litellm-provider-info-sync[bot]
41caaa301c
chore(prices): sync prices for 2 providers: 235 models, 59 new [enrichment failed: Google Gemini, 34 held]
...
azure_ai/Codestral-2501:
azure_ai/cohere-command-a:
azure_ai/deepseek-r1:
azure_ai/deepseek-v3:
azure_ai/deepseek-v3-0324:
azure_ai/deepseek-v3.1:
azure_ai/deepseek-v3.2:
azure_ai/deepseek-v3.2-speciale:
azure_ai/deepseek-v4-flash:
azure_ai/DeepSeek-V4-Flash-0731:
azure_ai/deepseek-v4-pro:
azure_ai/embed-v-4-0:
azure_ai/FW-DeepSeek-V3.2:
azure_ai/FW-DeepSeek-V4-Pro:
azure_ai/FW-GLM-5:
azure_ai/FW-GLM-5.1:
azure_ai/FW-GLM-5.2:
azure_ai/FW-GLM-5.2-Fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Inkling: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Kimi-K2.5:
azure_ai/FW-Kimi-K2.6:
azure_ai/FW-Kimi-K2.7-Code:
azure_ai/FW-Kimi-K3:
azure_ai/FW-MiniMax-M2.5:
azure_ai/FW-MiniMax-M3:
azure_ai/FW-Nemotron-3-Ultra-NVFP4: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B:
azure_ai/gpt-oss-120b:
azure_ai/grok-3:
azure_ai/global/grok-3:
azure_ai/grok-3-mini:
azure_ai/global/grok-3-mini:
azure_ai/grok-4:
azure_ai/grok-4-1-fast-non-reasoning:
azure_ai/grok-4-1-fast-reasoning:
azure_ai/grok-4-20-non-reasoning:
azure_ai/grok-4-20-reasoning:
azure_ai/grok-4-fast-non-reasoning:
azure_ai/grok-4-fast-reasoning:
azure_ai/grok-4.3:
azure_ai/grok-4.6:
azure_ai/grok-code-fast-1:
azure_ai/kimi-k2.5:
azure_ai/kimi-k2.6:
azure_ai/kimi-k2.7-code:
azure_ai/Llama-3.3-70B-Instruct:
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: input_cost_per_token, output_cost_per_token
azure_ai/MAI-DS-R1:
azure_ai/MAI-Image-2.5:
azure_ai/MAI-Image-2.5-Flash:
azure_ai/MAI-Image-2e:
azure_ai/MAI-Thinking-1:
azure_ai/mistral-large-3:
azure_ai/Phi-3-medium-128k-instruct:
azure_ai/Phi-3-medium-4k-instruct:
azure_ai/Phi-3-mini-128k-instruct:
azure_ai/Phi-3-mini-4k-instruct:
azure_ai/Phi-3-small-128k-instruct:
azure_ai/Phi-3-small-8k-instruct:
azure_ai/Phi-3.5-mini-instruct:
2026-09-15 18:46:36 +00:00
berriai-litellm-provider-info-sync[bot]
a511c9d45d
chore(prices): sync Google Gemini prices: 7 models [enrichment failed: Google Gemini, 86 held]
...
gemini/gemini-3.1-flash-image:
gemini/gemini-3.1-flash-lite: cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3.1-flash-lite-image:
gemini-3.1-flash-live-preview:
gemini/gemini-3.1-flash-live-preview:
gemini/gemini-3.1-flash-tts-preview: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-15 18:41:29 +00:00
berriai-litellm-provider-info-sync[bot]
1517f1205c
chore(prices): sync Google Gemini prices: 10 models [enrichment failed: Google Gemini, 177 held]
...
gemini/gemini-2.5-computer-use-preview-10-2025:
gemini/gemini-2.5-flash: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini/gemini-2.5-flash-image: input_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority
gemini/gemini-2.5-flash-lite: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini-2.5-flash-preview-tts:
gemini/gemini-2.5-flash-preview-tts:
gemini/gemini-2.5-pro: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority
gemini/gemini-2.5-pro-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_priority, output_cost_per_token_priority
2026-09-15 18:36:52 +00:00
berriai-litellm-provider-info-sync[bot]
363b7835a3
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
a8dd1394f9
chore(prices): drop source from us-gov rows again
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
1b69cfcf3c
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
68f97321dd
chore(prices): drop source from us-gov rows to match usgov pricing test
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
8a9a08caf0
chore(prices): sync prices for 2 providers: 14 models
...
chatgpt-image-latest: input_cost_per_image_token_batches
gemini-2.0-flash: input_cost_per_audio_token_batches
gemini-2.0-flash-lite: input_cost_per_audio_token_batches
gemini-2.5-flash: input_cost_per_audio_token_batches
gemini-2.5-flash-lite: input_cost_per_audio_token_batches
gemini-3-flash-preview: input_cost_per_audio_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_audio_token_batches
gemini-3.1-flash-lite: input_cost_per_audio_token_batches
vertex_ai/gemini-3.1-flash-lite: input_cost_per_audio_token_batches
gpt-image-1: input_cost_per_image_token_batches
gpt-image-1-mini: input_cost_per_image_token_batches
gpt-image-1.5: input_cost_per_image_token_batches
gpt-image-1.5-2025-12-16: input_cost_per_image_token_batches
gpt-image-2: input_cost_per_image_token_batches
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
01b069eb6c
chore(prices): sync AWS Bedrock prices: 25 models
...
anthropic.claude-fable-5:
anthropic.claude-fable-5-1:
anthropic.claude-opus-4-7:
anthropic.claude-opus-4-8:
anthropic.claude-opus-5:
anthropic.claude-sonnet-4-6:
anthropic.claude-sonnet-5:
global.anthropic.claude-fable-5:
global.anthropic.claude-fable-5-1:
global.anthropic.claude-opus-4-7:
global.anthropic.claude-opus-4-8:
global.anthropic.claude-opus-5:
global.anthropic.claude-sonnet-4-6:
global.anthropic.claude-sonnet-5:
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
us.anthropic.claude-fable-5:
us.anthropic.claude-fable-5-1:
us.anthropic.claude-opus-4-7:
us.anthropic.claude-opus-4-8:
us.anthropic.claude-opus-5:
us.anthropic.claude-sonnet-4-6:
us.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
dc3a0399f2
chore(prices): sync Vertex AI prices: 4 models
...
gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini-3.1-flash-live-preview: input_cost_per_second
gemini/gemini-3.1-flash-live-preview: input_cost_per_second
2026-09-15 17:42:56 +00:00
kerry-berri
b97bc10ec9
Merge pull request #41263 from BerriAI/litellm_usgov_rows_allow_price_list_source
...
test(pricing): let synced GovCloud Bedrock rows cite the AWS price list
2026-09-15 10:41:42 -07:00
Devin AI
9fce6e1989
test(pricing): let synced GovCloud Bedrock rows cite the AWS price list
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:09:12 +00:00
tin-berri
3ad9a7f336
Merge pull request #41174 from BerriAI/litellm_tier_model_affinity
...
fix(router): preserve session model choice within each complexity tier
2026-09-15 09:54:53 -07:00
Mateo Wang
c8114ba41f
Merge pull request #40921 from BerriAI/litellm_unified_key_policy_hook
...
feat(proxy): unified custom_key_policy hook for key generate, update and regenerate
2026-09-15 06:34:28 -07:00
Mateo Wang
9496f16f12
Merge pull request #41173 from BerriAI/litellm_realtime_health_check_credential_name
...
fix(health): resolve litellm_credential_name in realtime health checks
2026-09-15 05:24:29 -07:00
mateo-berri
1f2d050386
Merge remote-tracking branch 'origin/main' into litellm_unified_key_policy_hook
2026-09-15 05:11:29 -07:00
mateo-berri
87b630429d
fix(realtime): send openai and xai health check keys as bearer tokens
2026-09-15 02:20:09 -07:00
Mateo Wang
3ed6c19b8d
Merge pull request #40915 from BerriAI/litellm_internal_copy_37075
...
fix(vertex-live): bill Gemini Live sessions end to end (internal copy of #37075 )
2026-09-15 02:06:49 -07:00
Mateo Wang
c274fd8781
Merge pull request #41168 from BerriAI/litellm_bedrock_wif_session_policy_coverage
...
fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy
2026-09-15 01:49:41 -07:00
mateo-berri
5969fbb052
refactor(realtime): freeze the empty credential mapping in the health check
2026-09-15 01:44:46 -07:00
mateo-berri
0770f663c1
fix(vertex-live): report repeated search queries once per grounded turn
...
The session usage collapsed duplicate query strings across turns while the
price was per turn, so two turns asking the same question paid two fees yet
reported web_search_requests 1. Sum each turn's grounding requests so the
counter matches the bill; duplicates within one turn still collapse.
2026-09-15 01:35:36 -07:00
mateo-berri
1771255b32
Merge remote-tracking branch 'origin/main' into litellm_realtime_health_check_credential_name
2026-09-15 01:17:31 -07:00
mateo-berri
d5938ff886
Merge remote-tracking branch 'origin/main' into litellm_bedrock_wif_session_policy_coverage
2026-09-15 01:16:45 -07:00
Mateo Wang
80b9ed4f2c
Merge pull request #41191 from BerriAI/litellm_router_test_cap_resets_per_fallback_hop
...
fix(router): count num_retries_per_request across fallback hops
2026-09-15 01:13:41 -07:00
Mateo Wang
e5cb8b7534
Merge pull request #40984 from BerriAI/litellm_anthropic_guardrail_system_and_tool_use
...
fix(guardrails): scan the Anthropic top-level system prompt and tool_use arguments
2026-09-15 00:55:42 -07:00
mateo-berri
1b040af414
test(router): type the retry-cap tests this PR adds or touches
2026-09-15 00:34:38 -07:00
mateo-berri
4f27573424
merge: origin/main into litellm_internal_copy_37075
2026-09-15 00:34:18 -07:00
yuneng-jiang
81863c1b17
Merge pull request #41188 from BerriAI/litellm_spend_reconciliation
...
test(spend): reconcile concurrent requests and daily activity
2026-09-15 00:32:20 -07:00
tin-berri
feab83aae1
Merge pull request #41186 from BerriAI/litellm_statusline_router_cost_label
...
fix(cli): label savings cost bars with the auto-router name
2026-09-15 00:32:08 -07:00
Tin Chi Lo
81340439fc
fix(router): preserve session model choice within each complexity tier
2026-09-15 00:09:17 -07:00
mateo-berri
92714cac0c
fix(guardrails): validate tool_use rewrites before writing text rewrites back
...
A guardrail that rewrites text and hands back tool_use arguments that are
not a JSON object used to leave the text rewrite applied when the request
was rejected, so failure logging saw a half-rewritten request. Every
rejection now happens before any write to system or messages.
2026-09-15 00:07:45 -07:00
mateo-berri
f80cb5cb46
fix(router): ignore planted request_retry_count seeds and cover the rust OCR cap path
...
The router clamps a negative request_retry_count found in request metadata before counting a failure, and the proxy strips a client-supplied request_retry_count with the other router-reserved metadata fields. The rust OCR lifecycle test that trips the per-request cap now plants request_retry_count instead of attempted_retries, which the cap no longer reads since the previous commit
2026-09-15 00:04:02 -07:00
yuneng-jiang
7e3ca1421e
Merge pull request #41194 from BerriAI/litellm_stream_tool_contract
...
test(e2e): verify streamed answers and tool continuation
2026-09-15 00:01:49 -07:00
Tin Chi Lo
d0fdf1c237
fix(cli): show only the routed model in the footer header
2026-09-14 23:54:50 -07:00
Tin Chi Lo
4fca818f34
fix(cli): align wide and combining Unicode cost labels
2026-09-14 23:45:48 -07:00
tin-berri
4805c6d51f
Merge pull request #41072 from BerriAI/litellm_lit5201_provider_split
...
fix(router): honor team and key provider weights
2026-09-14 23:41:45 -07:00
Mateo Wang
1ce66e98a2
Merge pull request #40989 from BerriAI/litellm_responses_bridge_hoist_additional_tools
...
fix(responses): hoist Codex additional_tools input items into the chat bridge tools
2026-09-14 23:33:38 -07:00
Tin Chi Lo
398300c4e7
fix(router): honor team and key provider weights
2026-09-14 23:31:52 -07:00
Mateo Wang
b52de1675a
Merge pull request #40994 from BerriAI/litellm_sdk_exception_body_headers
...
fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm_proxy 400
2026-09-14 23:27:29 -07:00
mateo-berri
2bf44ed354
fix(guardrails): reject tool_use rewrites that are not JSON objects
2026-09-14 23:18:47 -07:00
mateo-berri
aaf924693a
fix(router): count num_retries_per_request across fallback hops
...
num_retries_per_request has always capped the retries of one request with its fallback hops included. #40930 started reading the per-hop attempted_retries counter instead, and every fallback hop restarts that counter at zero, so a request could spend a fresh retry budget on each hop and the legacy fallback cap test started seeing the hop run.
Router.log_retry now also keeps request_retry_count on the request metadata, incremented on every retry and fallback hop and never truncated the way previous_models is, and max_retries_per_request_hit reads that count. The flat retry records, the litellm_metadata coverage and caps above four from #40930 stay as they are, and the legacy test goes back to its previous_models == 0 assertion.
2026-09-14 23:13:50 -07:00
Mateo Wang
c93708b2a5
Merge pull request #40228 from AaronHowell/litellm_fix_responses_credentials_affinity
...
fix(responses): preserve provider affinity
2026-09-14 23:09:55 -07:00
mateo-berri
9fb94ea761
fix(exceptions): keep repeated litellm_proxy response headers on the rebuilt response
...
httpx.Headers.items() comma-joins repeated header names, so the rebuilt
response iterates multi_items() and keeps every value, matching what the
raw openai client exposes on e.response.headers
2026-09-14 22:59:40 -07:00
mateo-berri
e3152c011d
fix(responses): classify streamed tool calls on the chat name and strip guardrail edits around the grammar block
...
The streaming bridge restored the namespace before deciding whether a tool call was a custom tool, so a namespaced function sharing a short name with a nested custom tool streamed back as a custom_tool_call. Classify on the raw chat tool name first, the way the non-streaming path already does.
The guardrail merge only stripped the namespace prefix and grammar suffix from the ends of the edited description, so a guardrail appending text after the grammar block left the block in the member description and the chat conversion appended it a second time. Strip the first occurrence of each instead.
2026-09-14 22:58:25 -07:00
Yuneng Jiang
7a7770db0d
test(e2e): verify streamed answers and tool continuation
2026-09-14 22:46:55 -07:00
mateo-berri
62b2b36ce9
test(proxy): drop the reformat-only diff of the request processing tests
...
The proxy edge test file no longer carries any test of this change, and
the remaining diff was the scoped format gate reflowing the whole file to
the 120 limit, so it goes back to the merge base bytes
2026-09-14 22:39:19 -07:00
yucheng-berri
91588221cd
Merge pull request #39050 from BerriAI/litellm_lit6314_guardrail_metadata_transfer
...
fix(guardrails): record not_run evaluation when scoping leaves nothing to scan
2026-09-14 22:31:16 -07:00
mateo-berri
6764ab2673
test(router): assert num_retries_per_request as a per-group cap that resets per fallback hop
...
#40930 (LIT-7505) changed num_retries_per_request from a request-wide
cap to a per-model-group cap that resets on every fallback hop, and its
own comment in litellm/__init__.py names that contract. The legacy
test_async_fallbacks_max_retries_per_request still asserted the old
request-wide reading (previous_models == 0), so the CircleCI router
suite has been red on main since that merge for every run-ci PR.
The test now reads the flat RetryAttemptRecord entries the fallback
call carries and asserts the new contract directly: every record is
from the first group, the retry at attempted_retries 0 is the real
AuthenticationError, and each later attempt was refused with
"Max retries per request hit!".
2026-09-14 22:29:32 -07:00
Mateo Wang
a78b24c195
Merge pull request #41172 from BerriAI/litellm_azure_spend_log_zero_cost
...
fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate
2026-09-14 22:26:22 -07:00
Yuneng Jiang
80d804d6f9
test(spend): preserve multi-day coverage and immutable assertions
2026-09-14 22:26:20 -07:00