kerry
5e5e5086f0
Merge remote-tracking branch 'origin/main' into litellm-providers/price-sync
2026-09-17 19:29:18 +00:00
berriai-litellm-provider-info-sync[bot]
730195c603
chore(prices): sync Together AI prices: 6 models, 6 deprecated [sync failed: Google Gemini]
...
together_ai/deepseek-ai/DeepSeek-V4-Flash-0731: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4-Pro-0813: deprecation_date
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-17 18:01:11 +00:00
kerry
3d0fd127d5
feat(openrouter): add stealth/union-alpha to the model cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:20:16 +00:00
berriai-litellm-provider-info-sync[bot]
5a82ed9bca
chore(prices): sync prices for 2 providers: 27 models
...
fireworks_ai/accounts/fireworks/models/minimax-m3: supports_vision
fireworks_ai/minimax-m3: supports_vision
wandb/deepseek-ai/DeepSeek-V4-Flash: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Flash-0731: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro-0813: max_input_tokens
wandb/google/gemma-4-31B-it: max_input_tokens
wandb/ibm-granite/granite-4.1-8b: max_input_tokens
wandb/ibm-granite/granite-4.2-8b: max_input_tokens
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-70B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-8B-Instruct: max_input_tokens
wandb/MiniMaxAI/MiniMax-M3: max_input_tokens
wandb/moonshotai/Kimi-K2.6: max_input_tokens
wandb/moonshotai/Kimi-K2.7-Code: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: max_input_tokens
wandb/openai/gpt-oss-120b: max_input_tokens
wandb/openai/gpt-oss-20b: max_input_tokens
wandb/OpenPipe/Qwen3-14B-Instruct: max_input_tokens
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: max_input_tokens
wandb/Qwen/Qwen3.5-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.6-27B: max_input_tokens
wandb/Qwen/Qwen3.6-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.8-27B: max_input_tokens
wandb/zai-org/GLM-5.2: max_input_tokens
wandb/zai-org/GLM-5.3-Flash: max_input_tokens
2026-09-17 06:31:05 +00:00
berriai-litellm-provider-info-sync[bot]
b4212b949b
chore(prices): sync prices for 5 providers: 34 models, 1 new, 19 deprecated [1 with gaps]
...
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast:
azure_ai/FW-Kimi-K3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/deepseek-ai/DeepSeek-R1-0528: deprecation_date
wandb/deepseek-ai/DeepSeek-V3-0324: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Flash: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Pro: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4.1-Flash:
azure/eu/gpt-5.5-2026-04-24:
gemini-3.8-live: supports_response_schema
gemini-3.8-live-extended-thinking: supports_response_schema
azure/gpt-5.5-2026-04-24:
azure/gpt-5.6-luna-2026-07-09:
azure/gpt-5.6-sol-2026-07-09:
azure/gpt-5.6-terra-2026-07-09:
azure/gpt-6-astra-2026-09-03:
wandb/ibm-granite/granite-4.1-8b: deprecation_date
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: deprecation_date
wandb/meta-llama/Llama-3.1-70B-Instruct: deprecation_date
wandb/meta-llama/Llama-4-Scout-17B-16E-Instruct: deprecation_date
wandb/microsoft/Phi-4-mini-instruct: deprecation_date
wandb/MiniMaxAI/MiniMax-M2.5: deprecation_date
wandb/moonshotai/Kimi-K2-Instruct: deprecation_date
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/OpenPipe/Qwen3-14B-Instruct: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Thinking-2507: deprecation_date
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-Coder-480B-A35B-Instruct: deprecation_date
wandb/Qwen/Qwen3.5-35B-A3B: deprecation_date
wandb/Qwen/Qwen3.6-27B: deprecation_date
azure/us/gpt-5.5-2026-04-24:
wandb/zai-org/GLM-4.5: deprecation_date
wandb/zai-org/GLM-5.3-Flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, supports_function_calling, supports_tool_choice, supports_response_schema, supports_prompt_caching, supports_reasoning
2026-09-17 06:01:14 +00:00
Mateo Wang
319f427c40
Merge pull request #41201 from BerriAI/litellm_gemini_37_38_flash_no_minimal_thinking
...
fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash
2026-09-16 16:52:08 -07:00
mateo-berri
a1ad95dbbd
fix(gemini): read the minimal thinking floor from the cost map and cover the /v1/messages bridge
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
2026-09-16 16:13:41 -07:00
Yuneng Jiang
a6b10ad654
fix(prices): dedupe Nova cache_read_input_token_cost keys left by a text merge
...
PR #41343 and PR #41112 both added cache_read_input_token_cost to the six amazon.nova-{micro,lite,pro}-v1:0 and us.amazon.nova-* entries, one at the top of each entry and one at the bottom. The merge kept both, so every PR now fails test_price_map_has_no_duplicate_keys. Both copies carried the same value, so this only removes the trailing duplicate in both price files
2026-09-16 14:58:57 -07:00
Mateo Wang
6edb549dbd
Merge pull request #41112 from BerriAI/litellm_registry_audit_2026_09_14
...
fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching
2026-09-16 14:31:50 -07:00
Mateo Wang
8e524370e1
Merge pull request #41343 from BerriAI/litellm_lit5030_bedrock_invoke_nova_prompt_caching
...
fix(bedrock): make prompt caching work on the Nova InvokeModel route
2026-09-16 13:42:21 -07:00
Devin AI
ba03f60f71
fix(models): add batch and flex tier prices on Azure dated and Gemini latest alias rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:50:28 +00:00
Devin AI
93eefc5922
fix(models): dedupe Fireworks deprecation keys and sync Azure dated snapshot service tiers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:20:20 +00:00
Devin AI
8caf2cb61f
fix(models): flag prompt caching on Vertex and Azure AI Grok rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:17:45 +00:00
Devin AI
7734e3e186
fix(models): sonnet 4.5 1M input, daybreak alias, Mistral GLM 5.3, Azure dated snapshots, Together/OpenRouter sync
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:17:13 +00:00
berriai-litellm-provider-info-sync[bot]
b893e6b926
chore(prices): sync Google Gemini prices: 22 models
...
gemini/gemini-2.5-flash: max_tokens, max_output_tokens, supports_audio_input
gemini/gemini-2.5-flash-image: max_input_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-lite: max_tokens, max_output_tokens, supports_audio_input
gemini-2.5-flash-native-audio-preview-12-2025: supports_vision, max_input_tokens, supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-native-audio-preview-12-2025: supports_vision, max_input_tokens, supports_web_search, supports_response_schema, supports_function_calling
gemini-2.5-flash-preview-tts: max_tokens, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-preview-tts: max_tokens, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-pro: max_tokens, max_output_tokens
gemini/gemini-2.5-pro-preview-tts: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-3-flash-preview: max_tokens, max_output_tokens, supports_audio_input
gemini/gemini-3-pro-image: supports_response_schema
gemini/gemini-3.1-flash-image: max_input_tokens, supports_response_schema
gemini/gemini-3.1-flash-lite-image: supports_web_search, supports_function_calling
gemini-3.1-flash-live-preview: supports_response_schema
gemini/gemini-3.1-flash-live-preview: supports_response_schema
gemini/gemini-3.1-flash-tts-preview: supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-3.5-flash: max_tokens, max_output_tokens
gemini/gemini-3.5-live-translate-preview: supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-3.5-transcribe: supports_function_calling
gemini/gemini-3.5-transcribe-live: supports_function_calling
gemini/gemini-embedding-2: supports_vision, supports_audio_input
gemini/gemini-omni-1.1-flash: max_input_tokens
2026-09-16 19:06:19 +00:00
Devin AI
03931180d7
merge main into litellm_registry_audit_2026_09_14
2026-09-16 19:03:36 +00:00
berriai-litellm-provider-info-sync[bot]
f8fb31db3a
chore(prices): sync Azure prices: 2 models
...
azure_ai/grok-4.3: input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens
azure_ai/grok-4.6: input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens
2026-09-16 17:51:09 +00:00
kerry
b9ab36279b
fix(prices): add tpm and rpm to gemini 3.8 live rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:49:17 +00:00
berriai-litellm-provider-info-sync[bot]
809685603a
chore(prices): sync Azure prices: 247 models
...
azure_ai/Codestral-2501:
azure_ai/cohere-command-a:
azure_ai/deepseek-r1:
azure_ai/deepseek-v3:
azure_ai/deepseek-v3-0324:
azure_ai/deepseek-v3.1:
azure_ai/deepseek-v3.2:
azure_ai/deepseek-v3.2-speciale:
azure_ai/deepseek-v4-flash:
azure_ai/DeepSeek-V4-Flash-0731:
azure_ai/deepseek-v4-pro:
azure_ai/embed-v-4-0:
azure_ai/FW-DeepSeek-V3.2:
azure_ai/FW-DeepSeek-V4-Pro:
azure_ai/FW-GLM-5:
azure_ai/FW-GLM-5.1:
azure_ai/FW-GLM-5.2:
azure_ai/FW-GLM-5.2-Fast:
azure_ai/FW-Inkling:
azure_ai/FW-Kimi-K2.5:
azure_ai/FW-Kimi-K2.6:
azure_ai/FW-Kimi-K2.7-Code:
azure_ai/FW-Kimi-K3:
azure_ai/FW-MiniMax-M2.5:
azure_ai/FW-MiniMax-M3:
azure_ai/FW-Nemotron-3-Ultra-NVFP4:
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B:
azure_ai/gpt-oss-120b:
azure_ai/grok-3:
azure_ai/global/grok-3:
azure_ai/grok-3-mini:
azure_ai/global/grok-3-mini:
azure_ai/grok-4:
azure_ai/grok-4-1-fast-non-reasoning:
azure_ai/grok-4-1-fast-reasoning:
azure_ai/grok-4-20-non-reasoning:
azure_ai/grok-4-20-reasoning:
azure_ai/grok-4-fast-non-reasoning:
azure_ai/grok-4-fast-reasoning:
azure_ai/grok-4.3:
azure_ai/grok-4.6:
azure_ai/grok-code-fast-1:
azure_ai/kimi-k2.5:
azure_ai/kimi-k2.6:
azure_ai/kimi-k2.7-code:
azure_ai/Llama-3.3-70B-Instruct:
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8:
azure_ai/MAI-DS-R1:
azure_ai/MAI-Image-2.5:
azure_ai/MAI-Image-2.5-Flash:
azure_ai/MAI-Image-2e:
azure_ai/MAI-Thinking-1:
azure_ai/mistral-large-3:
azure_ai/Phi-3-medium-128k-instruct:
azure_ai/Phi-3-medium-4k-instruct:
azure_ai/Phi-3-mini-128k-instruct:
azure_ai/Phi-3-mini-4k-instruct:
azure_ai/Phi-3-small-128k-instruct:
azure_ai/Phi-3-small-8k-instruct:
azure_ai/Phi-3.5-mini-instruct:
2026-09-16 16:26:23 +00:00
Devin AI
642525f571
fix(registry): add azure gpt-image-2.5 entries and together/azure deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:19:53 +00:00
berriai-litellm-provider-info-sync[bot]
0443605c40
chore(prices): sync Together AI prices: 4 models, 4 deprecated
...
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-16 13:16:12 +00:00
Devin AI
4fd69b25e1
Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14
2026-09-16 13:05:28 +00:00
berriai-litellm-provider-info-sync[bot]
641dcd5f8f
chore(prices): sync Google Gemini prices: 1 model
...
gemini/gemini-robotics-er-2-preview: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
2026-09-16 03:46:34 +00:00
berriai-litellm-provider-info-sync[bot]
c5a388a94c
chore(prices): sync Google Gemini prices: 1 model [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 6 held]
...
gemini/gemini-3.1-pro-preview-customtools: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-16 01:46:17 +00:00
kerry
91c8d1cdc1
chore: merge origin/main into litellm-providers/price-sync
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:30:24 +00:00
mateo-berri
3c15f64fd4
fix(bedrock): make prompt caching work on the Nova InvokeModel route
...
Nova InvokeModel rejects the standalone cachePoint blocks the shared Converse transform emits, so each one is folded into the block it caches and tool_config injection points are dropped before the transform runs, since this route has no tool caching to credit. Usage reads Bedrock's Count-suffixed cache keys and adds cached tokens into prompt_tokens, streaming routes every wrapped InvokeModel event through the Converse chunk parser and tolerates the missing totalTokens, and the Nova 1 cost-map entries gain cache_read_input_token_cost at a quarter of the input rate
2026-09-15 18:01:21 -07:00
kerry-berri
daa98cd6ff
Merge pull request #41320 from BerriAI/litellm_gemini_fallback_generalizations
...
feat(model_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization
2026-09-15 17:28:58 -07:00
kerry
c9dd4b44f8
revert(model_info): keep Gemini baseline majors single-digit
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:51:56 +00:00
kerry
945adfb603
refactor(model_info): support multi-digit Gemini majors
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:49:18 +00:00
kerry
5b54bf2328
refactor(model_info): restore provider-neutral Gemini chat baseline
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:22 +00:00
berriai-litellm-provider-info-sync[bot]
083ecb3c0b
chore(prices): sync prices for 3 providers: 26 models, 26 deprecated [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 34 held]
...
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation_date
fireworks_ai/deepseek-v4-pro: deprecation_date
fireworks_ai/accounts/fireworks/models/minimax-m2p7: deprecation_date
fireworks_ai/minimax-m2p7: deprecation_date
together_ai/deepseek-ai/deepseek-coder-33b-instruct: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Llama-70B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B: deprecation_date
together_ai/google/gemma-2-27b-it: deprecation_date
vertex_ai/imagegeneration@006: deprecation_date
vertex_ai/imagen-3.0-capability-001: deprecation_date
vertex_ai/imagen-3.0-fast-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-002: deprecation_date
vertex_ai/imagen-4.0-fast-generate-001: deprecation_date
vertex_ai/imagen-4.0-generate-001: deprecation_date
vertex_ai/imagen-4.0-ultra-generate-001: deprecation_date
together_ai/meta-llama/Llama-3-8b-chat-hf: deprecation_date
together_ai/meta-llama/Meta-Llama-3-70B-Instruct-Turbo: deprecation_date
together_ai/meta-llama/Meta-Llama-3-8B-Instruct: deprecation_date
together_ai/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO: deprecation_date
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF: deprecation_date
together_ai/Qwen/Qwen2-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2-VL-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-Coder-32B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-VL-72B-Instruct: deprecation_date
2026-09-15 23:31:18 +00:00
kerry
726430c0e7
refactor(model_info): scope Gemini baseline to first-party chat ids
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:16:32 +00:00
tin-berri
8cd00d2d6e
Merge pull request #41282 from BerriAI/litellm_fast_mode_toggle_0915
...
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
ai-gateway image / ai-gateway release image (push) Has been cancelled
feat(auto-router): add per-model Fast mode toggle
2026-09-15 16:10:06 -07:00
kerry
082f02bd73
refactor(model_info): drop Gemini vertex routing rule, keep baseline to certain capabilities
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:55:29 +00:00
kerry
024a887521
feat(model_info): add Gemini fallback generalization rules (vertex routing + 2.5+ family baseline)
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:49:35 +00:00
Devin AI
11211cbf29
fix(registry): drop gemini/gemini-3.8-live entries pending published rate limits
...
The gemini/ realtime test requires tpm and rpm, and Google publishes the
Live model limits only behind the AI Studio login, so the direct-API keys
cannot be sourced yet. The Vertex keys stay.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 20:18:29 +00:00
Tin Chi Lo
864f4a7a0e
feat(auto-router): add per-model Fast mode toggle
2026-09-15 12:45:16 -07:00
Devin AI
93f27d6082
fix(registry): add gemini 3.8 live, azure gpt-5.5/luna snapshots, doubao seed 2.1, fix together v4.1 flash context
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 19:15:46 +00:00
berriai-litellm-provider-info-sync[bot]
9f4fbcfe41
chore(prices): sync Azure prices: 14 models [enrichment failed: Google Gemini, 34 held]
...
azure/eu/gpt-5-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/eu/gpt-5-mini-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/eu/gpt-5-nano-2025-08-07: input_cost_per_token_batches, output_cost_per_token_batches
azure/eu/o1-mini-2024-09-12:
azure/eu/o1-preview-2024-09-12:
azure/o1-mini-2024-09-12: input_cost_per_token_batches, output_cost_per_token_batches
azure/us/gpt-4.1-2025-04-14: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-4.1-mini-2025-04-14: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-4.1-nano-2025-04-14: cache_read_input_token_cost, input_cost_per_token_batches
azure/us/gpt-5-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-5-mini-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-5-nano-2025-08-07: input_cost_per_token_batches, output_cost_per_token_batches
azure/us/o1-mini-2024-09-12:
azure/us/o1-preview-2024-09-12:
2026-09-15 18:51:21 +00:00
berriai-litellm-provider-info-sync[bot]
41caaa301c
chore(prices): sync prices for 2 providers: 235 models, 59 new [enrichment failed: Google Gemini, 34 held]
...
azure_ai/Codestral-2501:
azure_ai/cohere-command-a:
azure_ai/deepseek-r1:
azure_ai/deepseek-v3:
azure_ai/deepseek-v3-0324:
azure_ai/deepseek-v3.1:
azure_ai/deepseek-v3.2:
azure_ai/deepseek-v3.2-speciale:
azure_ai/deepseek-v4-flash:
azure_ai/DeepSeek-V4-Flash-0731:
azure_ai/deepseek-v4-pro:
azure_ai/embed-v-4-0:
azure_ai/FW-DeepSeek-V3.2:
azure_ai/FW-DeepSeek-V4-Pro:
azure_ai/FW-GLM-5:
azure_ai/FW-GLM-5.1:
azure_ai/FW-GLM-5.2:
azure_ai/FW-GLM-5.2-Fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Inkling: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Kimi-K2.5:
azure_ai/FW-Kimi-K2.6:
azure_ai/FW-Kimi-K2.7-Code:
azure_ai/FW-Kimi-K3:
azure_ai/FW-MiniMax-M2.5:
azure_ai/FW-MiniMax-M3:
azure_ai/FW-Nemotron-3-Ultra-NVFP4: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B:
azure_ai/gpt-oss-120b:
azure_ai/grok-3:
azure_ai/global/grok-3:
azure_ai/grok-3-mini:
azure_ai/global/grok-3-mini:
azure_ai/grok-4:
azure_ai/grok-4-1-fast-non-reasoning:
azure_ai/grok-4-1-fast-reasoning:
azure_ai/grok-4-20-non-reasoning:
azure_ai/grok-4-20-reasoning:
azure_ai/grok-4-fast-non-reasoning:
azure_ai/grok-4-fast-reasoning:
azure_ai/grok-4.3:
azure_ai/grok-4.6:
azure_ai/grok-code-fast-1:
azure_ai/kimi-k2.5:
azure_ai/kimi-k2.6:
azure_ai/kimi-k2.7-code:
azure_ai/Llama-3.3-70B-Instruct:
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: input_cost_per_token, output_cost_per_token
azure_ai/MAI-DS-R1:
azure_ai/MAI-Image-2.5:
azure_ai/MAI-Image-2.5-Flash:
azure_ai/MAI-Image-2e:
azure_ai/MAI-Thinking-1:
azure_ai/mistral-large-3:
azure_ai/Phi-3-medium-128k-instruct:
azure_ai/Phi-3-medium-4k-instruct:
azure_ai/Phi-3-mini-128k-instruct:
azure_ai/Phi-3-mini-4k-instruct:
azure_ai/Phi-3-small-128k-instruct:
azure_ai/Phi-3-small-8k-instruct:
azure_ai/Phi-3.5-mini-instruct:
2026-09-15 18:46:36 +00:00
berriai-litellm-provider-info-sync[bot]
a511c9d45d
chore(prices): sync Google Gemini prices: 7 models [enrichment failed: Google Gemini, 86 held]
...
gemini/gemini-3.1-flash-image:
gemini/gemini-3.1-flash-lite: cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3.1-flash-lite-image:
gemini-3.1-flash-live-preview:
gemini/gemini-3.1-flash-live-preview:
gemini/gemini-3.1-flash-tts-preview: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-15 18:41:29 +00:00
berriai-litellm-provider-info-sync[bot]
1517f1205c
chore(prices): sync Google Gemini prices: 10 models [enrichment failed: Google Gemini, 177 held]
...
gemini/gemini-2.5-computer-use-preview-10-2025:
gemini/gemini-2.5-flash: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini/gemini-2.5-flash-image: input_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority
gemini/gemini-2.5-flash-lite: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini-2.5-flash-preview-tts:
gemini/gemini-2.5-flash-preview-tts:
gemini/gemini-2.5-pro: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority
gemini/gemini-2.5-pro-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_priority, output_cost_per_token_priority
2026-09-15 18:36:52 +00:00
berriai-litellm-provider-info-sync[bot]
363b7835a3
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
a8dd1394f9
chore(prices): drop source from us-gov rows again
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
1b69cfcf3c
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
68f97321dd
chore(prices): drop source from us-gov rows to match usgov pricing test
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
8a9a08caf0
chore(prices): sync prices for 2 providers: 14 models
...
chatgpt-image-latest: input_cost_per_image_token_batches
gemini-2.0-flash: input_cost_per_audio_token_batches
gemini-2.0-flash-lite: input_cost_per_audio_token_batches
gemini-2.5-flash: input_cost_per_audio_token_batches
gemini-2.5-flash-lite: input_cost_per_audio_token_batches
gemini-3-flash-preview: input_cost_per_audio_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_audio_token_batches
gemini-3.1-flash-lite: input_cost_per_audio_token_batches
vertex_ai/gemini-3.1-flash-lite: input_cost_per_audio_token_batches
gpt-image-1: input_cost_per_image_token_batches
gpt-image-1-mini: input_cost_per_image_token_batches
gpt-image-1.5: input_cost_per_image_token_batches
gpt-image-1.5-2025-12-16: input_cost_per_image_token_batches
gpt-image-2: input_cost_per_image_token_batches
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
01b069eb6c
chore(prices): sync AWS Bedrock prices: 25 models
...
anthropic.claude-fable-5:
anthropic.claude-fable-5-1:
anthropic.claude-opus-4-7:
anthropic.claude-opus-4-8:
anthropic.claude-opus-5:
anthropic.claude-sonnet-4-6:
anthropic.claude-sonnet-5:
global.anthropic.claude-fable-5:
global.anthropic.claude-fable-5-1:
global.anthropic.claude-opus-4-7:
global.anthropic.claude-opus-4-8:
global.anthropic.claude-opus-5:
global.anthropic.claude-sonnet-4-6:
global.anthropic.claude-sonnet-5:
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
us.anthropic.claude-fable-5:
us.anthropic.claude-fable-5-1:
us.anthropic.claude-opus-4-7:
us.anthropic.claude-opus-4-8:
us.anthropic.claude-opus-5:
us.anthropic.claude-sonnet-4-6:
us.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
dc3a0399f2
chore(prices): sync Vertex AI prices: 4 models
...
gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini-3.1-flash-live-preview: input_cost_per_second
gemini/gemini-3.1-flash-live-preview: input_cost_per_second
2026-09-15 17:42:56 +00:00
Devin AI
3000c00c50
fix(models): add supports_response_schema to fireworks deepseek-v4-flash-vision-exp short alias
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 14:03:24 +00:00