Commit graph

2157 commits

Author SHA1 Message Date
kerry
b9ab36279b fix(prices): add tpm and rpm to gemini 3.8 live rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:49:17 +00:00
berriai-litellm-provider-info-sync[bot]
809685603a
chore(prices): sync Azure prices: 247 models
azure_ai/Codestral-2501: 
azure_ai/cohere-command-a: 
azure_ai/deepseek-r1: 
azure_ai/deepseek-v3: 
azure_ai/deepseek-v3-0324: 
azure_ai/deepseek-v3.1: 
azure_ai/deepseek-v3.2: 
azure_ai/deepseek-v3.2-speciale: 
azure_ai/deepseek-v4-flash: 
azure_ai/DeepSeek-V4-Flash-0731: 
azure_ai/deepseek-v4-pro: 
azure_ai/embed-v-4-0: 
azure_ai/FW-DeepSeek-V3.2: 
azure_ai/FW-DeepSeek-V4-Pro: 
azure_ai/FW-GLM-5: 
azure_ai/FW-GLM-5.1: 
azure_ai/FW-GLM-5.2: 
azure_ai/FW-GLM-5.2-Fast: 
azure_ai/FW-Inkling: 
azure_ai/FW-Kimi-K2.5: 
azure_ai/FW-Kimi-K2.6: 
azure_ai/FW-Kimi-K2.7-Code: 
azure_ai/FW-Kimi-K3: 
azure_ai/FW-MiniMax-M2.5: 
azure_ai/FW-MiniMax-M3: 
azure_ai/FW-Nemotron-3-Ultra-NVFP4: 
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B: 
azure_ai/gpt-oss-120b: 
azure_ai/grok-3: 
azure_ai/global/grok-3: 
azure_ai/grok-3-mini: 
azure_ai/global/grok-3-mini: 
azure_ai/grok-4: 
azure_ai/grok-4-1-fast-non-reasoning: 
azure_ai/grok-4-1-fast-reasoning: 
azure_ai/grok-4-20-non-reasoning: 
azure_ai/grok-4-20-reasoning: 
azure_ai/grok-4-fast-non-reasoning: 
azure_ai/grok-4-fast-reasoning: 
azure_ai/grok-4.3: 
azure_ai/grok-4.6: 
azure_ai/grok-code-fast-1: 
azure_ai/kimi-k2.5: 
azure_ai/kimi-k2.6: 
azure_ai/kimi-k2.7-code: 
azure_ai/Llama-3.3-70B-Instruct: 
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: 
azure_ai/MAI-DS-R1: 
azure_ai/MAI-Image-2.5: 
azure_ai/MAI-Image-2.5-Flash: 
azure_ai/MAI-Image-2e: 
azure_ai/MAI-Thinking-1: 
azure_ai/mistral-large-3: 
azure_ai/Phi-3-medium-128k-instruct: 
azure_ai/Phi-3-medium-4k-instruct: 
azure_ai/Phi-3-mini-128k-instruct: 
azure_ai/Phi-3-mini-4k-instruct: 
azure_ai/Phi-3-small-128k-instruct: 
azure_ai/Phi-3-small-8k-instruct: 
azure_ai/Phi-3.5-mini-instruct:
2026-09-16 16:26:23 +00:00
berriai-litellm-provider-info-sync[bot]
0443605c40
chore(prices): sync Together AI prices: 4 models, 4 deprecated
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-16 13:16:12 +00:00
berriai-litellm-provider-info-sync[bot]
641dcd5f8f
chore(prices): sync Google Gemini prices: 1 model
gemini/gemini-robotics-er-2-preview: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
2026-09-16 03:46:34 +00:00
berriai-litellm-provider-info-sync[bot]
c5a388a94c
chore(prices): sync Google Gemini prices: 1 model [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 6 held]
gemini/gemini-3.1-pro-preview-customtools: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-16 01:46:17 +00:00
kerry
91c8d1cdc1 chore: merge origin/main into litellm-providers/price-sync
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:30:24 +00:00
kerry-berri
daa98cd6ff
Merge pull request #41320 from BerriAI/litellm_gemini_fallback_generalizations
feat(model_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization
2026-09-15 17:28:58 -07:00
kerry
c9dd4b44f8 revert(model_info): keep Gemini baseline majors single-digit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:51:56 +00:00
kerry
945adfb603 refactor(model_info): support multi-digit Gemini majors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:49:18 +00:00
kerry
5b54bf2328 refactor(model_info): restore provider-neutral Gemini chat baseline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:22 +00:00
berriai-litellm-provider-info-sync[bot]
083ecb3c0b
chore(prices): sync prices for 3 providers: 26 models, 26 deprecated [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 34 held]
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation_date
fireworks_ai/deepseek-v4-pro: deprecation_date
fireworks_ai/accounts/fireworks/models/minimax-m2p7: deprecation_date
fireworks_ai/minimax-m2p7: deprecation_date
together_ai/deepseek-ai/deepseek-coder-33b-instruct: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Llama-70B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B: deprecation_date
together_ai/google/gemma-2-27b-it: deprecation_date
vertex_ai/imagegeneration@006: deprecation_date
vertex_ai/imagen-3.0-capability-001: deprecation_date
vertex_ai/imagen-3.0-fast-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-002: deprecation_date
vertex_ai/imagen-4.0-fast-generate-001: deprecation_date
vertex_ai/imagen-4.0-generate-001: deprecation_date
vertex_ai/imagen-4.0-ultra-generate-001: deprecation_date
together_ai/meta-llama/Llama-3-8b-chat-hf: deprecation_date
together_ai/meta-llama/Meta-Llama-3-70B-Instruct-Turbo: deprecation_date
together_ai/meta-llama/Meta-Llama-3-8B-Instruct: deprecation_date
together_ai/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO: deprecation_date
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF: deprecation_date
together_ai/Qwen/Qwen2-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2-VL-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-Coder-32B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-VL-72B-Instruct: deprecation_date
2026-09-15 23:31:18 +00:00
kerry
726430c0e7 refactor(model_info): scope Gemini baseline to first-party chat ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:16:32 +00:00
tin-berri
8cd00d2d6e
Merge pull request #41282 from BerriAI/litellm_fast_mode_toggle_0915
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
ai-gateway image / ai-gateway release image (push) Has been cancelled
feat(auto-router): add per-model Fast mode toggle
2026-09-15 16:10:06 -07:00
kerry
082f02bd73 refactor(model_info): drop Gemini vertex routing rule, keep baseline to certain capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:55:29 +00:00
kerry
024a887521 feat(model_info): add Gemini fallback generalization rules (vertex routing + 2.5+ family baseline)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:49:35 +00:00
Tin Chi Lo
864f4a7a0e feat(auto-router): add per-model Fast mode toggle 2026-09-15 12:45:16 -07:00
berriai-litellm-provider-info-sync[bot]
9f4fbcfe41
chore(prices): sync Azure prices: 14 models [enrichment failed: Google Gemini, 34 held]
azure/eu/gpt-5-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/eu/gpt-5-mini-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/eu/gpt-5-nano-2025-08-07: input_cost_per_token_batches, output_cost_per_token_batches
azure/eu/o1-mini-2024-09-12: 
azure/eu/o1-preview-2024-09-12: 
azure/o1-mini-2024-09-12: input_cost_per_token_batches, output_cost_per_token_batches
azure/us/gpt-4.1-2025-04-14: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-4.1-mini-2025-04-14: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-4.1-nano-2025-04-14: cache_read_input_token_cost, input_cost_per_token_batches
azure/us/gpt-5-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-5-mini-2025-08-07: input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_priority
azure/us/gpt-5-nano-2025-08-07: input_cost_per_token_batches, output_cost_per_token_batches
azure/us/o1-mini-2024-09-12: 
azure/us/o1-preview-2024-09-12:
2026-09-15 18:51:21 +00:00
berriai-litellm-provider-info-sync[bot]
41caaa301c
chore(prices): sync prices for 2 providers: 235 models, 59 new [enrichment failed: Google Gemini, 34 held]
azure_ai/Codestral-2501: 
azure_ai/cohere-command-a: 
azure_ai/deepseek-r1: 
azure_ai/deepseek-v3: 
azure_ai/deepseek-v3-0324: 
azure_ai/deepseek-v3.1: 
azure_ai/deepseek-v3.2: 
azure_ai/deepseek-v3.2-speciale: 
azure_ai/deepseek-v4-flash: 
azure_ai/DeepSeek-V4-Flash-0731: 
azure_ai/deepseek-v4-pro: 
azure_ai/embed-v-4-0: 
azure_ai/FW-DeepSeek-V3.2: 
azure_ai/FW-DeepSeek-V4-Pro: 
azure_ai/FW-GLM-5: 
azure_ai/FW-GLM-5.1: 
azure_ai/FW-GLM-5.2: 
azure_ai/FW-GLM-5.2-Fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Inkling: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Kimi-K2.5: 
azure_ai/FW-Kimi-K2.6: 
azure_ai/FW-Kimi-K2.7-Code: 
azure_ai/FW-Kimi-K3: 
azure_ai/FW-MiniMax-M2.5: 
azure_ai/FW-MiniMax-M3: 
azure_ai/FW-Nemotron-3-Ultra-NVFP4: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B: 
azure_ai/gpt-oss-120b: 
azure_ai/grok-3: 
azure_ai/global/grok-3: 
azure_ai/grok-3-mini: 
azure_ai/global/grok-3-mini: 
azure_ai/grok-4: 
azure_ai/grok-4-1-fast-non-reasoning: 
azure_ai/grok-4-1-fast-reasoning: 
azure_ai/grok-4-20-non-reasoning: 
azure_ai/grok-4-20-reasoning: 
azure_ai/grok-4-fast-non-reasoning: 
azure_ai/grok-4-fast-reasoning: 
azure_ai/grok-4.3: 
azure_ai/grok-4.6: 
azure_ai/grok-code-fast-1: 
azure_ai/kimi-k2.5: 
azure_ai/kimi-k2.6: 
azure_ai/kimi-k2.7-code: 
azure_ai/Llama-3.3-70B-Instruct: 
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: input_cost_per_token, output_cost_per_token
azure_ai/MAI-DS-R1: 
azure_ai/MAI-Image-2.5: 
azure_ai/MAI-Image-2.5-Flash: 
azure_ai/MAI-Image-2e: 
azure_ai/MAI-Thinking-1: 
azure_ai/mistral-large-3: 
azure_ai/Phi-3-medium-128k-instruct: 
azure_ai/Phi-3-medium-4k-instruct: 
azure_ai/Phi-3-mini-128k-instruct: 
azure_ai/Phi-3-mini-4k-instruct: 
azure_ai/Phi-3-small-128k-instruct: 
azure_ai/Phi-3-small-8k-instruct: 
azure_ai/Phi-3.5-mini-instruct:
2026-09-15 18:46:36 +00:00
berriai-litellm-provider-info-sync[bot]
a511c9d45d
chore(prices): sync Google Gemini prices: 7 models [enrichment failed: Google Gemini, 86 held]
gemini/gemini-3.1-flash-image: 
gemini/gemini-3.1-flash-lite: cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3.1-flash-lite-image: 
gemini-3.1-flash-live-preview: 
gemini/gemini-3.1-flash-live-preview: 
gemini/gemini-3.1-flash-tts-preview: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-15 18:41:29 +00:00
berriai-litellm-provider-info-sync[bot]
1517f1205c
chore(prices): sync Google Gemini prices: 10 models [enrichment failed: Google Gemini, 177 held]
gemini/gemini-2.5-computer-use-preview-10-2025: 
gemini/gemini-2.5-flash: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini/gemini-2.5-flash-image: input_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority
gemini/gemini-2.5-flash-lite: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches, cache_read_input_token_cost_priority
gemini-2.5-flash-preview-tts: 
gemini/gemini-2.5-flash-preview-tts: 
gemini/gemini-2.5-pro: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority
gemini/gemini-2.5-pro-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost, input_cost_per_audio_token_batches
gemini/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_priority, output_cost_per_token_priority
2026-09-15 18:36:52 +00:00
berriai-litellm-provider-info-sync[bot]
363b7835a3 chore(prices): sync AWS Bedrock prices: 4 models
us-gov.anthropic.claude-fable-5-1: 
us-gov.anthropic.claude-opus-4-8: 
us-gov.anthropic.claude-opus-5: 
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
a8dd1394f9 chore(prices): drop source from us-gov rows again
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
1b69cfcf3c chore(prices): sync AWS Bedrock prices: 4 models
us-gov.anthropic.claude-fable-5-1: 
us-gov.anthropic.claude-opus-4-8: 
us-gov.anthropic.claude-opus-5: 
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
68f97321dd chore(prices): drop source from us-gov rows to match usgov pricing test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
8a9a08caf0 chore(prices): sync prices for 2 providers: 14 models
chatgpt-image-latest: input_cost_per_image_token_batches
gemini-2.0-flash: input_cost_per_audio_token_batches
gemini-2.0-flash-lite: input_cost_per_audio_token_batches
gemini-2.5-flash: input_cost_per_audio_token_batches
gemini-2.5-flash-lite: input_cost_per_audio_token_batches
gemini-3-flash-preview: input_cost_per_audio_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_audio_token_batches
gemini-3.1-flash-lite: input_cost_per_audio_token_batches
vertex_ai/gemini-3.1-flash-lite: input_cost_per_audio_token_batches
gpt-image-1: input_cost_per_image_token_batches
gpt-image-1-mini: input_cost_per_image_token_batches
gpt-image-1.5: input_cost_per_image_token_batches
gpt-image-1.5-2025-12-16: input_cost_per_image_token_batches
gpt-image-2: input_cost_per_image_token_batches
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
01b069eb6c chore(prices): sync AWS Bedrock prices: 25 models
anthropic.claude-fable-5: 
anthropic.claude-fable-5-1: 
anthropic.claude-opus-4-7: 
anthropic.claude-opus-4-8: 
anthropic.claude-opus-5: 
anthropic.claude-sonnet-4-6: 
anthropic.claude-sonnet-5: 
global.anthropic.claude-fable-5: 
global.anthropic.claude-fable-5-1: 
global.anthropic.claude-opus-4-7: 
global.anthropic.claude-opus-4-8: 
global.anthropic.claude-opus-5: 
global.anthropic.claude-sonnet-4-6: 
global.anthropic.claude-sonnet-5: 
us-gov.anthropic.claude-fable-5-1: 
us-gov.anthropic.claude-opus-4-8: 
us-gov.anthropic.claude-opus-5: 
us-gov.anthropic.claude-sonnet-5: 
us.anthropic.claude-fable-5: 
us.anthropic.claude-fable-5-1: 
us.anthropic.claude-opus-4-7: 
us.anthropic.claude-opus-4-8: 
us.anthropic.claude-opus-5: 
us.anthropic.claude-sonnet-4-6: 
us.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
dc3a0399f2 chore(prices): sync Vertex AI prices: 4 models
gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini-3.1-flash-live-preview: input_cost_per_second
gemini/gemini-3.1-flash-live-preview: input_cost_per_second
2026-09-15 17:42:56 +00:00
IToSSc
8607c49ea1 feat: add aihubmix provider pricing entries
Add 72 model price entries for the aihubmix openai_like provider so
cost tracking and budgets work for aihubmix/* model calls. The
provider is already registered in llms/openai_like/providers.json
but model_prices_and_context_window.json had zero entries for it.

The Anthropic-family entries (claude-fable-5, claude-haiku-4-5,
claude-opus-4-8, claude-opus-5, claude-sonnet-5) carry the same
supports_adaptive_thinking, thinking_always_on,
supports_sampling_params, and prompt_cache_min_tokens flags already
used by this repo's other Anthropic re-exports (azure_ai, databricks,
openrouter, and so on) for the same underlying models, since those
flags gate request shapes the provider otherwise rejects with a 400.

TASK-2BK38Y
2026-09-15 14:16:20 +08:00
kerry
ce83fac351 fix(cost): bill batch embeddings per modality token rate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:50:21 +00:00
kerry
a28ea22ec1 fix(cost): move gemini-embedding-2-preview to per-token rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:34:22 +00:00
kerry
4a8ec7b9d8 fix(cost): bill gemini-embedding-2-preview per token like the GA entries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:18:50 +00:00
kerry
d4f2119b03 fix(cost): bill gemini-embedding-2 per token and stop double charging audio
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:41:46 +00:00
berriai-litellm-provider-info-sync[bot]
c3f8c07c43
chore(prices): sync Vertex AI prices: 14 models
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost
vertex_ai/gemini-2.5-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost
vertex_ai/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority
vertex_ai/gemini-3.1-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex
vertex_ai/gemini-3.1-flash-lite: cache_read_input_audio_token_cost
vertex_ai/gemini-3.1-flash-lite-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
vertex_ai/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
vertex_ai/gemini-3.5-flash: 
vertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_priority
vertex_ai/gemini-3.6-flash: 
vertex_ai/gemini-3.7-flash: 
vertex_ai/gemini-3.8-flash: 
vertex_ai/gemini-embedding-2:
2026-09-14 22:00:57 +00:00
Devin AI
e4ebeae800 fix(prices): add text output rate to Vertex TTS entries, sync gemini-embedding-2 alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:49:15 +00:00
Devin AI
f45e20e6c2 fix(prices): mark Vertex Gemini TTS entries as audio_speech
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:41:43 +00:00
kerry
abbf5aae20 chore(prices): flag gpt-5.5-cyber as a reasoning model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:41:10 +00:00
berriai-litellm-provider-info-sync[bot]
93e65995d6
chore(prices): sync Vertex AI prices: 4 models
gemini-2.5-flash: cache_read_input_audio_token_cost
gemini-2.5-flash-lite: cache_read_input_audio_token_cost
gemini-3-flash-preview: cache_read_input_audio_token_cost
gemini-3.1-flash-lite: cache_read_input_audio_token_cost
2026-09-14 21:26:02 +00:00
kerry
ddf13e9505 Merge remote-tracking branch 'origin/main' into litellm-providers/price-sync 2026-09-14 21:21:55 +00:00
kerry
250ff03a04 feat(model_info): scope fill_missing rules to azure, bedrock and vertex hosts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
kerry
812bbee0b3 fix(model_info): scope fill_missing backfill to the rule's providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
fb2057fde7 refactor(model_info): rename backfill_exact_entries to fill_missing_fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
42c708670d fix(model_info): guard backfill by mode, drop provider key, tighten claude major regex
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
cee7215b24 feat(model_info): opt-in field-level backfill from fallback generalization rules
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Mateo Wang
cab1e113f7
Merge pull request #40976 from BerriAI/litellm_azure_gpt_chat_latest_pricing
feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries
2026-09-14 11:21:20 -07:00
mateo
1661e72c2c fix(registry): drop retired friendliai llama-3.1 serverless models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 13:21:46 +00:00
mateo-berri
4c022a3089 feat(pricing): add azure gpt-chat-latest global and data zone rates 2026-09-12 23:34:12 -07:00
berriai-litellm-provider-info-sync[bot]
6423acc11a
chore(prices): sync prices for 5 providers: 278 models, 34 new
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp: 
fireworks_ai/deepseek-v4-flash-vision-exp: 
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/muse-glimmer-30b: 
fireworks_ai/muse-glimmer-30b: 
fireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4: 
fireworks_ai/nemotron-3-ultra-nvfp4: 
fireworks_ai/accounts/fireworks/models/qwen3-embedding-8b: 
fireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token
fireworks_ai/accounts/fireworks/models/qwen3p7-plus: 
fireworks_ai/qwen3p7-plus: 
fireworks_ai/accounts/fireworks/models/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/routers/glm-5p2-fast: 
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
fireworks_ai/accounts/fireworks/routers/kimi-k3-fast: 
together_ai/arcee-ai/trinity-mini: input_cost_per_token, output_cost_per_token
together_ai/arize-ai/qwen-2-1.5b-instruct: 
babbage-002: input_cost_per_token_batches, output_cost_per_token_batches
chat-latest: 
chatgpt-image-latest: output_cost_per_token, input_cost_per_image_token, output_cost_per_image_token, input_cost_per_token_batches, output_cost_per_token_batches
claude-fable-5: 
claude-fable-5-1: 
claude-haiku-4-5: 
claude-mythos-5: 
claude-mythos-5-1: 
claude-opus-4-5: 
claude-opus-4-6: 
claude-opus-4-7: 
claude-opus-4-8: 
claude-opus-5: 
claude-sonnet-4-5: 
claude-sonnet-4-6: 
claude-sonnet-5: 
davinci-002: input_cost_per_token_batches, output_cost_per_token_batches
deep-research-pro-preview-12-2025: cache_read_input_token_cost
together_ai/deepseek-ai/deepseek-coder-33b-instruct: input_cost_per_token, output_cost_per_token
together_ai/deepseek-ai/DeepSeek-R1-0528:
2026-09-13 04:55:39 +00:00
shivam
a28e595a9d Merge remote-tracking branch 'origin/main' into litellm_fix_realtime_cached_audio_cost 2026-09-13 04:24:18 +00:00
Mateo Wang
b1a61f510c
Merge pull request #35918 from Lee-Si-Yoon/feat/friendli-model-metadata-sync
feat(friendli): auto-sync Friendli model metadata into price registry
2026-09-12 21:13:52 -07:00
mateo-berri
a94c060b84 fix(cost): fill the missing realtime cache-read rates 2026-09-12 17:48:00 -07:00