Commit graph

2403 commits

Author SHA1 Message Date
berriai-litellm-provider-info-sync[bot]
28755b98a0
chore(prices): sync Fireworks AI prices: 2 models, 2 new [2 with gaps] (#42590)
* chore(prices): sync Fireworks AI prices: 2 models, 2 new [2 with gaps]

fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: max_input_tokens, supports_tool_choice, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority, max_output_tokens, max_tokens, supports_vision, supports_reasoning
fireworks_ai/accounts/fireworks/models/minimax-m2p7: max_input_tokens, supports_tool_choice, supports_response_schema, supports_function_calling, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority, max_output_tokens, max_tokens

* chore(prices): sync Fireworks AI prices: 2 models, 2 deprecated

fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation_date
fireworks_ai/accounts/fireworks/models/minimax-m2p7: deprecation_date

Price-Sync: litellm-providers

* feat(prices): add fireworks_ai/accounts/fireworks/models/ember-1

Prices, context length and capability flags read from the Fireworks serverless models API on 2026-09-23. Smoke tested with a live completion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prices): resolve merge conflict markers left in the cost map merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 14:47:01 -07:00
devin-ai-integration[bot]
3220397ea2
feat(models): add together_ai/together/Tev1-4B-experimental (#42807)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 14:46:41 -07:00
devin-ai-integration[bot]
ccee9e77ce
feat(bedrock): add bare openai.gpt-6-sol and openai.gpt-6-luna cost map rows (#42798)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 20:43:22 +00:00
berriai-litellm-provider-info-sync[bot]
5a1e07797c
chore(prices): sync Baseten prices: 1 model (#42771)
baseten/zai-org/GLM-5.3-Fast:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-23 13:41:46 -07:00
berriai-litellm-provider-info-sync[bot]
5beac4f18d
chore(prices): sync OpenRouter prices: 2 models, 2 deprecated [20 held] (#42756)
* chore(prices): sync OpenRouter prices: 2 models, 2 deprecated [20 held]

openrouter/stealth/space-bunny-alpha: deprecation_date
openrouter/z-ai/glm-5.3-flashx: deprecation_date

Price-Sync: litellm-providers

* chore(prices): add openrouter/qwen/qwen3.8-max-prime from OpenRouter models API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): record video input for openrouter/qwen/qwen3.8-max-prime

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 13:29:39 -07:00
devin-ai-integration[bot]
1c289e5ecd
fix(prices): add baseten/zai-org/GLM-5.3-Fast pricing (#42764)
* fix(prices): add baseten/zai-org/GLM-5.3-Fast pricing with cost tracking e2e

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): assert message instead of comment on breakdown row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(baseten): drop the live e2e cost tracking test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-23 11:31:27 -07:00
devin-ai-integration[bot]
b41e6c966a
feat(models): add openrouter/stealth/space-bunny-alpha (#42759)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 16:52:34 +00:00
devin-ai-integration[bot]
21530d887b
feat(gemini): add gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts prices (#42752)
* feat(gemini): add gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts prices

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(gemini): bill tiered TTS output through output_cost_per_token tiers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 09:28:09 -07:00
devin-ai-integration[bot]
6b642f3648
feat(cost-map): add Azure Foundry pricing for gpt-6-sol and gpt-6-luna (#42747)
* feat(cost-map): add Azure Foundry pricing for gpt-6-sol and gpt-6-luna

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): give azure/eu gpt-6-sol and gpt-6-luna full model metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 15:17:36 +00:00
devin-ai-integration[bot]
75a6bca8b9
feat(bedrock): add gpt-6-sol and gpt-6-luna model pricing (#42746)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 08:13:41 -07:00
devin-ai-integration[bot]
2dccc0dc79
feat(models): add openrouter/aion-labs/aion-3.5-mini pricing (#42743)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 07:39:34 -07:00
devin-ai-integration[bot]
411fa04f86
fix(model_prices): add azure_ai gpt-image-2 and groq llama-guard-3-8b deprecation dates (#42738) 2026-09-23 07:30:17 -07:00
berriai-litellm-provider-info-sync[bot]
d525b0a8df
chore(prices): sync OpenRouter prices: 19 models, 9 new [18 held] (#42592)
* chore(prices): sync OpenRouter prices: 19 models, 9 new [18 held]

openrouter/~deepseek/deepseek-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~deepseek/deepseek-v4-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~moonshotai/kimi-latest: input_cost_per_token, output_cost_per_token
openrouter/~z-ai/glm-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/aion-labs/aion-2.0: max_input_tokens
openrouter/aion-labs/aion-3.0: max_input_tokens
openrouter/aion-labs/aion-3.0-mini: max_input_tokens
openrouter/anthropic/claude-opus-5.5:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, cache_creation_input_token_cost_above_1hr
openrouter/cohere/command-a-plus: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4.1-flash: max_tokens, max_output_tokens, off_peak_pricing
openrouter/deepseek/deepseek-v4.1-flash:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/openai/gpt-6-luna-pro:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-luna:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-sol-pro:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-sol:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-oss-20b:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3.8-omni-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

* chore(prices): sync OpenRouter prices: 1 model [9 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [12 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [9 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [12 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [10 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [13 held]

openrouter/deepseek/deepseek-v4-pro-0813: off_peak_pricing

Price-Sync: litellm-providers

* feat(prices): add openrouter/upstage/solar-mini4 from OpenRouter models API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync OpenRouter prices: 1 model [16 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* feat(prices): add openrouter/aion-labs/aion-3.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 07:23:42 -07:00
berriai-litellm-provider-info-sync[bot]
721d39f476
chore(prices): sync Vertex AI prices: 1 model (#42680)
gemini-live-2.5-flash-native-audio:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 23:22:13 -07:00
berriai-litellm-provider-info-sync[bot]
b2789d6268
chore(prices): sync AWS Bedrock prices: 1 model (#42685)
us.mistral.pixtral-large-2502-v1:0:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 22:57:51 -07:00
devin-ai-integration[bot]
7172dfc400
fix(prices): align regional Bedrock Mistral Large 24.02 keys with the AWS pricing page (#42684)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 22:57:12 -07:00
berriai-litellm-provider-info-sync[bot]
3190d93136
chore(prices): sync AWS Bedrock prices: 11 models (#42677)
* chore(prices): sync AWS Bedrock prices: 11 models

deepseek.v3-v1:0: 
global.openai.gpt-5.6-luna: 
global.openai.gpt-5.6-sol: 
global.openai.gpt-5.6-terra: 
global.openai.gpt-6-astra: 
us.deepseek.r1-v1:0: 
us.openai.gpt-5.6-luna: 
us.openai.gpt-5.6-sol: 
us.openai.gpt-5.6-terra: 
us.openai.gpt-6-astra: 
writer.palmyra-vision-7b:

Price-Sync: litellm-providers

* fix(prices): align Bedrock Mistral Large 24.02 and Small 24.02 with the AWS pricing page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 4 models

mistral.mistral-7b-instruct-v0:2: 
mistral.mixtral-8x7b-instruct-v0:1: 
openai.gpt-oss-120b-1:0: 
openai.gpt-oss-20b-1:0:

Price-Sync: litellm-providers

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 22:39:12 -07:00
berriai-litellm-provider-info-sync[bot]
d0040196fe
chore(prices): sync AWS Bedrock prices: 2 models (#42673)
global.xai.grok-4.6: 
us.xai.grok-4.6:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 21:41:13 -07:00
devin-ai-integration[bot]
fdbd8382a4
fix(pricing): align bedrock_mantle/openai.gpt-daybreak-blue-5.6-sol with its Bedrock model card (#42672)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:38:41 -07:00
berriai-litellm-provider-info-sync[bot]
80221911cd
chore(prices): sync Vertex AI prices: 2 models, 2 new [enrichment failed: Vertex AI, 224 held] (#42589)
* chore(prices): sync Vertex AI prices: 2 models, 2 new [enrichment failed: Vertex AI, 224 held]

vertex_ai/gemini-2.0-flash: input_cost_per_token, output_cost_per_token, input_cost_per_character, input_cost_per_audio_token, input_cost_per_token_batches, output_cost_per_token_batches, input_cost_per_audio_token_batches
vertex_ai/gemini-2.0-flash-lite: input_cost_per_token, output_cost_per_token, input_cost_per_character, input_cost_per_audio_token, input_cost_per_token_batches, output_cost_per_token_batches, input_cost_per_audio_token_batches

* chore(prices): sync Vertex AI prices: 52 models

vertex_ai/claude-fable-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5-1: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5-1@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-haiku-4-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-haiku-4-5@20251001: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-5@20251101: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-6: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-6@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-7: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-7@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-8: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-8@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-4-5: cache_read_input_token_cost_batches
vertex_ai/claude-sonnet-4-5@20250929: cache_read_input_token_cost_batches
vertex_ai/claude-sonnet-4-6: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-4-6@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-5: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/claude-sonnet-5@default: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/codestral-2: 
vertex_ai/codestral-2@001: 
vertex_ai/deepseek-ai/deepseek-ocr-maas: 
vertex_ai/deepseek-ai/deepseek-r1-0528-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/deepseek-ai/deepseek-v3.1-maas: cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/deepseek-ai/deepseek-v3.2-maas: cache_read_input_token_cost
vertex_ai/meta/llama-4-maverick-17b-128e-instruct-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/meta/llama-4-scout-17b-16e-instruct-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/minimaxai/minimax-m2-maas: cache_read_input_token_cost
vertex_ai/mistral-medium-3: 
vertex_ai/mistral-medium-3@001: 
vertex_ai/mistral-small-2503: 
vertex_ai/mistral-small-2503@001: 
vertex_ai/mistralai/codestral-2: 
vertex_ai/mistralai/codestral-2@001: 
vertex_ai/mistralai/mistral-medium-3: 
vertex_ai/mistralai/mistral-medium-3@001: 
vertex_ai/moonshotai/kimi-k2-thinking-maas: cache_read_input_token_cost
vertex_ai/openai/gpt-oss-120b-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/openai/gpt-oss-20b-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-235b-a22b-instruct-2507-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas: cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-next-80b-a3b-instruct-maas: 
vertex_ai/qwen/qwen3-next-80b-a3b-thinking-maas: 
vertex_ai/xai/grok-4.1-fast-non-reasoning: 
vertex_ai/xai/grok-4.1-fast-reasoning: 
vertex_ai/zai-org/glm-4.7-maas: cache_read_input_token_cost
vertex_ai/zai-org/glm-5-maas:

Price-Sync: litellm-providers

* chore(prices): add verified Vertex AI zai glm-5.2-maas entry and fix Gemini 2.0 mode

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prices): add Vertex AI claude-sonnet-4-5 1h cache write price above 200K

The Vertex AI pricing page prices Claude Sonnet 4.5's 1h Cache Write at
$6.00 up to 200K input tokens and $12.00 above, matching the value the
anthropic and bedrock entries already carry.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(prices): add vertex_ai/gemini-omni-1.1-flash-preview from the Vertex pricing page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:29:36 -07:00
devin-ai-integration[bot]
3eb7e45615
fix(pricing): drop the unpublished cached rate from the Gemini Live preview entries (#42651)
* fix(pricing): correct cached-token fields on realtime cost-map entries

azure/gpt-realtime-2 was the only member of the gpt-realtime-2 family priced
on one side of its cached-audio meter. Azure publishes that meter as
"gpt-realtime-2 Audio cd inp Gl 1M Tokens" at 0.4 per 1M and charges the
same rate for the write that populates the cache and the read that hits it,
so cache_creation_input_audio_token_cost lands at 4e-07, matching
azure/gpt-realtime-2.1, azure/gpt-realtime-2.1-mini and the openai
gpt-realtime-2 entry. No cost path reads that field yet, so this corrects
what get_model_info reports rather than what anything bills.

The gemini Live entries go the other way. Google's Vertex context-caching
page publishes separate supported-model lists for implicit and explicit
caching, and no Live or native-audio model is in either one. Its pricing
page prints N/A in both cached-input columns for every Gemini 2.5 Flash
Live API row, where plain 2.5 Flash and 2.5 Flash-Lite both carry real
cached prices, and the Vertex model card for the family marks context
caching not supported outright. Vertex never reports cachedContentTokenCount
on a Live session either, including for a byte-identical 7,021-token prefix
replayed across sessions minutes apart, which is well past the 2,048-token
minimum the same page sets for the Gemini 2 family.

So the 7.5e-08 on the two preview siblings priced something the provider does
not sell, and supports_prompt_caching on all three claimed a capability the
model does not have. The rate comes out. The flag is set to false rather than
removed, because get_model_info maps an absent key to None, and None is how
this map spells "nobody checked" across the 2,788 entries that omit it, where
false records the vendor's documented no. Both readers of the flag gate on
`is True`, so nothing bills or behaves differently either way.

Only the cached fields change on the two 09-2025 preview entries. Their
source field points at the Gemini API pricing page rather than the Vertex
one, so they describe a different surface with its own published limits, and
their context windows are left alone rather than assumed to match the Vertex
model card that drives the GA entry.

Tests cover all three halves: the family invariant that a cached audio read
implies an equal cached audio write, a cached count on a Live entry leaving
the bill at the fresh-input total instead of adding the old 7.5e-08, and
supports_prompt_caching answering false for all three entries while still
answering true for 2.5 Flash, so the false cannot be a swallowed lookup
error.

* fix(cost): correct gemini-live-2.5-flash-native-audio limits and capabilities

Google's model card for model ID gemini-live-2.5-flash-native-audio gives a
128K context window and 64K maximum output tokens, and marks structured
output, context caching and URL context as not supported. Its modality list
is text in and out, image in, audio in and out, and video in, with no
document input of any kind.

The entry advertised a 1M context window, an off-by-one 65535 output cap, and
three capability flags the vendor marks unsupported. Context caching is the
fourth and is handled in the cached-fields change alongside its two preview
siblings.

Both the bare id and vertex_ai/gemini-live-2.5-flash-native-audio resolve to
this single entry, so the test drives the corrected values through both.

* test(integration): cover live preview cached tokens billed at the fresh rate

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): cite dated sources for Live entry pins and drop restating docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:35:34 -07:00
berriai-litellm-provider-info-sync[bot]
5dfaa8d620
chore(prices): sync AWS Bedrock prices and sources from the AWS price list (#42632)
* chore(prices): sync AWS Bedrock prices: 3 models [sync failed: AWS Bedrock]

anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema

* chore(prices): sync AWS Bedrock prices: 32 models

ai21.j2-mid-v1: 
ai21.j2-ultra-v1: 
ai21.jamba-1-5-large-v1:0: 
ai21.jamba-1-5-mini-v1:0: 
ai21.jamba-instruct-v1:0: 
au.anthropic.claude-opus-4-7: 
au.anthropic.claude-opus-4-8: 
au.anthropic.claude-opus-5: 
au.anthropic.claude-sonnet-4-6: 
au.anthropic.claude-sonnet-5: 
cohere.command-light-text-v14: 
cohere.command-text-v14: input_cost_per_token
cohere.embed-english-v3: 
cohere.embed-multilingual-v3: 
cohere.embed-v4:0: 
eu.anthropic.claude-fable-5: 
eu.anthropic.claude-opus-4-7: 
eu.anthropic.claude-opus-4-8: 
eu.anthropic.claude-opus-5: 
eu.anthropic.claude-sonnet-4-6: 
eu.anthropic.claude-sonnet-5: 
jp.anthropic.claude-opus-4-7: cache_creation_input_token_cost_above_1hr
jp.anthropic.claude-opus-4-8: 
jp.anthropic.claude-opus-5: 
jp.anthropic.claude-sonnet-4-6: 
jp.anthropic.claude-sonnet-5: 
meta.llama2-13b-chat-v1: 
meta.llama2-70b-chat-v1: 
us.writer.palmyra-x4-v1:0: 
us.writer.palmyra-x5-v1:0: 
writer.palmyra-x4-v1:0: 
writer.palmyra-x5-v1:0:

Price-Sync: litellm-providers

* fix(bedrock): correct eu.anthropic.claude-opus-4-5 regional prices from the AWS price list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 30 models

anthropic.claude-haiku-4-5-20251001-v1:0: 
anthropic.claude-opus-4-5-20251101-v1:0: 
anthropic.claude-opus-4-6-v1: 
anthropic.claude-sonnet-4-20250514-v1:0: 
anthropic.claude-sonnet-4-5-20250929-v1:0: 
apac.anthropic.claude-sonnet-4-20250514-v1:0: 
au.anthropic.claude-haiku-4-5-20251001-v1:0: 
au.anthropic.claude-opus-4-6-v1: 
au.anthropic.claude-sonnet-4-5-20250929-v1:0: 
eu.anthropic.claude-haiku-4-5-20251001-v1:0: 
eu.anthropic.claude-opus-4-5-20251101-v1:0: 
eu.anthropic.claude-opus-4-6-v1: 
eu.anthropic.claude-sonnet-4-20250514-v1:0: 
eu.anthropic.claude-sonnet-4-5-20250929-v1:0: 
global.anthropic.claude-haiku-4-5-20251001-v1:0: 
global.anthropic.claude-opus-4-5-20251101-v1:0: 
global.anthropic.claude-opus-4-6-v1: 
global.anthropic.claude-sonnet-4-20250514-v1:0: 
global.anthropic.claude-sonnet-4-5-20250929-v1:0: 
jp.anthropic.claude-haiku-4-5-20251001-v1:0: 
jp.anthropic.claude-sonnet-4-5-20250929-v1:0: 
mistral.voxtral-mini-3b-2507: 
mistral.voxtral-small-24b-2507: 
us-gov.anthropic.claude-sonnet-4-5-20250929-v1:0: 
us.anthropic.claude-haiku-4-5-20251001-v1:0: 
us.anthropic.claude-opus-4-1-20250805-v1:0: 
us.anthropic.claude-opus-4-5-20251101-v1:0: 
us.anthropic.claude-opus-4-6-v1: 
us.anthropic.claude-sonnet-4-20250514-v1:0: 
us.anthropic.claude-sonnet-4-5-20250929-v1:0:

Price-Sync: litellm-providers

* fix(bedrock): keep Claude Opus 5.5 response schema support as a maintainer ruled in #42626

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): stop pinning the jp Opus 4.7 cache field absence in the ttl fallback test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 3 models

anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema

Price-Sync: litellm-providers

* fix(bedrock): keep Claude Opus 5.5 response schema support per the #42626 ruling, reverting the cron restack

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:24:28 -07:00
devin-ai-integration[bot]
5c24802fbd
fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse (#42644)
* fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse

Bedrock rejects outputConfig.textFormat on Opus 4.7 and 4.8 with
"output_config.format: Extra inputs are not permitted", and the AWS
model cards list structured outputs as not supported for both, so
their cost-map entries no longer claim supports_native_structured_output
and json_schema requests fall back to the json_tool_call tool.

Fixes #27846

* test(bedrock): assert Opus 4.7 and 4.8 inline the schema on Invoke, move the native case to Sonnet 4.6

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 20:09:21 -07:00
berriai-litellm-provider-info-sync[bot]
96a2015c83
chore(prices): sync OpenAI prices: 1 model (#42648)
sora-2-pro-high-res:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:38 -07:00
berriai-litellm-provider-info-sync[bot]
0135387abf
chore(prices): sync Google Gemini prices: 1 model (#42642)
gemini/gemini-robotics-er-2-streaming-preview: input_cost_per_token, output_cost_per_token

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:34 -07:00
yuneng-jiang
b395bfefdd
fix: repair seven regressions caught by CircleCI on main (#42640)
* fix: repair seven regressions caught by CircleCI on main

- vertex_ai: stop treating fine-tuned endpoint ids (numeric or
  vertex_ai/gemini/<id>) and gemma models as Gemini 3+, which injected
  temperature=1.0 and Gemini 3 thinking config into their requests (#42465)
- cost: price Azure DALL-E 3 from its azure/<quality>/<size>/dall-e-3 rows;
  it only worked through the OpenAI rows that #42435 removed
- bedrock: stream bedrock/invoke/moonshot through an OpenAI-shaped chunk
  decoder; the generic decoder dropped every chunk, which the
  supports_response_schema flag from #42338 un-skipped in CI
- proxy: keep the public model_group on pre-routing rejections so the Usage
  page groups them under the model name, not the deployment (#41077)
- cost map: mirror the base rows' capability flags onto Bedrock regional and
  cross-region copies (#42254 and later syncs)
- whitelist the new regional Bedrock rows from #42543 and #42588 for the
  converse routing check, following the existing regional-row convention

* fix(model-prices): mirror capability flags onto ap-southeast-3 bedrock rows

* refactor(bedrock): tighten types on the moonshot stream decoder and its tests
2026-09-23 02:26:36 +00:00
Emerson Gomes
30004f5f05
fix(bedrock): drop unsupported sampling params on converse reasoning models (#39834)
* fix(bedrock): drop unsupported sampling params on converse reasoning models

* test(bedrock): resolve duplicate import after rebase
2026-09-22 19:07:14 -07:00
devin-ai-integration[bot]
7688f56256
fix(pricing): drop the duplicate cache_read_input_token_cost_batches key from 23 entries (#42623)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 17:28:48 -07:00
berriai-litellm-provider-info-sync[bot]
95d4613bd3
chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps] (#42591)
* chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps]

xai/grok-code-fast: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1-0825: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema

* fix(prices): add context limits and reasoning flag to xai grok-code-fast aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prices): keep the xai sync diff limited to the alias fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:14:48 -07:00
devin-ai-integration[bot]
d31e8aac6d
feat(cost-map): add Claude Opus 5.5 for Vertex AI and Azure AI (#42599)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:48:07 +00:00
devin-ai-integration[bot]
9082f8e5d0
feat(bedrock): add Claude Opus 5.5 pricing and capabilities (#42588)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:37:12 +00:00
devin-ai-integration[bot]
075536eca1
chore(cost-map): remove models past their deprecation date (#42435)
* chore(cost-map): remove models past their deprecation date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop the empty parametrize left behind by the gemini web search removal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): drop merge base block left by conflict resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop gemini image cost tests pinned on removed model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:19:26 +00:00
berriai-litellm-provider-info-sync[bot]
cd69342eef
chore(prices): sync Vertex AI prices: 20 models (#42577)
deep-research-pro-preview-12-2025: cache_read_input_token_cost_batches
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost_batches
gemini-3-pro-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3-pro-image: cache_read_input_token_cost_batches
gemini-3.1-flash-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-image: cache_read_input_token_cost_batches
gemini-3.1-flash-lite: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-lite: cache_read_input_token_cost_batches
gemini-3.1-flash-lite-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-lite-image: cache_read_input_token_cost_batches
gemini-3.5-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.5-flash: cache_read_input_token_cost_batches
gemini-3.5-flash-lite: cache_read_input_token_cost_batches
vertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_batches
gemini-3.6-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.6-flash: cache_read_input_token_cost_batches
gemini-3.7-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.7-flash: cache_read_input_token_cost_batches
gemini-3.8-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.8-flash: cache_read_input_token_cost_batches

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 14:11:31 -07:00
devin-ai-integration[bot]
e30f9f578c
fix(model_prices): registry audit 2026-09-22, absorb open pricing PRs (#42543)
* fix(model_prices): registry audit 2026-09-22, absorb open pricing PRs

Rolls the open registry-only PRs into one PR after re-verifying every value against the official provider source: OpenAI, Azure, Vertex AI and Gemini batch cache-read prices, Baseten model metadata from the authenticated inference API, Bedrock eu-west-2 Nemotron Super 3 pricing from the AWS offer file, and OpenRouter prices refreshed from the live OpenRouter models API

Co-authored-by: sinksilk <785976238@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add groq/llama-guard-3-8b from the Groq model page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): refresh openrouter deepseek aliases from live api and drop stale off-peak windows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): resolve baseten merge conflicts against main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: sinksilk <785976238@qq.com>
2026-09-22 14:05:33 -07:00
berriai-litellm-provider-info-sync[bot]
b73696a15d
chore(prices): sync OpenAI prices: 3 models (#42557)
gpt-5-2025-08-07: cache_read_input_token_cost_batches
gpt-5-mini-2025-08-07: cache_read_input_token_cost_batches
gpt-5-nano-2025-08-07: cache_read_input_token_cost_batches

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:52:30 -07:00
berriai-litellm-provider-info-sync[bot]
e54f228ed8
chore(prices): sync Baseten prices: 4 models, 4 new [enrichment failed: Baseten, 3 held] (#42578)
baseten/deepseek-ai/DeepSeek-V4-Flash-0731: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/deepseek-ai/DeepSeek-V4-Pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/deepseek-ai/DeepSeek-V4-Pro-0813: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-5.2-Fast: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:46:15 -07:00
berriai-litellm-provider-info-sync[bot]
2727359a9e
chore(prices): sync Baseten prices: 12 models, 9 new [enrichment failed: Baseten, 15 held] (#42558)
baseten/deepseek-ai/DeepSeek-V4.1-Flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K2.6: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K2.7-Code: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K3: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/openai/gpt-oss-120b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, cache_read_input_token_cost
baseten/thinkingmachines/inkling: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/thinkingmachines/inkling-small: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-4.7: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, cache_read_input_token_cost
baseten/zai-org/GLM-5.2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-5.3: supports_reasoning
baseten/zai-org/GLM-5.3-Flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:24:49 -07:00
devin-ai-integration[bot]
9d299d016a
chore(pricing): remove retired models flagged by the provider sync (#42521)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 19:03:40 +00:00
devin-ai-integration[bot]
e0c2cbff21
feat(openai): add GPT-6 Sol and GPT-6 Luna (#42515)
Add gpt-6-sol and gpt-6-luna to the model cost map with pricing from the OpenAI pricing page and reasoning effort levels none through max. Extend the long-context priority pricing and reasoning effort capability tests to cover both models

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 11:34:01 -07:00
berriai-litellm-provider-info-sync[bot]
8332f0848a
chore(prices): sync Google Gemini prices: 13 models (#42500)
gemini/gemini-2.5-flash: cache_read_input_token_cost_batches
gemini/gemini-2.5-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-2.5-pro: cache_read_input_token_cost_batches
gemini/gemini-3-flash-preview: cache_read_input_token_cost_batches
gemini/gemini-3.1-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-3.1-pro-preview: cache_read_input_token_cost_batches
gemini/gemini-3.1-pro-preview-customtools: cache_read_input_token_cost_batches
gemini/gemini-3.5-flash: cache_read_input_token_cost_batches
gemini/gemini-3.5-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-3.6-flash: cache_read_input_token_cost_batches
gemini/gemini-3.7-flash: cache_read_input_token_cost_batches
gemini/gemini-3.8-flash: cache_read_input_token_cost_batches
gemini/gemini-robotics-er-2-preview: cache_read_input_token_cost_batches

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:29:38 -07:00
berriai-litellm-provider-info-sync[bot]
8a9305fa27
chore(prices): sync OpenAI prices: 25 models [enrichment failed: OpenAI, 32 held] (#42501)
chatgpt-image-latest: cache_read_input_token_cost_batches
ft:gpt-4.1-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4.1-mini-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4.1-nano-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4o-2024-08-06: cache_read_input_token_cost_batches
ft:gpt-4o-mini-2024-07-18: cache_read_input_token_cost_batches
ft:o4-mini-2025-04-16: cache_read_input_token_cost_batches
gpt-5: cache_read_input_token_cost_batches
gpt-5-mini: cache_read_input_token_cost_batches
gpt-5-nano: cache_read_input_token_cost_batches
gpt-5.1: cache_read_input_token_cost_batches
gpt-5.1-2025-11-13: cache_read_input_token_cost_batches
gpt-5.2: cache_read_input_token_cost_batches
gpt-5.2-2025-12-11: cache_read_input_token_cost_batches
gpt-5.4: cache_read_input_token_cost_batches
gpt-5.4-2026-03-05: cache_read_input_token_cost_batches
gpt-5.4-mini: cache_read_input_token_cost_batches
gpt-5.4-mini-2026-03-17: cache_read_input_token_cost_batches
gpt-5.4-nano: cache_read_input_token_cost_batches
gpt-5.4-nano-2026-03-17: cache_read_input_token_cost_batches
gpt-image-1: cache_read_input_token_cost_batches
gpt-image-1-mini: cache_read_input_token_cost_batches
gpt-image-1.5: cache_read_input_token_cost_batches
gpt-image-1.5-2025-12-16: cache_read_input_token_cost_batches
gpt-image-2: cache_read_input_token_cost_batches

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:21:11 -07:00
berriai-litellm-provider-info-sync[bot]
c3389d7901
chore(prices): sync Azure prices: 34 models (#42502)
azure/eu/gpt-5: cache_read_input_token_cost_batches
azure/eu/gpt-5-mini: cache_read_input_token_cost_batches
azure/eu/gpt-5-nano: cache_read_input_token_cost_batches
azure/eu/gpt-5.1: cache_read_input_token_cost_batches
azure/eu/gpt-5.2: cache_read_input_token_cost_batches
azure/eu/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/eu/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/eu/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/eu/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/eu/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/eu/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5: cache_read_input_token_cost_batches
azure/gpt-5-mini: cache_read_input_token_cost_batches
azure/gpt-5-nano: cache_read_input_token_cost_batches
azure/gpt-5.1: cache_read_input_token_cost_batches
azure/global/gpt-5.1: cache_read_input_token_cost_batches
azure/gpt-5.2: cache_read_input_token_cost_batches
azure/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5: cache_read_input_token_cost_batches
azure/us/gpt-5-mini: cache_read_input_token_cost_batches
azure/us/gpt-5-nano: cache_read_input_token_cost_batches
azure/us/gpt-5.1: cache_read_input_token_cost_batches
azure/us/gpt-5.2: cache_read_input_token_cost_batches
azure/us/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/us/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/us/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/us/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:21:06 -07:00
Mateo Wang
deba473821
fix(cost): bill batch prompts above 272K at OpenAI's long-context batch tier (#39861)
* fix(cost): bill batch prompts above 272K at OpenAI's long-context batch tier

* fix(cost): mirror batch long-context keys on custom pricing params

Register the two *_above_272k_tokens_batches keys on CustomPricingLiteLLMParams so a per-deployment override stays out of the shared backend key, add them to the inline model-info schema and alias-count tests, and build LiteLLM_Params and GenericLiteLLMParams through model_validate at the two dict-splat call sites so basedpyright's reportArgumentType budget ratchets down instead of blocking the new fields.

* fix(cost): add the gpt-5.5-pro batch long-context tier and ignore malformed batch tier keys

* fix(cost): bill cached batch tokens at OpenAI's cached batch rate

Adds cache_read_input_token_cost_batches and
cache_read_input_token_cost_above_272k_tokens_batches for the tiered
OpenAI entries at half the standard cached rate, bills cached batch
tokens at that rate per output line, and parses string-valued batch
rates in deployment-level model_info.

* fix(cost): bill batch cache writes at the batch cache-write rate and carry published batch rates for one-sided deployments

OpenAI's Batch table prices cache writes for gpt-6-astra, gpt-5.6, gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna at half the standard cache-write rate, so the cost map gains cache_creation_input_token_cost_batches and its above_272k tier for those entries and batch cost pulls written tokens out of the input bucket at that rate; models without the key keep billing writes at the batch input rate.

A deployment declaring only one side of its batch pricing now carries every published batch rate of the other side (tier, cached, cache write), its own keys win, and a lone tier, cached or cache-write batch key counts as declared pricing instead of being ignored.

* fix(cost): select the batch long-context tier from any batch tier key

A deployment that declares its own flat standard input rate keeps every
published batch rate of the output direction, including the 272K output
tier, but the tier was only ever selected when an input tier key was also
present. Detect the crossed tier from any of the four batch tier keys so
the carried output, cache-read, and cache-write tiers bill at their tier
rate above 272K tokens.

* chore(proxy): keep the OpenAPI snapshot as CI generates it

* fix(cost): pick each batch price component's tier from its own keys

The batch rate picker crossed one threshold for every component, so a
deployment declaring only an output tier also moved its input, cached, and
cache-write rates to that cutoff. Each component now crosses its own
*_above_<N>k_tokens_batches keys and falls back to its flat key.

The JSON schema is regenerated with the generator as it is on main:
cost-map-guard renders the PR's cost map with the base branch's generator,
so the descriptions for the new batch cache keys move to a follow-up.

* chore(proxy): restore the lazy OpenAPI snapshot to what CI's Python 3.12 generates

The merge commit carried a snapshot regenerated on a Python 3.14 venv, which dedents
docstrings at compile time, so one description line differed from the file CI regenerates
on 3.12 and the schema.d.ts sync check went red. The snapshot is byte-identical to main again
2026-09-22 10:22:41 -07:00
devin-ai-integration[bot]
c835a1a982
feat(anthropic): add Claude Opus 5.5 (#42489)
Adds the anthropic cost map entry for claude-opus-5-5 at $4/$20 per MTok
with $5 per MTok 5m cache writes, $8 per MTok 1h cache writes, $0.20 per
MTok cache reads (0.05x base), and fast mode at 2x. The entry sets
thinking_always_on (Opus 5.5 cannot turn thinking off) and
supports_forced_tool_use false (tool_choice required/named 400s, same as
Fable 5.1), mirrors that model by omitting thinking cache preservation,
and registers the model in the setup wizard provider list

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 09:52:35 -07:00
berriai-litellm-provider-info-sync[bot]
226fa7845a
chore(prices): sync OpenRouter prices: 1 model [enrichment failed: OpenRouter, 5 held] (#42485)
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 09:09:52 -07:00
berriai-litellm-provider-info-sync[bot]
4082523596
chore(prices): sync OpenRouter prices: 14 models, 1 deprecated [enrichment failed: OpenRouter, 5 held] (#42438)
openrouter/~z-ai/glm-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/aion-labs/aion-2.0: max_input_tokens
openrouter/aion-labs/aion-3.0: max_input_tokens
openrouter/aion-labs/aion-3.0-mini: max_input_tokens
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/dots-studio/dots-3-note-preview🆓 deprecation_date
openrouter/moonshotai/kimi-k2.7-code: output_cost_per_token
openrouter/moonshotai/kimi-k3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/openai/gpt-oss-20b: max_tokens, max_output_tokens, supports_prompt_caching, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3-next-80b-a3b-thinking: max_tokens, max_output_tokens
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3-flash:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 08:42:25 -07:00
tin-berri
5a8c4f48e4
feat(router): native compact-to-fit across conversation APIs (#42074)
* feat(router): native compact-to-fit across conversation APIs

* fix(router): preserve compaction admission and shared client boundaries

* fix(router): honor compaction fit fallbacks and router-scoped access

* fix(router): charge compaction usage to caller token limits

* test(http): keep FastAPI inside proxy tests

* fix(router): check compactor capacity before skipping escalation
2026-09-21 22:52:29 -07:00
berriai-litellm-provider-info-sync[bot]
798d45970b
chore(prices): sync OpenRouter prices: 2 models (#42418)
openrouter/nex-agi/nex-n2.5-mini: supports_response_schema
openrouter/nex-agi/nex-n2.5-pro: supports_tool_choice, supports_response_schema, supports_function_calling

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 22:29:26 -07:00
berriai-litellm-provider-info-sync[bot]
0132f34356
chore(prices): sync OpenRouter prices: 2 models, 2 new (#42407)
openrouter/nex-agi/nex-n2.5-mini: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/nex-agi/nex-n2.5-pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:40:25 -07:00
berriai-litellm-provider-info-sync[bot]
dfd8ffc545
chore(prices): sync OpenRouter prices: 2 models, 2 deprecated (#42381)
openrouter/nex-agi/nex-n2.5-mini🆓 deprecation_date
openrouter/nex-agi/nex-n2.5-pro🆓 deprecation_date

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:11:22 -07:00