berriai-litellm-provider-info-sync[bot]
5dfaa8d620
chore(prices): sync AWS Bedrock prices and sources from the AWS price list ( #42632 )
...
* chore(prices): sync AWS Bedrock prices: 3 models [sync failed: AWS Bedrock]
anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
* chore(prices): sync AWS Bedrock prices: 32 models
ai21.j2-mid-v1:
ai21.j2-ultra-v1:
ai21.jamba-1-5-large-v1:0:
ai21.jamba-1-5-mini-v1:0:
ai21.jamba-instruct-v1:0:
au.anthropic.claude-opus-4-7:
au.anthropic.claude-opus-4-8:
au.anthropic.claude-opus-5:
au.anthropic.claude-sonnet-4-6:
au.anthropic.claude-sonnet-5:
cohere.command-light-text-v14:
cohere.command-text-v14: input_cost_per_token
cohere.embed-english-v3:
cohere.embed-multilingual-v3:
cohere.embed-v4:0:
eu.anthropic.claude-fable-5:
eu.anthropic.claude-opus-4-7:
eu.anthropic.claude-opus-4-8:
eu.anthropic.claude-opus-5:
eu.anthropic.claude-sonnet-4-6:
eu.anthropic.claude-sonnet-5:
jp.anthropic.claude-opus-4-7: cache_creation_input_token_cost_above_1hr
jp.anthropic.claude-opus-4-8:
jp.anthropic.claude-opus-5:
jp.anthropic.claude-sonnet-4-6:
jp.anthropic.claude-sonnet-5:
meta.llama2-13b-chat-v1:
meta.llama2-70b-chat-v1:
us.writer.palmyra-x4-v1:0:
us.writer.palmyra-x5-v1:0:
writer.palmyra-x4-v1:0:
writer.palmyra-x5-v1:0:
Price-Sync: litellm-providers
* fix(bedrock): correct eu.anthropic.claude-opus-4-5 regional prices from the AWS price list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(prices): sync AWS Bedrock prices: 30 models
anthropic.claude-haiku-4-5-20251001-v1:0:
anthropic.claude-opus-4-5-20251101-v1:0:
anthropic.claude-opus-4-6-v1:
anthropic.claude-sonnet-4-20250514-v1:0:
anthropic.claude-sonnet-4-5-20250929-v1:0:
apac.anthropic.claude-sonnet-4-20250514-v1:0:
au.anthropic.claude-haiku-4-5-20251001-v1:0:
au.anthropic.claude-opus-4-6-v1:
au.anthropic.claude-sonnet-4-5-20250929-v1:0:
eu.anthropic.claude-haiku-4-5-20251001-v1:0:
eu.anthropic.claude-opus-4-5-20251101-v1:0:
eu.anthropic.claude-opus-4-6-v1:
eu.anthropic.claude-sonnet-4-20250514-v1:0:
eu.anthropic.claude-sonnet-4-5-20250929-v1:0:
global.anthropic.claude-haiku-4-5-20251001-v1:0:
global.anthropic.claude-opus-4-5-20251101-v1:0:
global.anthropic.claude-opus-4-6-v1:
global.anthropic.claude-sonnet-4-20250514-v1:0:
global.anthropic.claude-sonnet-4-5-20250929-v1:0:
jp.anthropic.claude-haiku-4-5-20251001-v1:0:
jp.anthropic.claude-sonnet-4-5-20250929-v1:0:
mistral.voxtral-mini-3b-2507:
mistral.voxtral-small-24b-2507:
us-gov.anthropic.claude-sonnet-4-5-20250929-v1:0:
us.anthropic.claude-haiku-4-5-20251001-v1:0:
us.anthropic.claude-opus-4-1-20250805-v1:0:
us.anthropic.claude-opus-4-5-20251101-v1:0:
us.anthropic.claude-opus-4-6-v1:
us.anthropic.claude-sonnet-4-20250514-v1:0:
us.anthropic.claude-sonnet-4-5-20250929-v1:0:
Price-Sync: litellm-providers
* fix(bedrock): keep Claude Opus 5.5 response schema support as a maintainer ruled in #42626
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): stop pinning the jp Opus 4.7 cache field absence in the ttl fallback test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(prices): sync AWS Bedrock prices: 3 models
anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
Price-Sync: litellm-providers
* fix(bedrock): keep Claude Opus 5.5 response schema support per the #42626 ruling, reverting the cron restack
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:24:28 -07:00
devin-ai-integration[bot]
5c24802fbd
fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse ( #42644 )
...
* fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse
Bedrock rejects outputConfig.textFormat on Opus 4.7 and 4.8 with
"output_config.format: Extra inputs are not permitted", and the AWS
model cards list structured outputs as not supported for both, so
their cost-map entries no longer claim supports_native_structured_output
and json_schema requests fall back to the json_tool_call tool.
Fixes #27846
* test(bedrock): assert Opus 4.7 and 4.8 inline the schema on Invoke, move the native case to Sonnet 4.6
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 20:09:21 -07:00
berriai-litellm-provider-info-sync[bot]
96a2015c83
chore(prices): sync OpenAI prices: 1 model ( #42648 )
...
sora-2-pro-high-res:
Price-Sync: litellm-providers
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:38 -07:00
berriai-litellm-provider-info-sync[bot]
0135387abf
chore(prices): sync Google Gemini prices: 1 model ( #42642 )
...
gemini/gemini-robotics-er-2-streaming-preview: input_cost_per_token, output_cost_per_token
Price-Sync: litellm-providers
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:34 -07:00
yuneng-jiang
b395bfefdd
fix: repair seven regressions caught by CircleCI on main ( #42640 )
...
* fix: repair seven regressions caught by CircleCI on main
- vertex_ai: stop treating fine-tuned endpoint ids (numeric or
vertex_ai/gemini/<id>) and gemma models as Gemini 3+, which injected
temperature=1.0 and Gemini 3 thinking config into their requests (#42465 )
- cost: price Azure DALL-E 3 from its azure/<quality>/<size>/dall-e-3 rows;
it only worked through the OpenAI rows that #42435 removed
- bedrock: stream bedrock/invoke/moonshot through an OpenAI-shaped chunk
decoder; the generic decoder dropped every chunk, which the
supports_response_schema flag from #42338 un-skipped in CI
- proxy: keep the public model_group on pre-routing rejections so the Usage
page groups them under the model name, not the deployment (#41077 )
- cost map: mirror the base rows' capability flags onto Bedrock regional and
cross-region copies (#42254 and later syncs)
- whitelist the new regional Bedrock rows from #42543 and #42588 for the
converse routing check, following the existing regional-row convention
* fix(model-prices): mirror capability flags onto ap-southeast-3 bedrock rows
* refactor(bedrock): tighten types on the moonshot stream decoder and its tests
2026-09-23 02:26:36 +00:00
Emerson Gomes
30004f5f05
fix(bedrock): drop unsupported sampling params on converse reasoning models ( #39834 )
...
* fix(bedrock): drop unsupported sampling params on converse reasoning models
* test(bedrock): resolve duplicate import after rebase
2026-09-22 19:07:14 -07:00
devin-ai-integration[bot]
7688f56256
fix(pricing): drop the duplicate cache_read_input_token_cost_batches key from 23 entries ( #42623 )
...
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 17:28:48 -07:00
berriai-litellm-provider-info-sync[bot]
95d4613bd3
chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps] ( #42591 )
...
* chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps]
xai/grok-code-fast: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1-0825: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
* fix(prices): add context limits and reasoning flag to xai grok-code-fast aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(prices): keep the xai sync diff limited to the alias fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:14:48 -07:00
devin-ai-integration[bot]
d31e8aac6d
feat(cost-map): add Claude Opus 5.5 for Vertex AI and Azure AI ( #42599 )
...
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:48:07 +00:00
devin-ai-integration[bot]
9082f8e5d0
feat(bedrock): add Claude Opus 5.5 pricing and capabilities ( #42588 )
...
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:37:12 +00:00
devin-ai-integration[bot]
075536eca1
chore(cost-map): remove models past their deprecation date ( #42435 )
...
* chore(cost-map): remove models past their deprecation date
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost-calc): drop the empty parametrize left behind by the gemini web search removal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): drop merge base block left by conflict resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost-calc): drop gemini image cost tests pinned on removed model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:19:26 +00:00
berriai-litellm-provider-info-sync[bot]
cd69342eef
chore(prices): sync Vertex AI prices: 20 models ( #42577 )
...
deep-research-pro-preview-12-2025: cache_read_input_token_cost_batches
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost_batches
gemini-3-pro-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3-pro-image: cache_read_input_token_cost_batches
gemini-3.1-flash-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-image: cache_read_input_token_cost_batches
gemini-3.1-flash-lite: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-lite: cache_read_input_token_cost_batches
gemini-3.1-flash-lite-image: cache_read_input_token_cost_batches
vertex_ai/gemini-3.1-flash-lite-image: cache_read_input_token_cost_batches
gemini-3.5-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.5-flash: cache_read_input_token_cost_batches
gemini-3.5-flash-lite: cache_read_input_token_cost_batches
vertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_batches
gemini-3.6-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.6-flash: cache_read_input_token_cost_batches
gemini-3.7-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.7-flash: cache_read_input_token_cost_batches
gemini-3.8-flash: cache_read_input_token_cost_batches
vertex_ai/gemini-3.8-flash: cache_read_input_token_cost_batches
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 14:11:31 -07:00
devin-ai-integration[bot]
e30f9f578c
fix(model_prices): registry audit 2026-09-22, absorb open pricing PRs ( #42543 )
...
* fix(model_prices): registry audit 2026-09-22, absorb open pricing PRs
Rolls the open registry-only PRs into one PR after re-verifying every value against the official provider source: OpenAI, Azure, Vertex AI and Gemini batch cache-read prices, Baseten model metadata from the authenticated inference API, Bedrock eu-west-2 Nemotron Super 3 pricing from the AWS offer file, and OpenRouter prices refreshed from the live OpenRouter models API
Co-authored-by: sinksilk <785976238@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): add groq/llama-guard-3-8b from the Groq model page
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): refresh openrouter deepseek aliases from live api and drop stale off-peak windows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): resolve baseten merge conflicts against main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: sinksilk <785976238@qq.com>
2026-09-22 14:05:33 -07:00
berriai-litellm-provider-info-sync[bot]
b73696a15d
chore(prices): sync OpenAI prices: 3 models ( #42557 )
...
gpt-5-2025-08-07: cache_read_input_token_cost_batches
gpt-5-mini-2025-08-07: cache_read_input_token_cost_batches
gpt-5-nano-2025-08-07: cache_read_input_token_cost_batches
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:52:30 -07:00
berriai-litellm-provider-info-sync[bot]
e54f228ed8
chore(prices): sync Baseten prices: 4 models, 4 new [enrichment failed: Baseten, 3 held] ( #42578 )
...
baseten/deepseek-ai/DeepSeek-V4-Flash-0731: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/deepseek-ai/DeepSeek-V4-Pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/deepseek-ai/DeepSeek-V4-Pro-0813: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-5.2-Fast: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:46:15 -07:00
berriai-litellm-provider-info-sync[bot]
2727359a9e
chore(prices): sync Baseten prices: 12 models, 9 new [enrichment failed: Baseten, 15 held] ( #42558 )
...
baseten/deepseek-ai/DeepSeek-V4.1-Flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K2.6: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K2.7-Code: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/moonshotai/Kimi-K3: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/openai/gpt-oss-120b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, cache_read_input_token_cost
baseten/thinkingmachines/inkling: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/thinkingmachines/inkling-small: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-4.7: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, cache_read_input_token_cost
baseten/zai-org/GLM-5.2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
baseten/zai-org/GLM-5.3: supports_reasoning
baseten/zai-org/GLM-5.3-Flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_reasoning, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 13:24:49 -07:00
devin-ai-integration[bot]
9d299d016a
chore(pricing): remove retired models flagged by the provider sync ( #42521 )
...
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 19:03:40 +00:00
devin-ai-integration[bot]
e0c2cbff21
feat(openai): add GPT-6 Sol and GPT-6 Luna ( #42515 )
...
Add gpt-6-sol and gpt-6-luna to the model cost map with pricing from the OpenAI pricing page and reasoning effort levels none through max. Extend the long-context priority pricing and reasoning effort capability tests to cover both models
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 11:34:01 -07:00
berriai-litellm-provider-info-sync[bot]
8332f0848a
chore(prices): sync Google Gemini prices: 13 models ( #42500 )
...
gemini/gemini-2.5-flash: cache_read_input_token_cost_batches
gemini/gemini-2.5-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-2.5-pro: cache_read_input_token_cost_batches
gemini/gemini-3-flash-preview: cache_read_input_token_cost_batches
gemini/gemini-3.1-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-3.1-pro-preview: cache_read_input_token_cost_batches
gemini/gemini-3.1-pro-preview-customtools: cache_read_input_token_cost_batches
gemini/gemini-3.5-flash: cache_read_input_token_cost_batches
gemini/gemini-3.5-flash-lite: cache_read_input_token_cost_batches
gemini/gemini-3.6-flash: cache_read_input_token_cost_batches
gemini/gemini-3.7-flash: cache_read_input_token_cost_batches
gemini/gemini-3.8-flash: cache_read_input_token_cost_batches
gemini/gemini-robotics-er-2-preview: cache_read_input_token_cost_batches
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:29:38 -07:00
berriai-litellm-provider-info-sync[bot]
8a9305fa27
chore(prices): sync OpenAI prices: 25 models [enrichment failed: OpenAI, 32 held] ( #42501 )
...
chatgpt-image-latest: cache_read_input_token_cost_batches
ft:gpt-4.1-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4.1-mini-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4.1-nano-2025-04-14: cache_read_input_token_cost_batches
ft:gpt-4o-2024-08-06: cache_read_input_token_cost_batches
ft:gpt-4o-mini-2024-07-18: cache_read_input_token_cost_batches
ft:o4-mini-2025-04-16: cache_read_input_token_cost_batches
gpt-5: cache_read_input_token_cost_batches
gpt-5-mini: cache_read_input_token_cost_batches
gpt-5-nano: cache_read_input_token_cost_batches
gpt-5.1: cache_read_input_token_cost_batches
gpt-5.1-2025-11-13: cache_read_input_token_cost_batches
gpt-5.2: cache_read_input_token_cost_batches
gpt-5.2-2025-12-11: cache_read_input_token_cost_batches
gpt-5.4: cache_read_input_token_cost_batches
gpt-5.4-2026-03-05: cache_read_input_token_cost_batches
gpt-5.4-mini: cache_read_input_token_cost_batches
gpt-5.4-mini-2026-03-17: cache_read_input_token_cost_batches
gpt-5.4-nano: cache_read_input_token_cost_batches
gpt-5.4-nano-2026-03-17: cache_read_input_token_cost_batches
gpt-image-1: cache_read_input_token_cost_batches
gpt-image-1-mini: cache_read_input_token_cost_batches
gpt-image-1.5: cache_read_input_token_cost_batches
gpt-image-1.5-2025-12-16: cache_read_input_token_cost_batches
gpt-image-2: cache_read_input_token_cost_batches
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:21:11 -07:00
berriai-litellm-provider-info-sync[bot]
c3389d7901
chore(prices): sync Azure prices: 34 models ( #42502 )
...
azure/eu/gpt-5: cache_read_input_token_cost_batches
azure/eu/gpt-5-mini: cache_read_input_token_cost_batches
azure/eu/gpt-5-nano: cache_read_input_token_cost_batches
azure/eu/gpt-5.1: cache_read_input_token_cost_batches
azure/eu/gpt-5.2: cache_read_input_token_cost_batches
azure/eu/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/eu/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/eu/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/eu/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/eu/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/eu/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5: cache_read_input_token_cost_batches
azure/gpt-5-mini: cache_read_input_token_cost_batches
azure/gpt-5-nano: cache_read_input_token_cost_batches
azure/gpt-5.1: cache_read_input_token_cost_batches
azure/global/gpt-5.1: cache_read_input_token_cost_batches
azure/gpt-5.2: cache_read_input_token_cost_batches
azure/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5: cache_read_input_token_cost_batches
azure/us/gpt-5-mini: cache_read_input_token_cost_batches
azure/us/gpt-5-nano: cache_read_input_token_cost_batches
azure/us/gpt-5.1: cache_read_input_token_cost_batches
azure/us/gpt-5.2: cache_read_input_token_cost_batches
azure/us/gpt-5.4: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5.4-mini: cache_read_input_token_cost_batches
azure/us/gpt-5.4-nano: cache_read_input_token_cost_batches
azure/us/gpt-5.4-pro: input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches
azure/us/gpt-5.5: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
azure/us/gpt-5.5-2026-04-23: cache_read_input_token_cost_batches, input_cost_per_token_above_272k_tokens_batches, output_cost_per_token_above_272k_tokens_batches, cache_read_input_token_cost_above_272k_tokens_batches
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 11:21:06 -07:00
Mateo Wang
deba473821
fix(cost): bill batch prompts above 272K at OpenAI's long-context batch tier ( #39861 )
...
* fix(cost): bill batch prompts above 272K at OpenAI's long-context batch tier
* fix(cost): mirror batch long-context keys on custom pricing params
Register the two *_above_272k_tokens_batches keys on CustomPricingLiteLLMParams so a per-deployment override stays out of the shared backend key, add them to the inline model-info schema and alias-count tests, and build LiteLLM_Params and GenericLiteLLMParams through model_validate at the two dict-splat call sites so basedpyright's reportArgumentType budget ratchets down instead of blocking the new fields.
* fix(cost): add the gpt-5.5-pro batch long-context tier and ignore malformed batch tier keys
* fix(cost): bill cached batch tokens at OpenAI's cached batch rate
Adds cache_read_input_token_cost_batches and
cache_read_input_token_cost_above_272k_tokens_batches for the tiered
OpenAI entries at half the standard cached rate, bills cached batch
tokens at that rate per output line, and parses string-valued batch
rates in deployment-level model_info.
* fix(cost): bill batch cache writes at the batch cache-write rate and carry published batch rates for one-sided deployments
OpenAI's Batch table prices cache writes for gpt-6-astra, gpt-5.6, gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna at half the standard cache-write rate, so the cost map gains cache_creation_input_token_cost_batches and its above_272k tier for those entries and batch cost pulls written tokens out of the input bucket at that rate; models without the key keep billing writes at the batch input rate.
A deployment declaring only one side of its batch pricing now carries every published batch rate of the other side (tier, cached, cache write), its own keys win, and a lone tier, cached or cache-write batch key counts as declared pricing instead of being ignored.
* fix(cost): select the batch long-context tier from any batch tier key
A deployment that declares its own flat standard input rate keeps every
published batch rate of the output direction, including the 272K output
tier, but the tier was only ever selected when an input tier key was also
present. Detect the crossed tier from any of the four batch tier keys so
the carried output, cache-read, and cache-write tiers bill at their tier
rate above 272K tokens.
* chore(proxy): keep the OpenAPI snapshot as CI generates it
* fix(cost): pick each batch price component's tier from its own keys
The batch rate picker crossed one threshold for every component, so a
deployment declaring only an output tier also moved its input, cached, and
cache-write rates to that cutoff. Each component now crosses its own
*_above_<N>k_tokens_batches keys and falls back to its flat key.
The JSON schema is regenerated with the generator as it is on main:
cost-map-guard renders the PR's cost map with the base branch's generator,
so the descriptions for the new batch cache keys move to a follow-up.
* chore(proxy): restore the lazy OpenAPI snapshot to what CI's Python 3.12 generates
The merge commit carried a snapshot regenerated on a Python 3.14 venv, which dedents
docstrings at compile time, so one description line differed from the file CI regenerates
on 3.12 and the schema.d.ts sync check went red. The snapshot is byte-identical to main again
2026-09-22 10:22:41 -07:00
devin-ai-integration[bot]
c835a1a982
feat(anthropic): add Claude Opus 5.5 ( #42489 )
...
Adds the anthropic cost map entry for claude-opus-5-5 at $4/$20 per MTok
with $5 per MTok 5m cache writes, $8 per MTok 1h cache writes, $0.20 per
MTok cache reads (0.05x base), and fast mode at 2x. The entry sets
thinking_always_on (Opus 5.5 cannot turn thinking off) and
supports_forced_tool_use false (tool_choice required/named 400s, same as
Fable 5.1), mirrors that model by omitting thinking cache preservation,
and registers the model in the setup wizard provider list
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 09:52:35 -07:00
berriai-litellm-provider-info-sync[bot]
226fa7845a
chore(prices): sync OpenRouter prices: 1 model [enrichment failed: OpenRouter, 5 held] ( #42485 )
...
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 09:09:52 -07:00
berriai-litellm-provider-info-sync[bot]
4082523596
chore(prices): sync OpenRouter prices: 14 models, 1 deprecated [enrichment failed: OpenRouter, 5 held] ( #42438 )
...
openrouter/~z-ai/glm-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/aion-labs/aion-2.0: max_input_tokens
openrouter/aion-labs/aion-3.0: max_input_tokens
openrouter/aion-labs/aion-3.0-mini: max_input_tokens
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/dots-studio/dots-3-note-preview🆓 deprecation_date
openrouter/moonshotai/kimi-k2.7-code: output_cost_per_token
openrouter/moonshotai/kimi-k3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/openai/gpt-oss-20b: max_tokens, max_output_tokens, supports_prompt_caching, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3-next-80b-a3b-thinking: max_tokens, max_output_tokens
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3-flash:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 08:42:25 -07:00
tin-berri
5a8c4f48e4
feat(router): native compact-to-fit across conversation APIs ( #42074 )
...
* feat(router): native compact-to-fit across conversation APIs
* fix(router): preserve compaction admission and shared client boundaries
* fix(router): honor compaction fit fallbacks and router-scoped access
* fix(router): charge compaction usage to caller token limits
* test(http): keep FastAPI inside proxy tests
* fix(router): check compactor capacity before skipping escalation
2026-09-21 22:52:29 -07:00
berriai-litellm-provider-info-sync[bot]
798d45970b
chore(prices): sync OpenRouter prices: 2 models ( #42418 )
...
openrouter/nex-agi/nex-n2.5-mini: supports_response_schema
openrouter/nex-agi/nex-n2.5-pro: supports_tool_choice, supports_response_schema, supports_function_calling
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 22:29:26 -07:00
berriai-litellm-provider-info-sync[bot]
0132f34356
chore(prices): sync OpenRouter prices: 2 models, 2 new ( #42407 )
...
openrouter/nex-agi/nex-n2.5-mini: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/nex-agi/nex-n2.5-pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:40:25 -07:00
berriai-litellm-provider-info-sync[bot]
dfd8ffc545
chore(prices): sync OpenRouter prices: 2 models, 2 deprecated ( #42381 )
...
openrouter/nex-agi/nex-n2.5-mini🆓 deprecation_date
openrouter/nex-agi/nex-n2.5-pro🆓 deprecation_date
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:11:22 -07:00
devin-ai-integration[bot]
0fd1c191ca
feat(fal_ai): add queue-only /fal_ai pass-through route with spend tracking ( #42360 )
2026-09-22 02:59:58 +00:00
devin-ai-integration[bot]
1106b16745
feat(openrouter): price typesafe/jev-1.13 and add an openrouter decisions pass-through ( #42301 )
2026-09-22 02:44:15 +00:00
devin-ai-integration[bot]
b833e1fc4c
feat(fal_ai): add flux-lora-depth image edits and moondream3 chat completions ( #42334 )
...
* feat(fal_ai): add flux-lora-depth image edits and moondream3 chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(fal_ai): retrigger codecov processing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject multi-turn and system messages for moondream3 chat
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): return 400 for invalid moondream3 chat requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject moondream3 responses missing output or usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject streaming moondream3 requests before dispatch
stream never reaches optional_params, so the transform_request check could not fire; reject in _complete_fal_ai on ctx.stream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 19:01:28 -07:00
kerry-berri
ee73e6391e
Merge pull request #42384 from BerriAI/litellm_xai_manual_price_sync
...
feat(pricing): add xai grok-4.20 aliases and image token prices from /v1/language-models
2026-09-21 18:32:10 -07:00
kerry
d3cf820c48
fix(pricing): keep function calling disabled on xai multi-agent rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:21:58 +00:00
kerry
839cb268fd
fix(openrouter): remove the retired stealth/union-alpha model from the cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:13:51 +00:00
kerry
52ae9534ad
feat(pricing): add xai grok-4.20 aliases and image token prices from /v1/language-models
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:08:11 +00:00
berriai-litellm-provider-info-sync[bot]
66d197f56c
chore(prices): sync OpenRouter prices: 4 models
...
openrouter/~deepseek/deepseek-pro-latest: off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~z-ai/glm-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-22 00:31:00 +00:00
berriai-litellm-provider-info-sync[bot]
c8252a50f5
chore(prices): sync OpenRouter prices: 4 models
...
openrouter/~deepseek/deepseek-pro-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
2026-09-22 00:01:01 +00:00
berriai-litellm-provider-info-sync[bot]
10ac81cfb5
chore(prices): sync OpenRouter prices: 2 models
...
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/meta-llama/llama-4-maverick: input_cost_per_token, output_cost_per_token
2026-09-21 23:30:58 +00:00
kerry-berri
7f550165df
Merge pull request #42363 from BerriAI/litellm_bedrock_kimi_k3_bare_key
...
fix(bedrock): add bare moonshotai.kimi-k3 cost map entry
2026-09-21 16:28:06 -07:00
kerry-berri
cf098ceccf
Merge pull request #42362 from BerriAI/litellm_xiaomi_mimo_v26
...
feat(xiaomi_mimo): add mimo-v2.6-pro and mimo-v2.6-flash cost map rows with live e2e coverage
2026-09-21 16:25:15 -07:00
kerry
9af3363c2f
fix(bedrock): add bare moonshotai.kimi-k3 cost map entry mirroring the global inference profile
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 23:18:18 +00:00
kerry
47a1053065
feat(xiaomi_mimo): add mimo-v2.6-pro and mimo-v2.6-flash cost map rows with live e2e coverage
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 23:12:57 +00:00
berriai-litellm-provider-info-sync[bot]
c371617786
chore(prices): sync OpenRouter prices: 4 models, 3 deprecated
...
openrouter/bytedance-seed/seed-1.6: deprecation_date
openrouter/bytedance-seed/seed-1.6-flash: deprecation_date
openrouter/bytedance-seed/seed-2.0-code: deprecation_date
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 23:01:01 +00:00
berriai-litellm-provider-info-sync[bot]
cf7988a2df
chore(prices): sync OpenRouter prices: 3 models
...
openrouter/~z-ai/glm-flash-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3-flash: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 22:31:01 +00:00
kerry-berri
0f47056d70
Merge pull request #42337 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 15:10:09 -07:00
berriai-litellm-provider-info-sync[bot]
b305928422
chore(prices): sync AWS Bedrock prices: 6 models [enrichment failed: AWS Bedrock, 4 held]
...
minimax.minimax-m2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
minimax.minimax-m2.1: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
minimax.minimax-m2.5: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
moonshot.kimi-k2-thinking: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
openai.gpt-oss-safeguard-120b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
openai.gpt-oss-safeguard-20b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
2026-09-21 22:01:14 +00:00
berriai-litellm-provider-info-sync[bot]
0eea01427d
chore(prices): sync OpenRouter prices: 3 models
...
openrouter/~deepseek/deepseek-v4-flash-latest: output_cost_per_token
openrouter/deepseek/deepseek-v4-flash-0731: output_cost_per_token
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 22:01:13 +00:00
berriai-litellm-provider-info-sync[bot]
8c64f82bb3
chore(prices): sync OpenRouter prices: 1 model
...
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:31:16 +00:00
berriai-litellm-provider-info-sync[bot]
a8cd414377
chore(prices): sync OpenRouter prices: 1 model
...
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:01:19 +00:00