Commit graph

2360 commits

Author SHA1 Message Date
devin-ai-integration[bot]
c835a1a982
feat(anthropic): add Claude Opus 5.5 (#42489)
Adds the anthropic cost map entry for claude-opus-5-5 at $4/$20 per MTok
with $5 per MTok 5m cache writes, $8 per MTok 1h cache writes, $0.20 per
MTok cache reads (0.05x base), and fast mode at 2x. The entry sets
thinking_always_on (Opus 5.5 cannot turn thinking off) and
supports_forced_tool_use false (tool_choice required/named 400s, same as
Fable 5.1), mirrors that model by omitting thinking cache preservation,
and registers the model in the setup wizard provider list

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 09:52:35 -07:00
berriai-litellm-provider-info-sync[bot]
226fa7845a
chore(prices): sync OpenRouter prices: 1 model [enrichment failed: OpenRouter, 5 held] (#42485)
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 09:09:52 -07:00
berriai-litellm-provider-info-sync[bot]
4082523596
chore(prices): sync OpenRouter prices: 14 models, 1 deprecated [enrichment failed: OpenRouter, 5 held] (#42438)
openrouter/~z-ai/glm-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/aion-labs/aion-2.0: max_input_tokens
openrouter/aion-labs/aion-3.0: max_input_tokens
openrouter/aion-labs/aion-3.0-mini: max_input_tokens
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/dots-studio/dots-3-note-preview🆓 deprecation_date
openrouter/moonshotai/kimi-k2.7-code: output_cost_per_token
openrouter/moonshotai/kimi-k3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/openai/gpt-oss-20b: max_tokens, max_output_tokens, supports_prompt_caching, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3-next-80b-a3b-thinking: max_tokens, max_output_tokens
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3-flash:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3:batch: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 08:42:25 -07:00
tin-berri
5a8c4f48e4
feat(router): native compact-to-fit across conversation APIs (#42074)
* feat(router): native compact-to-fit across conversation APIs

* fix(router): preserve compaction admission and shared client boundaries

* fix(router): honor compaction fit fallbacks and router-scoped access

* fix(router): charge compaction usage to caller token limits

* test(http): keep FastAPI inside proxy tests

* fix(router): check compactor capacity before skipping escalation
2026-09-21 22:52:29 -07:00
berriai-litellm-provider-info-sync[bot]
798d45970b
chore(prices): sync OpenRouter prices: 2 models (#42418)
openrouter/nex-agi/nex-n2.5-mini: supports_response_schema
openrouter/nex-agi/nex-n2.5-pro: supports_tool_choice, supports_response_schema, supports_function_calling

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 22:29:26 -07:00
berriai-litellm-provider-info-sync[bot]
0132f34356
chore(prices): sync OpenRouter prices: 2 models, 2 new (#42407)
openrouter/nex-agi/nex-n2.5-mini: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/nex-agi/nex-n2.5-pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:40:25 -07:00
berriai-litellm-provider-info-sync[bot]
dfd8ffc545
chore(prices): sync OpenRouter prices: 2 models, 2 deprecated (#42381)
openrouter/nex-agi/nex-n2.5-mini🆓 deprecation_date
openrouter/nex-agi/nex-n2.5-pro🆓 deprecation_date

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-21 21:11:22 -07:00
devin-ai-integration[bot]
0fd1c191ca
feat(fal_ai): add queue-only /fal_ai pass-through route with spend tracking (#42360) 2026-09-22 02:59:58 +00:00
devin-ai-integration[bot]
1106b16745
feat(openrouter): price typesafe/jev-1.13 and add an openrouter decisions pass-through (#42301) 2026-09-22 02:44:15 +00:00
devin-ai-integration[bot]
b833e1fc4c
feat(fal_ai): add flux-lora-depth image edits and moondream3 chat completions (#42334)
* feat(fal_ai): add flux-lora-depth image edits and moondream3 chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(fal_ai): retrigger codecov processing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fal_ai): reject multi-turn and system messages for moondream3 chat

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fal_ai): return 400 for invalid moondream3 chat requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fal_ai): reject moondream3 responses missing output or usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fal_ai): reject streaming moondream3 requests before dispatch

stream never reaches optional_params, so the transform_request check could not fire; reject in _complete_fal_ai on ctx.stream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 19:01:28 -07:00
kerry-berri
ee73e6391e
Merge pull request #42384 from BerriAI/litellm_xai_manual_price_sync
feat(pricing): add xai grok-4.20 aliases and image token prices from /v1/language-models
2026-09-21 18:32:10 -07:00
kerry
d3cf820c48 fix(pricing): keep function calling disabled on xai multi-agent rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:21:58 +00:00
kerry
839cb268fd fix(openrouter): remove the retired stealth/union-alpha model from the cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:13:51 +00:00
kerry
52ae9534ad feat(pricing): add xai grok-4.20 aliases and image token prices from /v1/language-models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:08:11 +00:00
berriai-litellm-provider-info-sync[bot]
66d197f56c
chore(prices): sync OpenRouter prices: 4 models
openrouter/~deepseek/deepseek-pro-latest: off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~z-ai/glm-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-22 00:31:00 +00:00
berriai-litellm-provider-info-sync[bot]
c8252a50f5
chore(prices): sync OpenRouter prices: 4 models
openrouter/~deepseek/deepseek-pro-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
2026-09-22 00:01:01 +00:00
berriai-litellm-provider-info-sync[bot]
10ac81cfb5
chore(prices): sync OpenRouter prices: 2 models
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/meta-llama/llama-4-maverick: input_cost_per_token, output_cost_per_token
2026-09-21 23:30:58 +00:00
kerry-berri
7f550165df
Merge pull request #42363 from BerriAI/litellm_bedrock_kimi_k3_bare_key
fix(bedrock): add bare moonshotai.kimi-k3 cost map entry
2026-09-21 16:28:06 -07:00
kerry-berri
cf098ceccf
Merge pull request #42362 from BerriAI/litellm_xiaomi_mimo_v26
feat(xiaomi_mimo): add mimo-v2.6-pro and mimo-v2.6-flash cost map rows with live e2e coverage
2026-09-21 16:25:15 -07:00
kerry
9af3363c2f fix(bedrock): add bare moonshotai.kimi-k3 cost map entry mirroring the global inference profile
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 23:18:18 +00:00
kerry
47a1053065 feat(xiaomi_mimo): add mimo-v2.6-pro and mimo-v2.6-flash cost map rows with live e2e coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 23:12:57 +00:00
berriai-litellm-provider-info-sync[bot]
c371617786
chore(prices): sync OpenRouter prices: 4 models, 3 deprecated
openrouter/bytedance-seed/seed-1.6: deprecation_date
openrouter/bytedance-seed/seed-1.6-flash: deprecation_date
openrouter/bytedance-seed/seed-2.0-code: deprecation_date
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 23:01:01 +00:00
berriai-litellm-provider-info-sync[bot]
cf7988a2df
chore(prices): sync OpenRouter prices: 3 models
openrouter/~z-ai/glm-flash-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3-flash: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 22:31:01 +00:00
kerry-berri
0f47056d70
Merge pull request #42337 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 15:10:09 -07:00
berriai-litellm-provider-info-sync[bot]
b305928422
chore(prices): sync AWS Bedrock prices: 6 models [enrichment failed: AWS Bedrock, 4 held]
minimax.minimax-m2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
minimax.minimax-m2.1: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
minimax.minimax-m2.5: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
moonshot.kimi-k2-thinking: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
openai.gpt-oss-safeguard-120b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
openai.gpt-oss-safeguard-20b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
2026-09-21 22:01:14 +00:00
berriai-litellm-provider-info-sync[bot]
0eea01427d
chore(prices): sync OpenRouter prices: 3 models
openrouter/~deepseek/deepseek-v4-flash-latest: output_cost_per_token
openrouter/deepseek/deepseek-v4-flash-0731: output_cost_per_token
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 22:01:13 +00:00
berriai-litellm-provider-info-sync[bot]
8c64f82bb3
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:31:16 +00:00
berriai-litellm-provider-info-sync[bot]
a8cd414377
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:01:19 +00:00
berriai-litellm-provider-info-sync[bot]
07aadfd192
chore(prices): sync OpenRouter prices: 4 models, 3 new
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-pro-ultraspeed: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 20:31:35 +00:00
kerry-berri
5216844c40
Merge pull request #42286 from BerriAI/litellm_fal_ai_minimax_h3
feat(fal_ai): add MiniMax H3 text-to-video and reference-to-video
2026-09-21 13:16:08 -07:00
kerry-berri
d79eadfe8c
Merge pull request #42297 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 13:11:01 -07:00
kerry-berri
dd73c9fe3d
Merge pull request #42298 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 1 model [enrichment failed: AWS Bedrock, 16 held]
2026-09-21 13:10:57 -07:00
kerry-berri
246a6ea54a
Merge pull request #42282 from BerriAI/litellm_fal_price_from_response_dims
fix(fal_ai): price images from the dimensions fal returns
2026-09-21 13:08:49 -07:00
berriai-litellm-provider-info-sync[bot]
4a825a7259
chore(prices): sync AWS Bedrock prices: 1 model [enrichment failed: AWS Bedrock, 16 held]
zai.glm-5: supports_vision, supports_audio_input, supports_response_schema
2026-09-21 20:01:35 +00:00
berriai-litellm-provider-info-sync[bot]
a2bb1f7e57
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 20:01:21 +00:00
Mateo Wang
2e35ae1065
Merge pull request #42049 from BerriAI/litellm_mantle_native_anthropic_messages
feat(bedrock_mantle): serve /v1/messages for Claude models on Mantle's native Anthropic Messages API
2026-09-21 12:52:16 -07:00
kerry-berri
c3663aaac8
Merge pull request #42289 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 12:41:02 -07:00
berriai-litellm-provider-info-sync[bot]
f851e6ddb7
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 18 held]
us.moonshotai.kimi-k3: 
zai.glm-4.7: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
zai.glm-4.7-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:31:34 +00:00
berriai-litellm-provider-info-sync[bot]
cbf5bb5e71
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:31:29 +00:00
kerry-berri
f92ff60ebd
Merge pull request #42280 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 12:11:43 -07:00
kerry
9b54c4b077 refactor(fal_ai): bill flux dev per 1024x1024 megapixel and drop ImageResponse retyping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 19:10:33 +00:00
kerry
adc4e6a132 fix(fal_ai): price images from the dimensions fal returns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 19:10:32 +00:00
kerry
8d73ce756a feat(fal_ai): add MiniMax H3 text-to-video and reference-to-video
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 19:03:00 +00:00
berriai-litellm-provider-info-sync[bot]
a6842da112
chore(prices): sync OpenRouter prices: 3 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:01:36 +00:00
berriai-litellm-provider-info-sync[bot]
ae06a6478f
chore(prices): sync AWS Bedrock prices: 4 models [enrichment failed: AWS Bedrock, 26 held]
global.moonshotai.kimi-k3: 
qwen.qwen3-coder-next: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-next-80b-a3b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-vl-235b-a22b: max_tokens, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:01:36 +00:00
mateo-berri
2fe5c8990e fix(bedrock_mantle): bill Mantle's un-versioned Claude ids from a Mantle cost row
Mantle serves anthropic.claude-haiku-4-5 without the dated -20251001-v1:0
suffix the Bedrock row carries, so the native route billed it at 0. Add a
bedrock_mantle/anthropic.claude-haiku-4-5 row and let a
bedrock_mantle/<region>/<model> name fall back to the region-free
bedrock_mantle/<model> row before the provider-prefixed lookup. Also
satisfy the mutable-collection gate in the native messages transformation.
2026-09-21 12:00:38 -07:00
kerry-berri
8ee8613b07
Merge pull request #42271 from BerriAI/litellm_bedrock_kimi_k3_us_cris
feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
2026-09-21 11:51:54 -07:00
kerry-berri
12379aa1e3
Merge pull request #42095 from BerriAI/litellm_fal_gpt_image_25_flux_dev_edits
feat(fal_ai): add gpt-image-2.5 flare/sunburst, flux/dev and image edits
2026-09-21 11:45:45 -07:00
kerry-berri
24f0f373fc
Merge pull request #42270 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 6 models
2026-09-21 11:41:51 -07:00
kerry
29837b422e feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 18:32:55 +00:00