Commit graph

1813 commits

Author SHA1 Message Date
Nicolas Herment
50f625433d
Revert incorrect changes to sonnet-4 max output tokens (#14933) 2025-09-26 11:19:52 -07:00
Ishaan Jaff
360befa216
[Feat] Add support for Gemini 2.5 Flash and Flash-lite preview models (09-2025 release) (#14948)
* add gemini-2.5-flash-preview-09-2025

* docs add gemini-2.5-flash-preview-09-2025 model family
2025-09-26 09:51:38 -07:00
Daniel Klein
de795a4531 Fix inconsistent token configs for gpt-5 models 2025-09-26 09:43:26 -04:00
Toy-97
6c95bd926f
update: DeepInfra model data refresh [2025-09-26]
Added models:
deepinfra/deepseek-ai/DeepSeek-V3.1-Terminus

Removed models:
deepinfra/zai-org/GLM-4.5-Air

Modified models:
deepinfra/NousResearch/Hermes-3-Llama-3.1-70B:
   - input_cost_per_token: 1.2e-07 → 3e-07

deepinfra/Qwen/Qwen3-32B:
   - output_cost_per_token: 3e-07 → 2.8e-07

deepinfra/Qwen/Qwen3-Next-80B-A3B-Instruct:
   - max_tokens: 4096 → 262144
   - max_output_tokens: 4096 → 262144
   - max_input_tokens: 4096 → 262144

deepinfra/Qwen/Qwen3-Next-80B-A3B-Thinking:
   - max_tokens: 4096 → 262144
   - max_output_tokens: 4096 → 262144
   - max_input_tokens: 4096 → 262144

deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507:
   - input_cost_per_token: 1.3e-07 → 9e-08

deepinfra/meta-llama/Meta-Llama-3.1-70B-Instruct:
   - input_cost_per_token: 2.3e-07 → 4e-07

deepinfra/google/gemini-2.5-flash:
   - output_cost_per_token: 1.75e-06 → 2.5e-06
   - input_cost_per_token: 2.1e-07 → 3e-07

deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo:
   - output_cost_per_token: 2e-08 → 3e-08
   - input_cost_per_token: 1.5e-08 → 2e-08

deepinfra/meta-llama/Llama-3.2-3B-Instruct:
   - output_cost_per_token: 2.4e-08 → 2e-08
   - input_cost_per_token: 1.2e-08 → 2e-08

deepinfra/Sao10K/L3-8B-Lunaris-v1-Turbo:
   - input_cost_per_token: 2e-08 → 4e-08

deepinfra/openai/gpt-oss-120b:
   - input_cost_per_token: 9e-08 → 5e-08

deepinfra/google/gemini-2.5-pro:
   - output_cost_per_token: 7e-06 → 1e-05
   - input_cost_per_token: 8.75e-07 → 1.25e-06

deepinfra/NousResearch/Hermes-3-Llama-3.1-405B:
   - output_cost_per_token: 8e-07 → 1e-06
   - input_cost_per_token: 7e-07 → 1e-06

deepinfra/Qwen/Qwen3-235B-A22B:
   - output_cost_per_token: 6e-07 → 5.4e-07
   - input_cost_per_token: 1.3e-07 → 1.8e-07

deepinfra/nvidia/Llama-3.1-Nemotron-70B-Instruct:
   - output_cost_per_token: 3e-07 → 6e-07
   - input_cost_per_token: 1.2e-07 → 6e-07

deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo:
   - output_cost_per_token: 1.2e-07 → 3.9e-07
   - input_cost_per_token: 3.8e-08 → 1.3e-07

deepinfra/deepseek-ai/DeepSeek-V3-0324:
   - input_cost_per_token: 2.8e-07 → 2.5e-07
   - cache_read_input_token_cost: 2.24e-07 → None

deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506:
   - output_cost_per_token: 1e-07 → 2e-07
   - input_cost_per_token: 5e-08 → 7.5e-08

deepinfra/Qwen/Qwen3-235B-A22B-Thinking-2507:
   - output_cost_per_token: 6e-07 → 2.9e-06
   - input_cost_per_token: 1.3e-07 → 3e-07

deepinfra/zai-org/GLM-4.5:
   - output_cost_per_token: 2e-06 → 1.6e-06
   - input_cost_per_token: 5.5e-07 → 4e-07

deepinfra/mistralai/Mixtral-8x7B-Instruct-v0.1:
   - output_cost_per_token: 2.4e-07 → 4e-07
   - input_cost_per_token: 8e-08 → 4e-07

deepinfra/openai/gpt-oss-20b:
   - output_cost_per_token: 1.6e-07 → 1.5e-07

deepinfra/google/gemma-3-27b-it:
   - output_cost_per_token: 1.7e-07 → 1.6e-07
2025-09-26 20:11:26 +08:00
Krish Dholakia
2f3155c2ee
Merge pull request #14858 from oytunkutrup1/litellm_fix_gpt3.5_price_fix
GPT-3.5-Turbo price updated.
2025-09-25 23:47:43 -07:00
Krish Dholakia
6a1be4722e
Merge pull request #14879 from huangyafei/update_price
Add gpt-5 and gpt-5-codex to OpenRouter cost map
2025-09-25 23:46:35 -07:00
Ishaan Jaffer
1585750c1b test fix 2025-09-24 21:45:42 -07:00
huangyafei
107c45da9d Add gpt-5-codex to OpenRouter cost map 2025-09-25 11:05:56 +08:00
huangyafei
1b95940d18 Add gpt-5 to OpenRouter cost map 2025-09-25 11:02:54 +08:00
Luis Felipe Salazar Ucros
c6cb36186c
Add sambanova deepseek v3.1 and gpt-oss-120b (#14866)
* add sambanova deepseek v3.1 and gpt-oss-120b

* add sambanova deepseek v3.1 and gpt-oss-120b
2025-09-24 10:54:22 -07:00
Mubashir Osmani
0cd91a82d2
Added vertex_ai/qwen models and azure/gpt-5-codex (#14844)
* added qwen models and gpt-5-codex

* fix flaky test

* fix failing test
2025-09-24 10:40:00 -07:00
onlylonly
22eef373eb
feat: New model - Add support for Qwen models family & Deepseek 3.1 to Amazon Bedrock (#14845)
* New model - Add Bedrock deepseek v3.1 model - "deepseek.v3-v1:0"

* New model - Add Bedrock Qwen models - "qwen.qwen3-coder-480b-a35b-v1:0", "qwen.qwen3-235b-a22b-2507-v1:0", "qwen.qwen3-coder-30b-a3b-v1:0", "qwen.qwen3-32b-v1:0"

* fix: add "deepseek.v3-v1:0" in  litellm/model_prices_and_context_window_backup.json
2025-09-24 09:58:23 -07:00
oytun.kutrup
410e4956c4
GPT-3.5-Turbo price updated. 2025-09-24 15:32:23 +03:00
Krish Dholakia
d5ced7e068
Merge pull request #14639 from danielmklein/main
fix: update sonnet 4 configs to reflect million-token context window pricing
2025-09-23 18:06:15 -07:00
Ishaan Jaff
fff1d1a9f9
Add Vertex AI Qwen3 models to pricing and context window (#14828)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: ishaan <ishaan@berri.ai>
2025-09-23 16:16:57 -07:00
Ishaan Jaff
8443eff4a2
feat: add xai/grok-4-fast models (#14833) 2025-09-23 16:01:44 -07:00
Krish Dholakia
88f9cad886
Merge pull request #14796 from BerriAI/litellm_anthopic_token_count_issue
Add service_tier based pricing support for openai[ BOTH Service & Priority Support]
2025-09-22 22:38:59 -07:00
Sameerlite
d2208023d2 Add service_tier based pricing support for openai 2025-09-23 10:48:27 +05:30
Gagan Raj Chinka
a5ed3446cd
opnerouter/x-ai/grok-4-fast:free" to model_prices_and_context_window.json (#14779)
* Update model_prices_and_context_window_backup.json

* Update model_prices_and_context_window.json
2025-09-22 21:29:15 -07:00
Anubhav Singh
fde4dcb0b8
Merge branch 'main' into wandb-inference 2025-09-23 00:14:11 +05:30
Tim Elfrink
282617f2bc Fix gemini-2.5-flash-image-preview model routing
- Update mode from 'chat' to 'image_generation' for both model variants
- Ensures correct routing to image generation endpoints
- Resolves 400 'request not supported' error for image generation
2025-09-19 07:58:58 +02:00
Ishaan Jaffer
114d077cc9 fix: model cost map check 2025-09-18 17:52:56 -07:00
Ishaan Jaffer
c1a967992f fix: model cost map check 2025-09-18 17:37:09 -07:00
Ishaan Jaffer
0626affa11 fix: model cost map check 2025-09-18 17:25:28 -07:00
Ishaan Jaff
4c983f985a
[Feat] Add Bedrock Twelve Labs embedding provider support (#14697)
* fix: add 12 labs to bedrock embedding

* fix: get_bedrock_embedding_provider

* test: test_text_embedding

* fix: 12 labs embedding transform

* fix: refactor 12 labs transform logic

* fix: test_e2e_bedrock_embedding

* fix: test_e2e_bedrock_embedding

* feat: add bedrock twelvelabs pricing

* DOCS: docs bedrock embedding

* DOCS: 12 labs bedrock overview

* fix: bedrock embeddings 12 labs
2025-09-18 17:16:45 -07:00
Sameer Kankute
36bedc69ff
Add TwelveLabs marengo model (#14674) 2025-09-18 11:21:35 -07:00
Daniel Klein
b8b30775a4 fix: update sonnet 4 configs to reflect million-context-window pricing 2025-09-17 11:33:23 -04:00
Krrish Dholakia
97cc5f55d6 build(model_prices_and_context_window.json): add claude-3-5 haiku, claude-3-5 sonnet, claude-opus 3, claude-haiku 3 "cache_creation_input_token_cost_above_1hr" pricing 2025-09-16 18:58:49 -07:00
Krrish Dholakia
3188ae9281 build(model_prices_and_context_window.json): add claude-opus-4, claude-3-7-sonnet, claude-sonnet-4 "cache_creation_input_token_cost_above_1hr" pricing 2025-09-16 18:56:24 -07:00
Krrish Dholakia
a2bfd3e476 build(model_prices_and_context_window.json): add claude-opus-4-1 "cache_creation_input_token_cost_above_1hr" pricing 2025-09-16 18:53:56 -07:00
Krrish Dholakia
0fa11c3c25 build(model_prices_and_context_window.json): add "cache_creation_input_token_cost_above_1hr" to all claude-3-5-sonnet models 2025-09-16 18:52:37 -07:00
Krrish Dholakia
1c855385c9 build(model_cost): add cache_creation_input_token_cost_above_1hr pricing 2025-09-16 18:43:57 -07:00
xprilion
470e1b7d2d tinyfix json error 2025-09-16 16:47:51 +05:30
Anubhav Singh
a9667e5930
Merge branch 'BerriAI:main' into wandb-inference 2025-09-16 16:46:58 +05:30
Tim Elfrink
30c3e7b3d3
Fix: Bedrock cross-region inference profile cost calculation (#14566)
* Add tests for Bedrock cross-region inference profile mapping

- Test model mapping lookup works correctly
- Test proxy cost calculation scenario reproduces original issue
- Verify cost calculation returns expected values
- Ensure compatibility with existing test patterns

* Fix Bedrock cross-region inference profile cost calculation

- Add mapping for bedrock/us.anthropic.claude-3-5-haiku-20241022-v1:0
- Sync backup file for local testing consistency
- Resolve proxy spend tracking failures for cross-region profiles
- Maintain identical configuration with standalone profile

Fixes #14458
2025-09-15 07:10:20 -07:00
Anubhav Singh
67276a8151
Merge branch 'main' into wandb-inference 2025-09-15 18:34:16 +05:30
Elias TOURNEUX
ef9d1ddc40
feat: Add OVHCloud AI Endpoints as a provider 2025-09-12 13:37:03 +02:00
Ishaan Jaff
69ef062f55 fix tiered_pricing test 2025-09-11 19:56:44 -07:00
Ishaan Jaff
dda115cc6d
[Feat] Cost Tracking - Add support for Tiered Cost Tracking for Qwen API (Dashscope) (#14471)
* add dashscope logo

* docs fix

* docs fix

* fix supports_batch_calling

* fix naming

* fix input_cost_per_audio_token

* use output_cost_per_reasoning_token

* add tiered_pricing in get_model_info

* test fixes

* fix cost calc

* ruff fix
2025-09-11 18:14:39 -07:00
Ishaan Jaff
258b674dbb fix deepinfra test 2025-09-10 19:39:23 -07:00
xprilion
667481e75b (feat): Add W&B Inference to LiteLLM 2025-09-11 00:07:30 +05:30
Krish Dholakia
e57a05b2dc
Merge pull request #14324 from Toy-97/patch-1
update: DeepInfra model data refresh [2025-09-08]
2025-09-09 22:31:44 -07:00
Krrish Dholakia
076b46c806 build: remove end of life bedrock model 2025-09-09 19:49:22 -07:00
Pedro Azevedo
44dc6d3aef
Fix: Add supports_function_calling for GPT OSS in Bedrock provider (#14375)
* Refactor JSON formatting and remove unnecessary whitespace in model prices and context window

* Fix formatting inconsistencies and remove unnecessary whitespace in model prices JSON
2025-09-09 13:49:32 -07:00
Krrish Dholakia
3dba103e2c fix: fix typo 2025-09-08 18:13:34 -07:00
Toy-97
890ee1abfa
update: DeepInfra model data refresh [2025-09-08]
Removed models:
deepinfra/Qwen/Qwen2.5-Coder-32B-Instruct
deepinfra/deepseek-ai/DeepSeek-V3-0324-Turbo
deepinfra/meta-llama/Llama-3.2-90B-Vision-Instruct
deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-Turbo
deepinfra/meta-llama/Meta-Llama-3-70B-Instruct
deepinfra/mistralai/Devstral-Small-2507
deepinfra/mistralai/Mistral-7B-Instruct-v0.3
deepinfra/mistralai/Mistral-Small-3.1-24B-Instruct-2503

Modified models:
deepinfra/deepseek-ai/DeepSeek-R1-0528:
   - cache_read_input_token_cost: None → 4e-07

deepinfra/google/gemma-3-4b-it:
   - input_cost_per_token: 2e-08 → 4e-08
   - output_cost_per_token: 4e-08 → 8e-08

deepinfra/Qwen/QwQ-32B:
   - input_cost_per_token: 7.5e-08 → 1.5e-07
   - output_cost_per_token: 1.5e-07 → 4e-07

deepinfra/deepseek-ai/DeepSeek-V3-0324:
   - cache_read_input_token_cost: None → 2.24e-07

deepinfra/deepseek-ai/DeepSeek-V3.1:
   - input_cost_per_token: 3e-07 → 2.7e-07
   - cache_read_input_token_cost: None → 2.16e-07

deepinfra/Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo:
   - cache_read_input_token_cost: 1.5e-07 → 2.4e-07

deepinfra/NousResearch/Hermes-3-Llama-3.1-70B:
   - supports_tool_choice: True → False

deepinfra/deepseek-ai/DeepSeek-R1-Turbo:
   - max_output_tokens: 163840 → 40960
   - max_input_tokens: 163840 → 40960
   - max_tokens: 163840 → 40960
2025-09-08 19:17:58 +08:00
Krish Dholakia
ba10173ec7
Merge branch 'main' into heroku-llms 2025-09-06 22:10:20 -07:00
Thomas Rehn
d88771ca49 fix: correct output pricing for gemini-2.5-flash-image-preview https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-flash-image-preview 2025-09-05 16:22:07 +02:00
Krish Dholakia
f67339a86c
Merge pull request #14028 from onlylhf/volcengine-embedding-support
Add Volcengine embedding module with handler and transformation logic
2025-09-04 21:01:25 -07:00
Ishaan Jaff
23ae7170d1
[Feat] Allow using Veo Video Generation through LiteLLM Pass through routes (#14228)
* fix: add follow_redirects=True,

* test_pass_through_with_httpbin_redirect

* cook book veo video

* docs Veo Video Generation with Google AI Studio

* add veo-3.0-generate-preview cost tracking details

* track vertex_video_models
2025-09-03 18:25:43 -07:00