Commit graph

1260 commits

Author SHA1 Message Date
Sameer Kankute
18a524a1dc
Merge pull request #20408 from BerriAI/litellm_fix_model+map_streaming
Fix supports_native_streaming for gemini and vertex ai and claude models
2026-02-04 17:56:48 +05:30
Sameer Kankute
b2feedc469
Merge pull request #20318 from BerriAI/litellm_oss_staging_02_03_2026
feat(guardrails): implement team-based isolation guardrails mgmnt (#1…
2026-02-04 17:49:30 +05:30
Sameer Kankute
5cec5c3bef fix model map json 2026-02-04 14:48:49 +05:30
Sameer Kankute
0ed5e51f9b Fix supports_native_streaming for anthropic models 2026-02-04 14:44:31 +05:30
Sameer Kankute
20c21c63d9 Fix supports_native_streaming for gemini and vertex ai models 2026-02-04 14:41:32 +05:30
Sameer Kankute
216cf4cac5 Add gemini deep research in model cost map 2026-02-04 13:38:13 +05:30
Sameer Kankute
25fa1ad4e7
Merge pull request #20386 from naaa760/fix/extra-head-chat-comp-brid
fix(proxy): forward extra headers in chat
2026-02-04 09:11:43 +05:30
Felipe Rodrigues Gare Carnielli
ea19d8dbf6 fixing glm-4.7 input cost per token 2026-02-03 09:57:00 -03:00
krauckbot
17c0a88a60
fix: add missing capability flags to vercel_ai_gateway models (#20276)
67 vercel_ai_gateway models were missing capability flags (supports_vision,
supports_function_calling, supports_tool_choice, supports_response_schema).

These capabilities were inferred from the corresponding direct provider entries
for the same models (e.g., vercel_ai_gateway/anthropic/claude-3.5-sonnet now has
the same capabilities as anthropic/claude-3.5-sonnet).

Models fixed include:
- Claude 3/3.5/3.7 (Anthropic)
- GPT-4/5 variants (OpenAI)
- Gemini 2.0/2.5 (Google)
- Grok 3/4 (xAI)
- Mistral/Mixtral variants
- Qwen models
- DeepSeek models
- And more

This ensures consistent capability reporting across providers for the same
underlying models.

Co-authored-by: krauckbot <krauckbot123@gmail.com>
2026-02-02 22:07:17 -08:00
Felipe Rodrigues Gare Carnielli
f32bd8474e adding together ai models to litellm models json 2026-02-03 00:25:58 -03:00
krauckbot
c4bbd56a56
feat: add Kimi K2.5 model entry for Moonshot provider (#20273)
Add moonshot/kimi-k2.5 model with:
- Input cost: $0.60/M tokens (6e-07)
- Output cost: $3.00/M tokens (3e-06)
- Cache read cost: $0.10/M tokens (1e-07)
- 256K context window
- Vision, function calling, tool choice, web search support

Reference: https://huggingface.co/moonshotai/Kimi-K2.5

Note: K2.5 thinking mode is controlled via API parameters, not a separate model ID.

Co-authored-by: krauckbot <krauckbot123@gmail.com>
2026-02-02 13:28:57 -08:00
Chesars
3215dc4d4e feat(vertex_ai): add global endpoint support for Qwen MaaS models
Fixes #19788

- Add `supported_regions: ["global"]` to Qwen MaaS models in model_prices_and_context_window.json
- Update `get_supported_regions()` to read directly from `model_cost` dict
- Update `get_complete_vertex_url()` to use `get_vertex_region()` for global-only models
- Update `create_vertex_url()` to generate correct URL for global location (without region prefix)
- Add tests for Qwen global endpoint support
2026-02-02 18:13:10 +05:30
Sameer Kankute
7773a92069
Merge pull request #20258 from BerriAI/litellm_cerebras_reasoning
fix: add reasoning param support for GPT OSS cerebras
2026-02-02 17:42:28 +05:30
Sameer Kankute
415c26f281 fix: add reasoning param support for GPT OSS cerebras 2026-02-02 17:20:04 +05:30
cscguochang-agent
76407bcf37 feat(bedrock): add base cache costs for sonnet v1 (#20214) 2026-02-01 09:36:51 +08:00
Ishaan Jaff
5345a763c2
[Feat] v2 - Logs view with side panel and improved UX (#20091)
* init: azure_ai/azure-model-router

* show additional_costs in CostBreakdown

* UI show cost breakdown fields

* feat: dedicated cost calc for azure ai

* test_azure_ai_model_router

* docs azure model router

* test azure model router

* fix transfrom

* Add transform file

* fix:feat: route to config

* v0 - looks decen view

* refactored code

* fix ui

* fixes ui

* complete v2 viewer

* address feedback

* address feedback
2026-01-30 18:34:13 -08:00
Sameer Kankute
a8054264ae
Merge pull request #19975 from BerriAI/litellm_oss_staging_01_29_2026
Litellm oss staging 01 29 2026
2026-01-30 16:58:28 +05:30
Sameer Kankute
fdb4b54add
Merge pull request #20009 from genga6/fix/#20006-update-max-input-tokens-for-gpt-5.2-codex
Fix `max_input_tokens` for `gpt-5.2-codex`
2026-01-30 16:18:53 +05:30
Sameer Kankute
eb50c780e9
Merge branch 'main' into litellm_oss_staging_01_29_2026 2026-01-30 09:03:05 +05:30
Ishaan Jaff
f7e1a22947
[Feat] New Model - amazon.nova-2-pro-preview-20251202-v1:0 (#20033)
* init: amazon.nova-2-pro-preview-20251202-v1:0

* init: nova amazon.nova-2-pro

* add s3_vectors
2026-01-29 16:55:55 -08:00
Takumi Matsuzawa
fa54c241e0 Fix max_input_tokens for gpt-5.2-codex 2026-01-29 15:39:17 +00:00
Sameer Kankute
df072979e5
Merge branch 'main' into litellm_oss_staging_01_28_2026 2026-01-29 17:39:42 +05:30
Sameer Kankute
ef15861fde Fix: litellm_fix_robotic_model_map_entry 2026-01-29 17:26:50 +05:30
Aaron Yim
d4031c8ba6
Add OpenRouter Kimi K2.5 (#19872)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 22:34:48 -08:00
Sameer Kankute
be8a76f270 fix gemini gemini-robotics-er-1.5-preview entry 2026-01-29 09:06:44 +05:30
rushilchugh01
562f0a0282
feat: Add new OpenRouter models: xiaomi/mimo-v2-flash, z-ai/glm-4.7, z-ai/glm-4.7-flash, and minimax/minimax-m2.1. to model prices and context window (#19938)
Co-authored-by: Rushil Chugh <Rushil>
2026-01-28 18:56:20 -08:00
Sameer Kankute
3ab1b9f543 Fix gemini-robotics-er-1.5-preview name 2026-01-28 21:13:37 +05:30
Sameer Kankute
9fe8b12f44
Merge pull request #19924 from BerriAI/litellm_minimax_reasoning_caching_1
Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi
2026-01-28 18:04:13 +05:30
Sameer Kankute
f6ead49afe Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi 2026-01-28 17:25:26 +05:30
Sameer Kankute
42a0d576f3
Merge pull request #19910 from BerriAI/main
merge 01 27
2026-01-28 08:30:47 +05:30
Jay Prajapati
6a9d41234f
fix: allow tool_choice for Azure GPT-5 chat models (#19813)
* fix: don't treat gpt-5-chat as GPT-5 reasoning

* fix: mark azure gpt-5-chat as supporting tool_choice

* test: cover gpt-5-chat params on azure/openai
2026-01-27 17:51:13 -08:00
Sameer Kankute
c8a93c7d81
Merge pull request #19845 from BerriAI/litellm_gemini-robotics-er-1.5-preview2
Add  Gemini Robotics-ER 1.5 preview support
2026-01-27 17:47:44 +05:30
Sameer Kankute
c834d7d1fe
Merge branch 'main' into litellm_oss_staging_01_27_2026 2026-01-27 17:11:15 +05:30
Sameer Kankute
9a2750f8ec
Merge pull request #19617 from BerriAI/litellm_oss_staging_01_23_2026
Litellm oss staging 01 23 2026
2026-01-27 16:55:32 +05:30
Sameer Kankute
cf012a2f65 Add gemini-robotics-er-1.5-preview model in model map 2026-01-27 13:58:03 +05:30
Cesar Garcia
e4a557d95f
fix(xai): correct cached token cost calculation for xAI models (#19772)
* fix(azure): use generic cost calculator for audio token pricing

Azure audio models were charging audio output tokens at the text token
rate instead of the correct audio token rate. This resulted in costs
being ~6.65x lower than expected.

The fix replaces Azure's custom cost calculation logic with the generic
cost calculator that properly handles text, audio, cached, reasoning,
and image tokens.

Fixes #19764

* fix(xai): correct cached token cost calculation for xAI models

- Fix double-counting issue where xAI reports text_tokens = prompt_tokens
  (including cached), causing tokens to be charged twice
- Add cache_read_input_token_cost to xAI grok-3 and grok-3-mini model variants
- Detection: when text_tokens + cached_tokens > prompt_tokens, recalculate
  text_tokens = prompt_tokens - cached_tokens

xAI pricing (25% of input for cached):
- grok-3 variants: $0.75/M cached (input $3/M)
- grok-3-mini variants: $0.075/M cached (input $0.30/M)
2026-01-26 21:00:35 -08:00
Cesar Garcia
0d45b01069
fix(models): set gpt-5.2-codex mode to responses for Azure and OpenRouter (#19770)
Fixes #19754

The gpt-5.2-codex model only supports the responses API, not chat completions.
Updated azure/gpt-5.2-codex and openrouter/openai/gpt-5.2-codex entries to use
mode: "responses" and supported_endpoints: ["/v1/responses"].
2026-01-26 20:36:10 -08:00
Cesar Garcia
31a8d76d11
Update Gemini 2.0 Flash deprecation dates to March 31, 2026 (#19592)
Google announced that Gemini 2.0 Flash and Flash Lite models will be discontinued on March 31, 2026. Updated deprecation_date field for all affected model variants across different providers (vertex_ai, gemini, deepinfra, openrouter, vercel_ai_gateway).

Models updated:
- gemini-2.0-flash (added deprecation date)
- gemini-2.0-flash-001 (updated from 2026-02-05)
- gemini-2.0-flash-lite (added deprecation date)
- gemini-2.0-flash-lite-001 (updated from 2026-02-25)

All variants now correctly reflect the March 31, 2026 shutdown date.
2026-01-23 20:36:36 -08:00
John Greek
26a2c90818
[Fix] Anthropic models on Azure AI cache pricing (#19532) (#19614) 2026-01-22 20:00:40 -08:00
拐爷&&老拐瘦
3372430d40
Add pricing for volcengine models (deepseek-v3-2, glm-4-7, kimi-k2-thinking) (#19335)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-22 19:47:21 -08:00
Sameer Kankute
ad1edd38d5
Merge branch 'main' into litellm_staging_01_21_2026 2026-01-22 17:56:40 +05:30
Sameer Kankute
36f3250016
Merge pull request #19500 from Chesars/fix/audio-model-pricing
fix(pricing): correct audio token costs for gpt-4o-audio-preview models
2026-01-22 09:18:55 +05:30
Cesar Garcia
4106d24215
feat: add GMI Cloud provider support (#19376)
* feat: add GMI Cloud provider support

Add GMI Cloud as an OpenAI-compatible provider with:
- Provider configuration in providers.json
- Documentation page with usage examples
- Model pricing for 16 models (Claude, GPT, DeepSeek, Gemini, etc.)
- Sidebar entry for docs navigation

* Add gmi_cloud to provider_endpoints_support.json

Add provider entry to pass CI validation check that ensures all
providers in openai_like/providers.json are documented.

* Fix provider key: gmi_cloud -> gmi

Match the provider key with providers.json

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-01-21 15:48:15 -08:00
Chesars
796e93552d Fix gpt-audio models pricing and add dated snapshots
- Fix audio token pricing for gpt-audio ($32/$64 per 1M, not $40/$80)
- Add gpt-audio-2025-08-28 snapshot (OpenAI returns this in responses)
- Add gpt-audio-mini-2025-10-06 and gpt-audio-mini-2025-12-15 snapshots
- Add missing fields: supported_endpoints, supported_modalities,
  supported_output_modalities, supports_native_streaming, etc.
2026-01-21 13:05:38 -03:00
Chesars
da8770004c Add gpt-audio and gpt-audio-mini models to pricing
Fixes #19490 - adds missing OpenAI audio models with correct pricing:

gpt-audio:
- Text: $2.50/$10.00 per 1M tokens (input/output)
- Audio: $40/$80 per 1M tokens (input/output)

gpt-audio-mini:
- Text: $0.60/$2.40 per 1M tokens (input/output)
- Audio: $10/$20 per 1M tokens (input/output)
2026-01-21 12:47:10 -03:00
Chesars
a0cfb56801 fix(pricing): correct audio token costs for gpt-4o-audio-preview models
Update audio token pricing for gpt-4o-audio-preview and
gpt-4o-audio-preview-2024-10-01 to match OpenAI's official pricing:

- input_cost_per_audio_token: 0.0001 -> 4e-05 ($40/1M tokens)
- output_cost_per_audio_token: 0.0002 -> 8e-05 ($80/1M tokens)

The previous values were 2.5x higher than OpenAI's actual pricing.
2026-01-21 10:46:03 -03:00
Sameer Kankute
540370a1aa
Merge pull request #19479 from BerriAI/litellm_sarvam_int
Add support for sarvam models
2026-01-21 19:03:52 +05:30
Connor Luebbehusen
b810a68f89
fix: correct gemini-2.5-flash-lite audio input and cache read pricing 2026-01-21 05:58:38 -05:00
Sameer Kankute
46c0f903a3 Add support for sarvam models 2026-01-21 15:17:26 +05:30
Sameer Kankute
9e1275b76c
Merge branch 'main' into litellm_staging_01_19_2026 2026-01-20 19:19:36 +05:30