mateo-berri
51514b9123
fix(cost-map): azure/gpt-6-astra accepts reasoning_effort none on Foundry
2026-09-04 17:29:55 -07:00
mateo-berri
3202963f25
feat(cost-map): add azure/gpt-6-astra and azure/us/gpt-6-astra Foundry pricing
2026-09-04 16:57:40 -07:00
mateo
50d6b26a86
fix(registry): mark baseten GLM-5.3 as vision-capable per Baseten vision docs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 20:10:58 +00:00
mateo
1f0611a8b9
fix(registry): drop Together MiniMax M2.7 and revert Qwen2.5 7B Turbo pricing, both non-serverless
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:25:10 +00:00
mateo
7b96a11e5f
feat(registry): add OpenRouter catalog gaps, Fireworks DeepSeek V4 Flash Vision, Together MiniMax M2.7 and Qwen2.5 7B Turbo pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:19:08 +00:00
mateo
0c29f510bc
fix(registry): drop gpt-image-2 text output price, add openrouter minimax-m3 and qwen3.7-plus
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 15:03:29 +00:00
mateo
c707f2fe5d
fix(model_prices): databricks gpt-5-3-codex is served via the Responses API
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:32:11 +00:00
mateo
8e83d6d63d
fix(model_prices): add Databricks Sep-2026 catalog, Azure gpt-realtime-2.x, per-token realtime image pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:06:54 +00:00
mateo
08bfdadb10
chore: merge litellm_internal_staging into litellm_registry_audit_2026_09_02, drop the deleted ocr ledger
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 04:18:09 +00:00
mateo
f292667601
fix(registry): mark gpt-daybreak-*-latest as responses mode to match their Responses-only endpoints
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 21:41:12 +00:00
Mateo Wang
29ea2bd2cd
Merge pull request #39426 from BerriAI/litellm_azure_ai_grok_4_6_cost_map
...
feat(azure_ai): add grok-4.6 to the model cost map
2026-09-03 14:36:07 -07:00
mateo
00bdfe797a
fix(registry): mark gemini-3.5-live-translate-preview as realtime with official token limits
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 21:16:55 +00:00
mateo
5a3a2f3d0a
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 21:09:30 +00:00
mateo
eae7b806e3
fix(registry): point Bedrock Qwen3 Coder 480B source at the us-west-2 on-demand price list
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 20:49:44 +00:00
mateo
080e364d5e
fix(registry): carry Anthropic thinking/sampling flags on new Perplexity and OpenRouter Claude entries
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 20:28:48 +00:00
Mateo Wang
828d561fcc
Merge pull request #39622 from BerriAI/litellm_gpt_6_astra
...
feat(models): add gpt-6-astra pricing and metadata
2026-09-03 13:18:34 -07:00
mateo
2c4eb693ed
fix(model_prices): absorb Baseten GLM-5.3 and OpenRouter live prices, fix Bedrock Qwen3 Coder 480B input price and Gemini Live image price
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 19:53:58 +00:00
mateo-berri
4991d0bf3e
fix(models): match gpt-6-astra reasoning effort levels to OpenAI docs
...
OpenAI documents low, medium, high, xhigh, and max for gpt-6-astra, with no none level, so the entry stops advertising none and starts advertising max.
2026-09-03 12:47:25 -07:00
mateo-berri
897fba08c8
feat(models): add gpt-6-astra pricing and metadata
...
Adds the OpenAI gpt-6-astra entry to both price files with standard, flex, priority (fast mode), batch, and above-272K long-context rates, and regression tests covering each tier and the batch rates.
2026-09-03 12:47:25 -07:00
mateo
32a3a65312
fix(model_prices): add Lyria 3.5, Perplexity Agent API and OpenRouter first-party models, fix Nebius, Mistral, OpenRouter metadata
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 19:35:30 +00:00
mateo-berri
51d821ae45
fix(cost): bill bedrock_mantle web search at $12 per 1k queries using Bedrock's reported count
2026-09-03 12:21:54 -07:00
mateo
2c63095e83
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 19:02:19 +00:00
mateo-berri
a1e58aabe7
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_mantle_web_search
2026-09-03 09:50:42 -07:00
mateo
840173e778
feat(registry): add azure_ai/mistral-ocr-4-0 page and annotation prices from Azure Retail Prices
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:48:25 +00:00
mateo
f26407aa8c
feat(registry): add azure_ai/MAI-Thinking-1 from Azure Retail Prices and Foundry docs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:43:48 +00:00
mateo
55a5f142e6
fix(model_prices): add azure_ai Codestral-2501 and FW-Nemotron-Lightning-3.5, sync Azure and Vertex deprecation dates, fix novita gpt-oss vision flags
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:23:02 +00:00
mateo
1a39275cb3
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 13:05:42 +00:00
mateo-berri
c6b48af0d7
chore: merge litellm_internal_staging into litellm_azure_ai_grok_4_6_cost_map
2026-09-02 16:38:16 -07:00
mateo-berri
fe34124610
feat(azure_ai): add grok-4.6 to the model cost map
2026-09-02 16:03:45 -07:00
mateo
6fa02887c4
feat(model_prices): add meta/muse-spark-1.3 and its contributor tier
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 22:05:17 +00:00
mateo
e148868f0c
fix(model_prices): set watsonx max_output_tokens from IBM's documented maximum new tokens
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:38:45 +00:00
mateo
671559e591
fix(model_prices): set watsonx max_tokens equal to max_output_tokens per registry convention
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:35:29 +00:00
mateo
60ffde65e0
fix(model_prices): drop unpriced Volcengine Seed 2.1 entries, they would record zero spend
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:08:43 +00:00
mateo
9c5b20abdd
fix(model_prices): add Nebius, watsonx and Volcengine models and correct watsonx list prices
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:07:21 +00:00
mateo
7a8226e752
fix(model_prices): registry audit 2026-09-02, add claude-mythos-5-1 and gpt-daybreak aliases, fix gpt-5.5 Fast and W&B pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 19:40:53 +00:00
mateo-berri
f38a1ec129
Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_azure_deepseek_v4_flash_0731
...
# Conflicts:
# litellm/model_prices_and_context_window_backup.json
# model_prices_and_context_window.json
2026-09-02 11:43:22 -07:00
Mateo Wang
4049a075bd
Merge pull request #39170 from BerriAI/litellm_registry_audit_2026_09_01
...
fix(models): registry audit 2026-09-01: openai realtime and long-context tiers, mistral aliases, voyage, xai, fireworks, together, scaleway, azure ai, govcloud, azure gov, cloudflare whisper, deprecation dates
2026-09-02 11:02:43 -07:00
mateo-berri
a9d3a0746c
fix(models): price Azure DeepSeek V4 Flash 0731 from its own meters under the catalog id
2026-09-02 10:59:11 -07:00
mateo-berri
7a35c34303
fix(models): add the us-gov. geo inference profile keys for Claude Sonnet 5 and Opus 4.8
2026-09-02 10:30:14 -07:00
mateo
da23e0241d
fix(models): add cloudflare whisper transcription pricing and pin govcloud pricing tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 15:28:53 +00:00
Yujong Lee
07cf9dc46f
feat(models): add Azure DeepSeek V4 Flash 0731
2026-09-02 08:10:52 -07:00
mateo
b761277740
fix(models): drop inherited retirement dates from azure/us-gov entries pending a Government schedule source
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 15:09:49 +00:00
mateo-berri
6b83b16559
feat(gemini): day-0 pricing for gemini-3.8-flash
...
Gemini 3.8 Flash launches today with the same promotional pricing, limits,
and thinking settings as Gemini 3.7 Flash, so the gemini/, vertex_ai/, and
bare cost map entries mirror the 3.7 Flash ones. Regression tests lock the
launch prices, the 4096-token cache minimum, and the gemini-3 thought
signature gate in for the new model.
2026-09-02 08:04:14 -07:00
mateo
a7836ede15
fix(models): absorb open registry PRs: govcloud bedrock and mantle, azure gov, openai tiered long-context, scaleway, together qwen3.8, azure ai cache and kimi k2.7 code, azure mai deprecations
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 14:48:51 +00:00
mateo
23977fc290
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_01
2026-09-02 13:03:24 +00:00
James Liounis
e2c3f51c46
fix(search): forward search-tool params through the router, complete Parallel AI v1 param mapping ( #37883 )
...
* fix(search): forward search-tool params through the router, complete Parallel AI v1 param mapping
SearchAPIRouter dropped every parameter configured on a search tool, forwarding
only per-request kwargs. Any tool-level setting (mode, max_results, ...) was
silently lost on the way to the adapter, for every search provider.
Also completes the Parallel AI v1 search surface: after_date, fetch_policy,
location and include_domains now nest under advanced_settings instead of being
sent as unknown top-level fields, responses preserve search_id / session_id /
warnings / raw excerpts, and search cost is derived from the request mode and
the provider's reported usage rather than a single flat rate.
* fix(parallel_ai): stop a caller from pricing its own search request
`_parallel_ai_usage` carries the provider's reported usage into cost
calculation. It was only written when the response contained a usage block, so
a caller could pass `_parallel_ai_usage=[{"name": "sku_search", "count": 0}]`
and, whenever the provider omitted usage, bill $0.00 instead of $0.005 — the
value also reached the upstream request body as an unknown field.
The key is now stripped from inbound params and written unconditionally from
the parsed response, so only the provider can populate it.
* fix(parallel_ai): price fast search mode correctly
* test(parallel_ai): fake search at HTTP boundary
* fix(parallel_ai): tolerate null search result fields
---------
Co-authored-by: khushishelat <shelatkhushi@gmail.com>
2026-09-01 21:46:46 -07:00
mateo
d568bbe58d
fix(bedrock): use tool fallback without forced tool_choice for claude-fable-5-1 structured output
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:19:46 +00:00
mateo
3e3e4d6970
fix(anthropic): use native structured output for claude-fable-5-1 on Vertex AI and Bedrock Invoke
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:01:03 +00:00
Devin AI
818fc5b913
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_01
2026-09-01 19:40:50 +00:00
Devin AI
24ee419c85
fix(models): registry audit 2026-09-01 for openai realtime, mistral aliases, voyage, xai, fireworks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:28:04 +00:00