Mateo Wang
0318b4acdc
Merge pull request #30856 from emerzon/litellm_vertex_lyria_models
...
feat(vertex): add Lyria model support
2026-09-05 23:12:25 -07:00
mateo-berri
11e45ad953
fix(vertex_ai): mark the Lyria 3 catalog entries text-only
...
`vertex_ai/lyria-3-clip-preview` and `vertex_ai/lyria-3-pro-preview` were
registered with `supports_vision`, `supports_image_input`, and an `image`
modality, which contradicts their `gemini/lyria-3-*` siblings and makes
/model/info advertise image input on text-to-music models.
2026-09-05 23:00:01 -07:00
mateo-berri
6be78fa850
fix(vertex_ai): bill Lyria per generation, not per audio second
...
Google prices Lyria per generated clip, so every Vertex Lyria entry in the
price map now carries a single output_cost_per_image and both the speech
and the passthrough cost paths read that one field. The old
output_cost_per_second and audio_seconds_per_prediction pair assumed a
30 second clip, which does not match the 32.768 second WAV Vertex returns,
and no other model in the map priced audio that way
Drops max_audio_length_hours and max_audio_per_prompt from the price map,
its schema, the generator, and ModelInfo, since nothing reads them, and
drops the audio_mime_type hidden param for the same reason: the response
already carries the resolved content type on its own header
Folds the per-model bundled catalog lookups into one cached parse of the
local cost map, validated with a TypeAdapter over a ReadOnly TypedDict
2026-09-05 22:34:31 -07:00
Mateo Wang
56a61cf016
Merge pull request #39764 from BerriAI/litellm_govcloud_profiles_lit6421
...
feat(pricing): add GovCloud pricing for every live but unpriced Bedrock model
2026-09-05 17:15:22 -07:00
mateo-berri
8426235290
feat(ocr): add Cohere Parse support for cohere and azure_ai
2026-09-04 21:25:46 -07:00
mateo-berri
036d104533
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_govcloud_profiles_lit6421
2026-09-04 20:08:16 -07:00
mateo-berri
51514b9123
fix(cost-map): azure/gpt-6-astra accepts reasoning_effort none on Foundry
2026-09-04 17:29:55 -07:00
mateo-berri
3202963f25
feat(cost-map): add azure/gpt-6-astra and azure/us/gpt-6-astra Foundry pricing
2026-09-04 16:57:40 -07:00
mateo
50d6b26a86
fix(registry): mark baseten GLM-5.3 as vision-capable per Baseten vision docs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 20:10:58 +00:00
mateo
1f0611a8b9
fix(registry): drop Together MiniMax M2.7 and revert Qwen2.5 7B Turbo pricing, both non-serverless
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:25:10 +00:00
mateo
7b96a11e5f
feat(registry): add OpenRouter catalog gaps, Fireworks DeepSeek V4 Flash Vision, Together MiniMax M2.7 and Qwen2.5 7B Turbo pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:19:08 +00:00
mateo-berri
bef3585d82
feat(pricing): add GovCloud rows for every live but unpriced Bedrock model
...
Every model bedrock list-foundation-models and list-inference-profiles
report as live in us-gov-west-1 or us-gov-east-1 now has a priced row:
Claude Fable 5.1 (profile plus in-region), Nemotron Nano 9B (profile plus
in-region), Grok 4.6 (profile plus Mantle in both regions), the us-gov.
Claude 3 Haiku profile in the east, Nova Lite, Micro and the Nova 2
multimodal embeddings in the west, and the Gemma 4 and gpt-oss Mantle
SKUs the GovCloud offer files price. Offer-file rates are used where AWS
publishes them; Claude rows carry the 1.2x GovCloud premium.
2026-09-04 10:03:43 -07:00
mateo
0c29f510bc
fix(registry): drop gpt-image-2 text output price, add openrouter minimax-m3 and qwen3.7-plus
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 15:03:29 +00:00
mateo
c707f2fe5d
fix(model_prices): databricks gpt-5-3-codex is served via the Responses API
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:32:11 +00:00
mateo
8e83d6d63d
fix(model_prices): add Databricks Sep-2026 catalog, Azure gpt-realtime-2.x, per-token realtime image pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 14:06:54 +00:00
mateo
08bfdadb10
chore: merge litellm_internal_staging into litellm_registry_audit_2026_09_02, drop the deleted ocr ledger
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 04:18:09 +00:00
mateo-berri
43f31b0a4b
feat(pricing): add GovCloud Claude Opus 5 and us-gov. inference profile rows
2026-09-03 18:40:59 -07:00
mateo
f292667601
fix(registry): mark gpt-daybreak-*-latest as responses mode to match their Responses-only endpoints
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 21:41:12 +00:00
Mateo Wang
29ea2bd2cd
Merge pull request #39426 from BerriAI/litellm_azure_ai_grok_4_6_cost_map
...
feat(azure_ai): add grok-4.6 to the model cost map
2026-09-03 14:36:07 -07:00
mateo
00bdfe797a
fix(registry): mark gemini-3.5-live-translate-preview as realtime with official token limits
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 21:16:55 +00:00
mateo
5a3a2f3d0a
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 21:09:30 +00:00
mateo
eae7b806e3
fix(registry): point Bedrock Qwen3 Coder 480B source at the us-west-2 on-demand price list
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 20:49:44 +00:00
mateo
080e364d5e
fix(registry): carry Anthropic thinking/sampling flags on new Perplexity and OpenRouter Claude entries
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 20:28:48 +00:00
Mateo Wang
828d561fcc
Merge pull request #39622 from BerriAI/litellm_gpt_6_astra
...
feat(models): add gpt-6-astra pricing and metadata
2026-09-03 13:18:34 -07:00
mateo
2c4eb693ed
fix(model_prices): absorb Baseten GLM-5.3 and OpenRouter live prices, fix Bedrock Qwen3 Coder 480B input price and Gemini Live image price
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 19:53:58 +00:00
mateo-berri
4991d0bf3e
fix(models): match gpt-6-astra reasoning effort levels to OpenAI docs
...
OpenAI documents low, medium, high, xhigh, and max for gpt-6-astra, with no none level, so the entry stops advertising none and starts advertising max.
2026-09-03 12:47:25 -07:00
mateo-berri
897fba08c8
feat(models): add gpt-6-astra pricing and metadata
...
Adds the OpenAI gpt-6-astra entry to both price files with standard, flex, priority (fast mode), batch, and above-272K long-context rates, and regression tests covering each tier and the batch rates.
2026-09-03 12:47:25 -07:00
mateo
32a3a65312
fix(model_prices): add Lyria 3.5, Perplexity Agent API and OpenRouter first-party models, fix Nebius, Mistral, OpenRouter metadata
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 19:35:30 +00:00
mateo-berri
51d821ae45
fix(cost): bill bedrock_mantle web search at $12 per 1k queries using Bedrock's reported count
2026-09-03 12:21:54 -07:00
mateo
2c63095e83
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 19:02:19 +00:00
mateo-berri
a1e58aabe7
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_mantle_web_search
2026-09-03 09:50:42 -07:00
mateo
840173e778
feat(registry): add azure_ai/mistral-ocr-4-0 page and annotation prices from Azure Retail Prices
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:48:25 +00:00
mateo
f26407aa8c
feat(registry): add azure_ai/MAI-Thinking-1 from Azure Retail Prices and Foundry docs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:43:48 +00:00
mateo
55a5f142e6
fix(model_prices): add azure_ai Codestral-2501 and FW-Nemotron-Lightning-3.5, sync Azure and Vertex deprecation dates, fix novita gpt-oss vision flags
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 13:23:02 +00:00
mateo
1a39275cb3
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_02
2026-09-03 13:05:42 +00:00
Emerson Gomes
f18cb0cdb4
fix(vertex): make Lyria routing and billing data-driven
2026-09-02 19:22:42 -05:00
Emerson Gomes
b96844dd0c
feat(vertex): expose Lyria through audio speech
2026-09-02 19:22:15 -05:00
Emerson Gomes
514e9a1ee6
fix(vertex): address lyria review feedback
2026-09-02 19:19:53 -05:00
Emerson Gomes
3ead9d1688
feat(vertex): add Lyria model support
2026-09-02 19:19:52 -05:00
mateo-berri
c6b48af0d7
chore: merge litellm_internal_staging into litellm_azure_ai_grok_4_6_cost_map
2026-09-02 16:38:16 -07:00
mateo-berri
fe34124610
feat(azure_ai): add grok-4.6 to the model cost map
2026-09-02 16:03:45 -07:00
mateo
6fa02887c4
feat(model_prices): add meta/muse-spark-1.3 and its contributor tier
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 22:05:17 +00:00
mateo
e148868f0c
fix(model_prices): set watsonx max_output_tokens from IBM's documented maximum new tokens
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:38:45 +00:00
mateo
671559e591
fix(model_prices): set watsonx max_tokens equal to max_output_tokens per registry convention
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:35:29 +00:00
mateo
60ffde65e0
fix(model_prices): drop unpriced Volcengine Seed 2.1 entries, they would record zero spend
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:08:43 +00:00
mateo
9c5b20abdd
fix(model_prices): add Nebius, watsonx and Volcengine models and correct watsonx list prices
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 20:07:21 +00:00
mateo
7a8226e752
fix(model_prices): registry audit 2026-09-02, add claude-mythos-5-1 and gpt-daybreak aliases, fix gpt-5.5 Fast and W&B pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 19:40:53 +00:00
mateo-berri
f38a1ec129
Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_azure_deepseek_v4_flash_0731
...
# Conflicts:
# litellm/model_prices_and_context_window_backup.json
# model_prices_and_context_window.json
2026-09-02 11:43:22 -07:00
Mateo Wang
4049a075bd
Merge pull request #39170 from BerriAI/litellm_registry_audit_2026_09_01
...
fix(models): registry audit 2026-09-01: openai realtime and long-context tiers, mistral aliases, voyage, xai, fireworks, together, scaleway, azure ai, govcloud, azure gov, cloudflare whisper, deprecation dates
2026-09-02 11:02:43 -07:00
mateo-berri
a9d3a0746c
fix(models): price Azure DeepSeek V4 Flash 0731 from its own meters under the catalog id
2026-09-02 10:59:11 -07:00