mateo
d568bbe58d
fix(bedrock): use tool fallback without forced tool_choice for claude-fable-5-1 structured output
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:19:46 +00:00
mateo
3e3e4d6970
fix(anthropic): use native structured output for claude-fable-5-1 on Vertex AI and Bedrock Invoke
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 20:01:03 +00:00
mateo
d6005a1876
merge: resolve conflict with litellm_internal_staging in anthropic transformation tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 19:23:47 +00:00
Mateo Wang
c50d83ece2
Merge pull request #39070 from BerriAI/litellm_bedrock_invoke_native_structured_output
...
fix(bedrock): forward native structured outputs on Invoke instead of silently inlining the schema
2026-09-01 12:19:15 -07:00
Mateo Wang
435433fa07
Merge pull request #39149 from BerriAI/litellm_qwencloud_provider_aliases
...
feat(dashscope): add QwenCloud and Qwen AI Platform provider aliases
2026-09-01 12:18:05 -07:00
Mateo Wang
4c3ef9ae0a
Merge pull request #39023 from BerriAI/litellm_add_azure_deepseek_v4_flash_0731
...
feat: add Azure AI DeepSeek V4 Flash 0731 pricing
2026-09-01 11:54:25 -07:00
mateo-berri
0042493bca
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_invoke_native_structured_output
2026-09-01 11:50:05 -07:00
mateo
3c9ce458fd
feat(anthropic): gate forced tool_choice for Fable 5.1 behind supports_forced_tool_use
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:46:03 +00:00
mateo-berri
f3792fb700
feat(dashscope): add qwencloud and qwen_ai_platform provider aliases
2026-09-01 11:20:36 -07:00
mateo
fb93db7791
feat(models): add Claude Fable 5.1 across Anthropic, Bedrock, Vertex AI, and Azure AI
...
Adds claude-fable-5-1 cost map entries on the Anthropic API, Bedrock converse
(base, global, and us/eu geo inference profiles at the 10% regional premium),
Vertex AI, and Azure AI. Specs match Fable 5 (1M context, 128K output, $10/$50
per MTok, adaptive thinking always on, xhigh and max effort), except cache reads
land at $0.25 per MTok, a quarter of Fable 5's price and 0.025x base input
instead of the usual 0.1x.
Registers anthropic.claude-fable-5-1 in BEDROCK_CONVERSE_MODELS, lists the model
in the setup wizard, and extends the reasoning effort e2e grid. The partner cells
carry fail_reason markers until access on the CI accounts is confirmed.
Partner entries deliberately carry no deprecation_date: Anthropic publishes
retirement no sooner than 2027-09-01 for the first-party model, and the Foundry
and Vertex dates are not published yet.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:07:06 +00:00
mateo-berri
06d4521fc0
fix(registry): add vertex veo 3.1 resolution tier pricing per vertex pricing page
2026-09-01 10:15:18 -07:00
Devin AI
9a1aebc146
fix(registry): declare databricks deepseek cache-write rate at the input rate per repo convention
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 13:50:21 +00:00
Devin AI
5263570e68
Add cerebras/zai-glm-4.7 deprecation_date per Cerebras deprecations page
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 13:33:31 +00:00
Devin AI
7bfa0d7fb4
Registry audit: Fireworks DeepSeek V4 Flash 0731 pricing, Databricks DeepSeek V4 entries, provider deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 13:15:22 +00:00
Devin AI
40738355b4
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_1788201394-veo31-pricing-tiers
2026-09-01 13:03:19 +00:00
Tin Chi Lo
a27e12367e
fix(bedrock): forward native structured outputs on Invoke instead of silently inlining the schema
2026-08-31 23:19:22 -07:00
Mateo Wang
7a02e4163f
Merge pull request #38913 from BerriAI/litellm_gigachat_passthrough_25886
...
feat(gigachat): add native API passthrough routes with spend logging
2026-08-31 15:47:27 -07:00
mateo-berri
eb00986f18
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
...
# Conflicts:
# osv-scanner.toml
2026-08-31 15:25:10 -07:00
Mateo Wang
ac206518a0
Merge pull request #38881 from Lee-Si-Yoon/friendli/glm-5.3
...
feat(friendli): add zai-org/GLM-5.3 model pricing
2026-08-31 15:18:45 -07:00
Mateo Wang
336269cfec
Merge pull request #38597 from BerriAI/litellm_fix_nova_sonic_realtime_user_asr_usage
...
fix(bedrock): surface Nova Sonic user transcripts, speech events, and usage in realtime API
2026-08-31 15:11:31 -07:00
mateo-berri
abbccd3fd6
Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into friendli/glm-5.3
...
# Conflicts:
# litellm/model_prices_and_context_window_backup.json
# model_prices_and_context_window.json
2026-08-31 15:11:06 -07:00
Mateo Wang
b3a1dd1115
Merge pull request #38880 from Lee-Si-Yoon/friendli/glm-5.3-flash-v2
...
feat(friendli): add zai-org/GLM-5.3-Flash model pricing
2026-08-31 15:07:06 -07:00
Yujong Lee
be5997f366
feat: add Azure AI DeepSeek V4 Flash 0731 pricing
2026-08-31 14:49:43 -07:00
Devin AI
6809d537f0
merge: litellm_internal_staging into litellm_fix_nova_sonic_realtime_user_asr_usage
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 21:42:06 +00:00
mateo-berri
bd794f9f18
fix(friendli): track GLM-5.3 discounted live pricing and declare effort levels
2026-08-31 13:26:37 -07:00
mateo-berri
a90fb538bf
fix(friendli): declare GLM-5.3-Flash reasoning efforts as explicit levels
2026-08-31 13:25:51 -07:00
mateo-berri
59732f068b
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
2026-08-31 13:16:41 -07:00
Devin AI
c344c7a66b
fix(registry): add zai/glm-5.2, together Qwen3.8-Flash, cerebras/gemma-4-31b, elevenlabs/scribe_v2
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 19:11:04 +00:00
Devin AI
7b4b92f54f
fix(registry): update veo 3.1 pricing with resolution tiers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 18:37:26 +00:00
mateo-berri
8a6f47a6d4
merge: litellm_internal_staging into litellm_fix_nova_sonic_realtime_user_asr_usage
2026-08-31 10:17:12 -07:00
Devin AI
4291afbfa5
fix(registry): correct OpenAI preview shutdown dates, add whisper/transcribe and Bedrock/Vertex deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 13:08:06 +00:00
Devin AI
f111262e54
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_1788036971-stale-cost-map-sources
2026-08-31 13:03:35 +00:00
siyoon
e7bfe99cd3
feat(friendli): add zai-org/GLM-5.3 model pricing
...
Per https://api.friendli.ai/serverless/v1/models :
- $1.40 input / $4.40 output / $0.26 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
(per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- text-only (no vision), flagship GLM model
2026-08-30 14:23:36 +09:00
siyoon
3ea4b715ba
feat(friendli): add zai-org/GLM-5.3-Flash model pricing
...
Per https://api.friendli.ai/serverless/v1/models :
- $0.15 input / $0.50 output / $0.03 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
(per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- image + video input (native multimodal)
2026-08-30 14:22:41 +09:00
mateo-berri
70e2f4e68f
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
...
# Conflicts:
# litellm/llms/gigachat/chat/transformation.py
2026-08-29 22:08:54 -07:00
Mateo Wang
42d8360f29
Merge pull request #38820 from BerriAI/litellm_fix_together_sync_output_ceiling
...
fix(together_ai): stop writing context_length as max_output_tokens in the serverless sync
2026-08-29 16:44:56 -07:00
Mateo Wang
a979c89b88
Merge pull request #38804 from BerriAI/litellm_registry_audit_rolling_38693
...
fix(models): registry audit: new Together/Fireworks/Gemini/Mistral/xAI models, xai retirement repricing, bedrock grok-4.6 caching, deprecation dates
2026-08-29 16:44:45 -07:00
Mateo Wang
6bc8dafa99
Merge pull request #38740 from BerriAI/litellm_vertex_gemini_35_transcribe
...
feat(vertex_ai): support gemini-3.5-transcribe on /v1/audio/transcriptions
2026-08-29 16:19:56 -07:00
mateo-berri
af179be681
fix(together_ai): stop writing context_length as max_output_tokens in the serverless sync
...
The Together catalog exposes only context_length, so the sync was recording
every chat model's context window as its output ceiling. New entries now carry
max_input_tokens and the legacy max_tokens from the catalog and get an output
ceiling only from a reviewed capability rule. GLM-5.2 and GLM-5.3-Flash rules
carry the documented 128K ceiling, and the 26 other inflated together_ai chat
entries drop max_output_tokens in both registry copies.
2026-08-29 15:26:44 -07:00
mateo-berri
cfb7a26327
fix(registry): add gemma 4 capability flags verified against the gemini api
2026-08-29 14:38:21 -07:00
mateo-berri
68404d8ff4
fix(registry): drop xai video entries that have no video adapter
2026-08-29 14:20:48 -07:00
Devin AI
5024c4d520
fix: update stale source URLs in model cost map
...
119 entries pointed at 404ing or permanently-moved pages (Pylon #7777 ).
Replaced with verified working equivalents (200-checked or permanent
redirect targets).
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 20:56:21 +00:00
mateo-berri
2bd7b58640
fix(registry): correct xai retired slug pricing, bedrock grok caching, and unsourced entries
...
Reprice ten more retired xAI slugs (grok-3 and grok-3-mini families,
grok-4-1-fast) to the grok-4.3 rates they now bill at, with family-correct
deprecation dates. Restore cache_read_input_token_cost on the Bedrock Grok 4.6
entries so implicit cache hits bill at the cache-read rate while explicit
cachePoint stays unsupported. Drop the unsourced 1080p video rate and the
gemini/ live native-audio entry the Gemini API 404s on. Add Groq qwen3.8-27b
tool-use flags per Groq docs. Extend the xai and gemini tests to lock all of
this in
2026-08-29 13:24:09 -07:00
Devin AI
d77b4be31d
fix(models): align GLM-5.3 max_tokens with max_output_tokens
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 19:21:51 +00:00
Devin AI
df3d37db2a
fix(models): add verified Gemini, Mistral, Fireworks, xAI registry entries
...
- gemini: nano-banana-pro-preview, gemma-4-26b-a4b-it, gemma-4-31b-it
- mistral: 14 official aliases from api.mistral.ai/v1/models
- fireworks_ai: glm-5p3, qwen3-embedding-8b
- xai: grok-imagine-video, grok-imagine-video-1.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 19:14:09 +00:00
mateo-berri
a007fa49e5
Merge branch 'litellm_internal_staging' into litellm_veo_31_lite
2026-08-29 12:04:42 -07:00
Devin AI
f0849eb0c9
fix(models): xai retirement repricing, bedrock grok-4.6 caching, openai/gemini deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 13:11:26 +00:00
Devin AI
401b12e64b
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin/1787944648-registry-audit-rolling
2026-08-29 13:02:37 +00:00
mateo-berri
1a26608769
feat(vertex_ai): route gemini transcribe models to generateContent on /v1/audio/transcriptions
2026-08-29 00:46:03 -07:00
Mateo Wang
27c09248e4
Merge pull request #38593 from BerriAI/litellm_gpt5_default_reasoning_effort
...
fix(gpt-5): stop forwarding temperature and top_p to reasoning models that reject them
2026-08-28 15:20:41 -07:00