kerry
ce83fac351
fix(cost): bill batch embeddings per modality token rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:50:21 +00:00
kerry
a28ea22ec1
fix(cost): move gemini-embedding-2-preview to per-token rates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:34:22 +00:00
kerry
4a8ec7b9d8
fix(cost): bill gemini-embedding-2-preview per token like the GA entries
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:18:50 +00:00
kerry
d4f2119b03
fix(cost): bill gemini-embedding-2 per token and stop double charging audio
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:41:46 +00:00
berriai-litellm-provider-info-sync[bot]
c3f8c07c43
chore(prices): sync Vertex AI prices: 14 models
...
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost
vertex_ai/gemini-2.5-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost
vertex_ai/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority
vertex_ai/gemini-3.1-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex
vertex_ai/gemini-3.1-flash-lite: cache_read_input_audio_token_cost
vertex_ai/gemini-3.1-flash-lite-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
vertex_ai/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
vertex_ai/gemini-3.5-flash:
vertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_priority
vertex_ai/gemini-3.6-flash:
vertex_ai/gemini-3.7-flash:
vertex_ai/gemini-3.8-flash:
vertex_ai/gemini-embedding-2:
2026-09-14 22:00:57 +00:00
Devin AI
e4ebeae800
fix(prices): add text output rate to Vertex TTS entries, sync gemini-embedding-2 alias
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:49:15 +00:00
Devin AI
f45e20e6c2
fix(prices): mark Vertex Gemini TTS entries as audio_speech
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:41:43 +00:00
kerry
abbf5aae20
chore(prices): flag gpt-5.5-cyber as a reasoning model
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:41:10 +00:00
berriai-litellm-provider-info-sync[bot]
93e65995d6
chore(prices): sync Vertex AI prices: 4 models
...
gemini-2.5-flash: cache_read_input_audio_token_cost
gemini-2.5-flash-lite: cache_read_input_audio_token_cost
gemini-3-flash-preview: cache_read_input_audio_token_cost
gemini-3.1-flash-lite: cache_read_input_audio_token_cost
2026-09-14 21:26:02 +00:00
kerry
ddf13e9505
Merge remote-tracking branch 'origin/main' into litellm-providers/price-sync
2026-09-14 21:21:55 +00:00
kerry
250ff03a04
feat(model_info): scope fill_missing rules to azure, bedrock and vertex hosts
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
kerry
812bbee0b3
fix(model_info): scope fill_missing backfill to the rule's providers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
fb2057fde7
refactor(model_info): rename backfill_exact_entries to fill_missing_fields
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
42c708670d
fix(model_info): guard backfill by mode, drop provider key, tighten claude major regex
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Devin AI
cee7215b24
feat(model_info): opt-in field-level backfill from fallback generalization rules
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 18:58:53 +00:00
Mateo Wang
cab1e113f7
Merge pull request #40976 from BerriAI/litellm_azure_gpt_chat_latest_pricing
...
feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries
2026-09-14 11:21:20 -07:00
mateo
1661e72c2c
fix(registry): drop retired friendliai llama-3.1 serverless models
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 13:21:46 +00:00
mateo-berri
4c022a3089
feat(pricing): add azure gpt-chat-latest global and data zone rates
2026-09-12 23:34:12 -07:00
berriai-litellm-provider-info-sync[bot]
6423acc11a
chore(prices): sync prices for 5 providers: 278 models, 34 new
...
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp:
fireworks_ai/deepseek-v4-flash-vision-exp:
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/muse-glimmer-30b:
fireworks_ai/muse-glimmer-30b:
fireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4:
fireworks_ai/nemotron-3-ultra-nvfp4:
fireworks_ai/accounts/fireworks/models/qwen3-embedding-8b:
fireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token
fireworks_ai/accounts/fireworks/models/qwen3p7-plus:
fireworks_ai/qwen3p7-plus:
fireworks_ai/accounts/fireworks/models/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/routers/glm-5p2-fast:
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
fireworks_ai/accounts/fireworks/routers/kimi-k3-fast:
together_ai/arcee-ai/trinity-mini: input_cost_per_token, output_cost_per_token
together_ai/arize-ai/qwen-2-1.5b-instruct:
babbage-002: input_cost_per_token_batches, output_cost_per_token_batches
chat-latest:
chatgpt-image-latest: output_cost_per_token, input_cost_per_image_token, output_cost_per_image_token, input_cost_per_token_batches, output_cost_per_token_batches
claude-fable-5:
claude-fable-5-1:
claude-haiku-4-5:
claude-mythos-5:
claude-mythos-5-1:
claude-opus-4-5:
claude-opus-4-6:
claude-opus-4-7:
claude-opus-4-8:
claude-opus-5:
claude-sonnet-4-5:
claude-sonnet-4-6:
claude-sonnet-5:
davinci-002: input_cost_per_token_batches, output_cost_per_token_batches
deep-research-pro-preview-12-2025: cache_read_input_token_cost
together_ai/deepseek-ai/deepseek-coder-33b-instruct: input_cost_per_token, output_cost_per_token
together_ai/deepseek-ai/DeepSeek-R1-0528:
2026-09-13 04:55:39 +00:00
shivam
a28e595a9d
Merge remote-tracking branch 'origin/main' into litellm_fix_realtime_cached_audio_cost
2026-09-13 04:24:18 +00:00
Mateo Wang
b1a61f510c
Merge pull request #35918 from Lee-Si-Yoon/feat/friendli-model-metadata-sync
...
feat(friendli): auto-sync Friendli model metadata into price registry
2026-09-12 21:13:52 -07:00
mateo-berri
a94c060b84
fix(cost): fill the missing realtime cache-read rates
2026-09-12 17:48:00 -07:00
Mateo Wang
9d984371fd
Merge pull request #40909 from BerriAI/litellm_databricks_reasoning_effort_thinking
...
fix(databricks): translate reasoning_effort to thinking for Gemini 2.5
2026-09-12 17:46:44 -07:00
kerry-berri
d565031b9c
Merge pull request #40902 from BerriAI/litellm_openai_reasoning_fallback_rule
...
feat(registry): add openai reasoning-family fallback generalization
2026-09-12 16:27:16 -07:00
shivam
3aeae3c7fe
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_realtime_cached_audio_cost
...
ai-gateway image / ai-gateway release image (push) Waiting to run
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# tests/test_litellm/test_cost_calculator.py
2026-09-12 22:59:49 +00:00
ryan-crabbe-berri
c134fb7a38
Merge pull request #39395 from seyeong-han/litellm_meta_muse_voice_realtime
...
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
feat(realtime): add Meta Muse Voice transcription
2026-09-12 15:58:24 -07:00
kerry
df272d7e2f
fix(registry): limit reasoning fallback to single-digit gpt majors
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 22:01:58 +00:00
kerry
543ed2f6da
fix(registry): scope codex/deep-research/chat-latest markers to gpt bases
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:23:24 +00:00
kerry
db6b851884
feat(registry): add openai reasoning-family fallback generalization
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:14:24 +00:00
mateo-berri
c7b607c46e
fix(databricks): keep the Claude fallback when gating the anthropic thinking payload
...
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Gate the reasoning_effort translation on the cost-map flag or the model name containing
claude, so unmapped Claude serving endpoints keep translating. Flag the newer Claude
entries that were missing it. Expose supports_anthropic_thinking_payload as a public
helper next to the other supports_* wrappers instead of importing the private factory.
Drop the adaptive-only guard, since the adaptive flags only ever match Claude ids, and
add regression tests for an unmapped Claude endpoint and an adaptive Claude model
2026-09-12 13:13:38 -07:00
mateo-berri
10a0da7a32
Merge litellm_internal_staging into devin/1784568628-databricks-gemini-reasoning-effort
2026-09-12 13:08:22 -07:00
mateo-berri
9fd1ef01fe
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_realtime_cached_audio_cost
...
# Conflicts:
# litellm/responses/litellm_completion_transformation/transformation.py
2026-09-12 12:46:37 -07:00
mateo
414442cd06
fix(registry): mark computer-use-preview as supporting pdf input
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:46:13 +00:00
mateo
9aa06cb4a3
fix(registry): correct computer-use-preview provider/schema flag and OpenRouter deepseek-v3.2 / claude-opus-4.6 metadata
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:20:05 +00:00
mateo
db8201df8f
fix(registry): sync azure/us and azure/eu o-series deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 14:44:52 +00:00
mateo
a6418b3ff5
fix(registry): keep gpt-5.4-mini/nano max_input_tokens at 272000
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 13:45:15 +00:00
mateo
6c974df9a0
fix(registry): sync Azure o-series and Together deprecation dates, gpt-5.4-mini/nano context window
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 13:24:00 +00:00
ryan-crabbe-berri
17fde7a261
refactor(realtime): move Meta Muse Voice onto BaseRealtimeConfig
...
Replace the hand-rolled Meta realtime handler with a MetaRealtimeConfig
that plugs into the shared realtime handler and RealTimeStreaming relay.
Clients keep speaking the OpenAI realtime wire: session.update,
input_audio_buffer.append/commit and the OpenAI transcription events.
Unsupported transcription settings are logged and dropped, matching the
Gemini realtime precedent, and the Meta-specific session.mode, keywords,
language_bias, DIARIZATION and speaker extensions are removed.
Drop the MODEL_API_KEY env var in favor of the standard META_API_KEY,
remove the private-logging flag so spend logs record the transcript the
same way other realtime models do, and add per-second pricing for
muse-voice-transcribe-1.0.
The relay now sends raw bytes from transform_realtime_request straight to
the backend after pace_backend_send, and transcription sessions never
trigger response.create.
2026-09-11 20:12:15 -07:00
Young Han
b82b31a44f
feat(realtime): add Meta Muse Voice transcription
2026-09-11 20:10:53 -07:00
mateo
6cffb31e5c
feat(model_prices): add DeepSeek V4.1 Flash on Fireworks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 00:36:20 +00:00
mateo-berri
71f45683d7
fix(cost-map): keep minimal withheld on Bedrock gpt-5.4 and gpt-5.5
...
LiteLLM sends the Bedrock Mantle GPT rows through Bedrock's Responses endpoint, which refuses minimal on gpt-5.4 and gpt-5.5 like every other Bedrock GPT row. The earlier commit measured the raw chat endpoint, which accepts it, and dropped the flag by mistake. The ladder test now matches what the proxy path returns
2026-09-11 13:46:28 -07:00
mateo-berri
dbc57c13d4
fix(cost-map): match Bedrock GPT effort flags to what Bedrock accepts
...
Live calls to Bedrock Mantle and Converse on 2026-09-11: the gpt-5.6 luna, sol, and terra rows and gpt-6-astra return 200 on reasoning_effort=max, gpt-6-astra returns 400 on none, and Mantle gpt-5.4 and gpt-5.5 return 200 on minimal. The commercial Bedrock rows now carry exactly those flags, and the schema test asserts the measured ladder per row instead of a blanket mirror of the direct OpenAI rows
2026-09-11 13:37:50 -07:00
mateo
4ffd4ecc83
fix(model_prices): absorb cerebras/inception PRs, fix vertex/openai/together/openrouter pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 19:37:45 +00:00
Devin AI
4422be0f28
fix(cost-map): declare minimal unsupported on bedrock-hosted openai gpt rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 15:52:43 +00:00
Devin AI
43b56e8707
fix(cost-map): advertise xhigh reasoning effort on bedrock-hosted openai gpt rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 15:36:39 +00:00
mateo
4382b86b0f
registry audit 2026-09-11: xai/groq deprecation dates, deepseek-v4-flash vision, perplexity nemotron reasoning
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 13:14:37 +00:00
mateo
598e863510
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_10b
2026-09-11 13:04:19 +00:00
ryan-crabbe-berri
033f2e5e2a
feat(wandb): default unmapped W&B models to reasoning-capable
...
W&B's serverless catalog grows faster than the registry names it, so a model
they ship today resolves as non-reasoning here until someone edits the cost map,
and the caller's reasoning_effort is dropped or rejected.
Add a wandb-reasoning-baseline capability rule to fallback_generalizations so any
wandb/ id the map has not described defaults to supports_reasoning. Rules lose to
exact entries, so mapped non-reasoning models such as
wandb/meta-llama/Llama-3.1-8B-Instruct are unaffected.
The rule carries no mode and no pricing, so cost stays on the standard unpriced
behavior and the deployment does not read as catalog-mapped to the router's
reasoning-effort resolver.
Claude-Session: https://claude.ai/code/session_01A6SkwJdfZUmkzfUkrEkqX8
2026-09-10 16:40:07 -07:00
shivam
302a8d43da
fix(cost): bill cached realtime audio tokens at the audio cache-read rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 22:02:58 +00:00
ryan-crabbe-berri
2f114d44ed
Merge pull request #39190 from WolframRavenwolf/litellm_wandb_reasoning_effort
...
fix(wandb): preserve reasoning_effort in chat completions
2026-09-10 14:02:49 -07:00