Commit graph

31602 commits

Author SHA1 Message Date
Seven
53b2299016
Merge b95a9f1fe5 into 2dccc0dc79 2026-09-23 14:42:37 +00:00
devin-ai-integration[bot]
2dccc0dc79
feat(models): add openrouter/aion-labs/aion-3.5-mini pricing (#42743)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 07:39:34 -07:00
devin-ai-integration[bot]
5028f9ec59
fix(proxy): validate model credential name only when it changes (#42701)
PATCH /model/{id}/update rejected read-modify-write edits that resent an unchanged but dangling litellm_credential_name. Existence validation now runs only when the requested name differs from the stored one; empty string and non-admin detach rejections are unchanged

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 07:34:32 -07:00
devin-ai-integration[bot]
411fa04f86
fix(model_prices): add azure_ai gpt-image-2 and groq llama-guard-3-8b deprecation dates (#42738) 2026-09-23 07:30:17 -07:00
berriai-litellm-provider-info-sync[bot]
d525b0a8df
chore(prices): sync OpenRouter prices: 19 models, 9 new [18 held] (#42592)
* chore(prices): sync OpenRouter prices: 19 models, 9 new [18 held]

openrouter/~deepseek/deepseek-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~deepseek/deepseek-v4-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/~moonshotai/kimi-latest: input_cost_per_token, output_cost_per_token
openrouter/~z-ai/glm-flash-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/aion-labs/aion-2.0: max_input_tokens
openrouter/aion-labs/aion-3.0: max_input_tokens
openrouter/aion-labs/aion-3.0-mini: max_input_tokens
openrouter/anthropic/claude-opus-5.5:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, cache_creation_input_token_cost_above_1hr
openrouter/cohere/command-a-plus: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4.1-flash: max_tokens, max_output_tokens, off_peak_pricing
openrouter/deepseek/deepseek-v4.1-flash:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/openai/gpt-6-luna-pro:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-luna:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-sol-pro:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-6-sol:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, input_cost_per_token_above_272k_tokens, output_cost_per_token_above_272k_tokens, cache_read_input_token_cost_above_272k_tokens, cache_creation_input_token_cost_above_272k_tokens
openrouter/openai/gpt-oss-20b:batch: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3.8-omni-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost

* chore(prices): sync OpenRouter prices: 1 model [9 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [12 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [9 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [12 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [10 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* chore(prices): sync OpenRouter prices: 1 model [13 held]

openrouter/deepseek/deepseek-v4-pro-0813: off_peak_pricing

Price-Sync: litellm-providers

* feat(prices): add openrouter/upstage/solar-mini4 from OpenRouter models API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync OpenRouter prices: 1 model [16 held]

openrouter/deepseek/deepseek-v4.1-flash: off_peak_pricing

Price-Sync: litellm-providers

* feat(prices): add openrouter/aion-labs/aion-3.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 07:23:42 -07:00
devin-ai-integration[bot]
860bc7811d
refactor(types): replace Any with proven types in 5 files (#42722)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 03:14:45 -07:00
Oliver Jensen
bc3b5b1d5b
fix(proxy): revoke UI session tokens on logout and password change (#42463)
* fix(proxy): revoke UI session tokens on logout and password change

Adds POST /session/logout to revoke the presented UI session key server
side (previously logout was client-side only and the key stayed valid
until expiry). Password changes now revoke the user's other UI sessions:
self-change keeps the caller's session, admin reset and onboarding claim
revoke all. The BYOK OAuth cookie auth now re-resolves the embedded key
against the DB so revoked sessions get a 401.

* fix(proxy): satisfy B008 budget and backend allowlist for /session/logout

* refactor(proxy): satisfy type-discipline budget in session_endpoints
2026-09-23 10:31:38 +02:00
devin-ai-integration[bot]
a3196907e4
feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans (#42486)
* feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep the caller's header session under missing_session_id: generate and read replayed payload session ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): drop only the proxy-minted session id so a caller id on the other metadata key survives

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep a replayed session id hidden when it only echoes the payload trace id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep a replayed session id even when the payload trace id fell back to it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): stop reading the replayed payload's session id, the generated marker does not survive replay

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit gen_ai.conversation.id on otel v2 spans through a real proxy, sink and postgres

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep otel conversation rigs alive for the whole session so shuffled shards do not reboot the proxy per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): stop the audit rig proxies from probing sibling test peers for model info

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): record accepted OTLP batches in the sink instead of mutating the collector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): guard the accepted batch deque so snapshots cannot race sink appends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-23 00:49:29 -07:00
devin-ai-integration[bot]
40ec84caa2
fix(proxy): publish auth cache invalidations in the background so a wedged coordination Redis cannot stall user updates (#42534)
* fix(proxy): bound auth cache invalidation publish so a wedged coordination Redis cannot stall user updates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve publish callable at call time in evict_and_broadcast

The keyword-only default bound publish_auth_cache_invalidation at
function-definition time, so tests patching the module attribute observed
zero calls. Default to None, resolve the real publisher inside the body,
and keep the keyword-shaped cache_key call the existing contract asserts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): publish auth cache invalidations in the background so a wedged coordination Redis costs handlers nothing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): cap in-flight auth cache invalidation publishes so a wedge cannot drain the redis pool

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 23:33:32 -07:00
devin-ai-integration[bot]
320ad73f56
fix(policy_engine): keep inherited parent guardrails when a child policy condition misses (#42548)
* fix(policy_engine): keep inherited parent guardrails when a child policy condition misses

Attachment applicability now walks the policy inheritance chain, so an attached child whose own condition does not match still contributes the guardrails of its unconditional ancestors, and a non-default attachment that applies through an ancestor still suppresses default attachments. The resolver continues to skip only the chain members whose own condition fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(policy_engine): skip a policy's pipeline when its own condition misses

resolve_pipelines_for_context returned the pipeline of a matched policy without evaluating its own condition, so a condition-missing child admitted by the chain-aware matcher still ran its pipeline. It now mirrors resolve_policy_guardrails and drops the pipeline when the policy's own condition does not match.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(policy_engine): property test that chain matching only widens to applicable ancestors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(policy_engine): log policies admitted only through an inherited ancestor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(policy_engine): log ancestor admissions once per attachment scan

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 23:27:38 -07:00
berriai-litellm-provider-info-sync[bot]
721d39f476
chore(prices): sync Vertex AI prices: 1 model (#42680)
gemini-live-2.5-flash-native-audio:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 23:22:13 -07:00
berriai-litellm-provider-info-sync[bot]
b2789d6268
chore(prices): sync AWS Bedrock prices: 1 model (#42685)
us.mistral.pixtral-large-2502-v1:0:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 22:57:51 -07:00
devin-ai-integration[bot]
7172dfc400
fix(prices): align regional Bedrock Mistral Large 24.02 keys with the AWS pricing page (#42684)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 22:57:12 -07:00
berriai-litellm-provider-info-sync[bot]
3190d93136
chore(prices): sync AWS Bedrock prices: 11 models (#42677)
* chore(prices): sync AWS Bedrock prices: 11 models

deepseek.v3-v1:0: 
global.openai.gpt-5.6-luna: 
global.openai.gpt-5.6-sol: 
global.openai.gpt-5.6-terra: 
global.openai.gpt-6-astra: 
us.deepseek.r1-v1:0: 
us.openai.gpt-5.6-luna: 
us.openai.gpt-5.6-sol: 
us.openai.gpt-5.6-terra: 
us.openai.gpt-6-astra: 
writer.palmyra-vision-7b:

Price-Sync: litellm-providers

* fix(prices): align Bedrock Mistral Large 24.02 and Small 24.02 with the AWS pricing page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 4 models

mistral.mistral-7b-instruct-v0:2: 
mistral.mixtral-8x7b-instruct-v0:1: 
openai.gpt-oss-120b-1:0: 
openai.gpt-oss-20b-1:0:

Price-Sync: litellm-providers

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 22:39:12 -07:00
berriai-litellm-provider-info-sync[bot]
d0040196fe
chore(prices): sync AWS Bedrock prices: 2 models (#42673)
global.xai.grok-4.6: 
us.xai.grok-4.6:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 21:41:13 -07:00
devin-ai-integration[bot]
fdbd8382a4
fix(pricing): align bedrock_mantle/openai.gpt-daybreak-blue-5.6-sol with its Bedrock model card (#42672)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:38:41 -07:00
berriai-litellm-provider-info-sync[bot]
80221911cd
chore(prices): sync Vertex AI prices: 2 models, 2 new [enrichment failed: Vertex AI, 224 held] (#42589)
* chore(prices): sync Vertex AI prices: 2 models, 2 new [enrichment failed: Vertex AI, 224 held]

vertex_ai/gemini-2.0-flash: input_cost_per_token, output_cost_per_token, input_cost_per_character, input_cost_per_audio_token, input_cost_per_token_batches, output_cost_per_token_batches, input_cost_per_audio_token_batches
vertex_ai/gemini-2.0-flash-lite: input_cost_per_token, output_cost_per_token, input_cost_per_character, input_cost_per_audio_token, input_cost_per_token_batches, output_cost_per_token_batches, input_cost_per_audio_token_batches

* chore(prices): sync Vertex AI prices: 52 models

vertex_ai/claude-fable-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5-1: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-fable-5-1@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-haiku-4-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-haiku-4-5@20251001: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-5@20251101: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-6: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-6@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-7: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-7@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-8: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-4-8@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5-5: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-opus-5-5@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-4-5: cache_read_input_token_cost_batches
vertex_ai/claude-sonnet-4-5@20250929: cache_read_input_token_cost_batches
vertex_ai/claude-sonnet-4-6: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-4-6@default: input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_batches, cache_creation_input_token_cost_batches
vertex_ai/claude-sonnet-5: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/claude-sonnet-5@default: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/codestral-2: 
vertex_ai/codestral-2@001: 
vertex_ai/deepseek-ai/deepseek-ocr-maas: 
vertex_ai/deepseek-ai/deepseek-r1-0528-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/deepseek-ai/deepseek-v3.1-maas: cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/deepseek-ai/deepseek-v3.2-maas: cache_read_input_token_cost
vertex_ai/meta/llama-4-maverick-17b-128e-instruct-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/meta/llama-4-scout-17b-16e-instruct-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/minimaxai/minimax-m2-maas: cache_read_input_token_cost
vertex_ai/mistral-medium-3: 
vertex_ai/mistral-medium-3@001: 
vertex_ai/mistral-small-2503: 
vertex_ai/mistral-small-2503@001: 
vertex_ai/mistralai/codestral-2: 
vertex_ai/mistralai/codestral-2@001: 
vertex_ai/mistralai/mistral-medium-3: 
vertex_ai/mistralai/mistral-medium-3@001: 
vertex_ai/moonshotai/kimi-k2-thinking-maas: cache_read_input_token_cost
vertex_ai/openai/gpt-oss-120b-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/openai/gpt-oss-20b-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-235b-a22b-instruct-2507-maas: input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas: cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
vertex_ai/qwen/qwen3-next-80b-a3b-instruct-maas: 
vertex_ai/qwen/qwen3-next-80b-a3b-thinking-maas: 
vertex_ai/xai/grok-4.1-fast-non-reasoning: 
vertex_ai/xai/grok-4.1-fast-reasoning: 
vertex_ai/zai-org/glm-4.7-maas: cache_read_input_token_cost
vertex_ai/zai-org/glm-5-maas:

Price-Sync: litellm-providers

* chore(prices): add verified Vertex AI zai glm-5.2-maas entry and fix Gemini 2.0 mode

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prices): add Vertex AI claude-sonnet-4-5 1h cache write price above 200K

The Vertex AI pricing page prices Claude Sonnet 4.5's 1h Cache Write at
$6.00 up to 200K input tokens and $12.00 above, matching the value the
anthropic and bedrock entries already carry.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(prices): add vertex_ai/gemini-omni-1.1-flash-preview from the Vertex pricing page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:29:36 -07:00
devin-ai-integration[bot]
f44052d87b
fix(vector_stores): keep config-defined vector stores listed and read-only (#42574)
* fix(vector_stores): keep config-defined vector stores listed and read-only

Vector stores declared in config.yaml were purged from the in-memory registry by /vector_store/list because the database was treated as the only source of truth. Config-defined stores now carry is_config=True, stay in the list beside database rows, are never overwritten or evicted by database state, and reject /vector_store/new, /vector_store/update and /vector_store/delete with 400. The Admin UI renders them read-only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): show vector store source and read-only state for config-defined stores

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit config-owned vector stores across list, writes, search, authz, peers and redis outage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): show a visible read-only hint in the config vector store actions menu

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 04:02:16 +00:00
devin-ai-integration[bot]
3eb7e45615
fix(pricing): drop the unpublished cached rate from the Gemini Live preview entries (#42651)
* fix(pricing): correct cached-token fields on realtime cost-map entries

azure/gpt-realtime-2 was the only member of the gpt-realtime-2 family priced
on one side of its cached-audio meter. Azure publishes that meter as
"gpt-realtime-2 Audio cd inp Gl 1M Tokens" at 0.4 per 1M and charges the
same rate for the write that populates the cache and the read that hits it,
so cache_creation_input_audio_token_cost lands at 4e-07, matching
azure/gpt-realtime-2.1, azure/gpt-realtime-2.1-mini and the openai
gpt-realtime-2 entry. No cost path reads that field yet, so this corrects
what get_model_info reports rather than what anything bills.

The gemini Live entries go the other way. Google's Vertex context-caching
page publishes separate supported-model lists for implicit and explicit
caching, and no Live or native-audio model is in either one. Its pricing
page prints N/A in both cached-input columns for every Gemini 2.5 Flash
Live API row, where plain 2.5 Flash and 2.5 Flash-Lite both carry real
cached prices, and the Vertex model card for the family marks context
caching not supported outright. Vertex never reports cachedContentTokenCount
on a Live session either, including for a byte-identical 7,021-token prefix
replayed across sessions minutes apart, which is well past the 2,048-token
minimum the same page sets for the Gemini 2 family.

So the 7.5e-08 on the two preview siblings priced something the provider does
not sell, and supports_prompt_caching on all three claimed a capability the
model does not have. The rate comes out. The flag is set to false rather than
removed, because get_model_info maps an absent key to None, and None is how
this map spells "nobody checked" across the 2,788 entries that omit it, where
false records the vendor's documented no. Both readers of the flag gate on
`is True`, so nothing bills or behaves differently either way.

Only the cached fields change on the two 09-2025 preview entries. Their
source field points at the Gemini API pricing page rather than the Vertex
one, so they describe a different surface with its own published limits, and
their context windows are left alone rather than assumed to match the Vertex
model card that drives the GA entry.

Tests cover all three halves: the family invariant that a cached audio read
implies an equal cached audio write, a cached count on a Live entry leaving
the bill at the fresh-input total instead of adding the old 7.5e-08, and
supports_prompt_caching answering false for all three entries while still
answering true for 2.5 Flash, so the false cannot be a swallowed lookup
error.

* fix(cost): correct gemini-live-2.5-flash-native-audio limits and capabilities

Google's model card for model ID gemini-live-2.5-flash-native-audio gives a
128K context window and 64K maximum output tokens, and marks structured
output, context caching and URL context as not supported. Its modality list
is text in and out, image in, audio in and out, and video in, with no
document input of any kind.

The entry advertised a 1M context window, an off-by-one 65535 output cap, and
three capability flags the vendor marks unsupported. Context caching is the
fourth and is handled in the cached-fields change alongside its two preview
siblings.

Both the bare id and vertex_ai/gemini-live-2.5-flash-native-audio resolve to
this single entry, so the test drives the corrected values through both.

* test(integration): cover live preview cached tokens billed at the fresh rate

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): cite dated sources for Live entry pins and drop restating docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:35:34 -07:00
berriai-litellm-provider-info-sync[bot]
5dfaa8d620
chore(prices): sync AWS Bedrock prices and sources from the AWS price list (#42632)
* chore(prices): sync AWS Bedrock prices: 3 models [sync failed: AWS Bedrock]

anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema

* chore(prices): sync AWS Bedrock prices: 32 models

ai21.j2-mid-v1: 
ai21.j2-ultra-v1: 
ai21.jamba-1-5-large-v1:0: 
ai21.jamba-1-5-mini-v1:0: 
ai21.jamba-instruct-v1:0: 
au.anthropic.claude-opus-4-7: 
au.anthropic.claude-opus-4-8: 
au.anthropic.claude-opus-5: 
au.anthropic.claude-sonnet-4-6: 
au.anthropic.claude-sonnet-5: 
cohere.command-light-text-v14: 
cohere.command-text-v14: input_cost_per_token
cohere.embed-english-v3: 
cohere.embed-multilingual-v3: 
cohere.embed-v4:0: 
eu.anthropic.claude-fable-5: 
eu.anthropic.claude-opus-4-7: 
eu.anthropic.claude-opus-4-8: 
eu.anthropic.claude-opus-5: 
eu.anthropic.claude-sonnet-4-6: 
eu.anthropic.claude-sonnet-5: 
jp.anthropic.claude-opus-4-7: cache_creation_input_token_cost_above_1hr
jp.anthropic.claude-opus-4-8: 
jp.anthropic.claude-opus-5: 
jp.anthropic.claude-sonnet-4-6: 
jp.anthropic.claude-sonnet-5: 
meta.llama2-13b-chat-v1: 
meta.llama2-70b-chat-v1: 
us.writer.palmyra-x4-v1:0: 
us.writer.palmyra-x5-v1:0: 
writer.palmyra-x4-v1:0: 
writer.palmyra-x5-v1:0:

Price-Sync: litellm-providers

* fix(bedrock): correct eu.anthropic.claude-opus-4-5 regional prices from the AWS price list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 30 models

anthropic.claude-haiku-4-5-20251001-v1:0: 
anthropic.claude-opus-4-5-20251101-v1:0: 
anthropic.claude-opus-4-6-v1: 
anthropic.claude-sonnet-4-20250514-v1:0: 
anthropic.claude-sonnet-4-5-20250929-v1:0: 
apac.anthropic.claude-sonnet-4-20250514-v1:0: 
au.anthropic.claude-haiku-4-5-20251001-v1:0: 
au.anthropic.claude-opus-4-6-v1: 
au.anthropic.claude-sonnet-4-5-20250929-v1:0: 
eu.anthropic.claude-haiku-4-5-20251001-v1:0: 
eu.anthropic.claude-opus-4-5-20251101-v1:0: 
eu.anthropic.claude-opus-4-6-v1: 
eu.anthropic.claude-sonnet-4-20250514-v1:0: 
eu.anthropic.claude-sonnet-4-5-20250929-v1:0: 
global.anthropic.claude-haiku-4-5-20251001-v1:0: 
global.anthropic.claude-opus-4-5-20251101-v1:0: 
global.anthropic.claude-opus-4-6-v1: 
global.anthropic.claude-sonnet-4-20250514-v1:0: 
global.anthropic.claude-sonnet-4-5-20250929-v1:0: 
jp.anthropic.claude-haiku-4-5-20251001-v1:0: 
jp.anthropic.claude-sonnet-4-5-20250929-v1:0: 
mistral.voxtral-mini-3b-2507: 
mistral.voxtral-small-24b-2507: 
us-gov.anthropic.claude-sonnet-4-5-20250929-v1:0: 
us.anthropic.claude-haiku-4-5-20251001-v1:0: 
us.anthropic.claude-opus-4-1-20250805-v1:0: 
us.anthropic.claude-opus-4-5-20251101-v1:0: 
us.anthropic.claude-opus-4-6-v1: 
us.anthropic.claude-sonnet-4-20250514-v1:0: 
us.anthropic.claude-sonnet-4-5-20250929-v1:0:

Price-Sync: litellm-providers

* fix(bedrock): keep Claude Opus 5.5 response schema support as a maintainer ruled in #42626

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): stop pinning the jp Opus 4.7 cache field absence in the ttl fallback test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(prices): sync AWS Bedrock prices: 3 models

anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
global.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema
us.anthropic.claude-opus-5-5: supports_audio_input, supports_response_schema

Price-Sync: litellm-providers

* fix(bedrock): keep Claude Opus 5.5 response schema support per the #42626 ruling, reverting the cron restack

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:24:28 -07:00
devin-ai-integration[bot]
5c24802fbd
fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse (#42644)
* fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse

Bedrock rejects outputConfig.textFormat on Opus 4.7 and 4.8 with
"output_config.format: Extra inputs are not permitted", and the AWS
model cards list structured outputs as not supported for both, so
their cost-map entries no longer claim supports_native_structured_output
and json_schema requests fall back to the json_tool_call tool.

Fixes #27846

* test(bedrock): assert Opus 4.7 and 4.8 inline the schema on Invoke, move the native case to Sonnet 4.6

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 20:09:21 -07:00
devin-ai-integration[bot]
975bd28549
fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker (#42607)
* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): apply ruff format to invoke messages stream passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): drop drive-by reformat of existing invoke messages tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): collect streamed chunks into a tuple in passthrough regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): give the passthrough regression test a 10s first-chunk budget

* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]

* test(bedrock): take the gated byte stream's chunks as an immutable Sequence

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 20:01:56 -07:00
berriai-litellm-provider-info-sync[bot]
96a2015c83
chore(prices): sync OpenAI prices: 1 model (#42648)
sora-2-pro-high-res:

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:38 -07:00
berriai-litellm-provider-info-sync[bot]
0135387abf
chore(prices): sync Google Gemini prices: 1 model (#42642)
gemini/gemini-robotics-er-2-streaming-preview: input_cost_per_token, output_cost_per_token

Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-22 19:49:34 -07:00
yuneng-jiang
b395bfefdd
fix: repair seven regressions caught by CircleCI on main (#42640)
* fix: repair seven regressions caught by CircleCI on main

- vertex_ai: stop treating fine-tuned endpoint ids (numeric or
  vertex_ai/gemini/<id>) and gemma models as Gemini 3+, which injected
  temperature=1.0 and Gemini 3 thinking config into their requests (#42465)
- cost: price Azure DALL-E 3 from its azure/<quality>/<size>/dall-e-3 rows;
  it only worked through the OpenAI rows that #42435 removed
- bedrock: stream bedrock/invoke/moonshot through an OpenAI-shaped chunk
  decoder; the generic decoder dropped every chunk, which the
  supports_response_schema flag from #42338 un-skipped in CI
- proxy: keep the public model_group on pre-routing rejections so the Usage
  page groups them under the model name, not the deployment (#41077)
- cost map: mirror the base rows' capability flags onto Bedrock regional and
  cross-region copies (#42254 and later syncs)
- whitelist the new regional Bedrock rows from #42543 and #42588 for the
  converse routing check, following the existing regional-row convention

* fix(model-prices): mirror capability flags onto ap-southeast-3 bedrock rows

* refactor(bedrock): tighten types on the moonshot stream decoder and its tests
2026-09-23 02:26:36 +00:00
Emerson Gomes
30004f5f05
fix(bedrock): drop unsupported sampling params on converse reasoning models (#39834)
* fix(bedrock): drop unsupported sampling params on converse reasoning models

* test(bedrock): resolve duplicate import after rebase
2026-09-22 19:07:14 -07:00
devin-ai-integration[bot]
b0ac23d385
feat(logger): dispatch Python logging through the Rust diagnostics processor (#42616)
* feat(logger): add shared Rust diagnostics and Python logging bridge

* feat(logger): dispatch diagnostic processing through Rust

* chore: regenerate Cargo.lock after rebase

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: allowlist bounded logging tree walkers in recursive detector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(logger): skip decoding plain access arguments

* test(logger): skip embedded-python logger test when litellm deps are absent

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: cargo fmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: expect NativeDiagnosticProcessor in the native public surface

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(stub): export NativeDiagnosticProcessor via __new__ in _native.pyi

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): rename logger crate and document host sink contract

* test(logger): cover exc, stack, and nested extras in the diagnostic filter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logger): keep rendered redacted line when template scan flags a key pattern

The blanket REDACTED for a changed msg/color template discarded lines
whose rendered form was already redacted by the same pipeline, e.g.
'password=%s' became 'REDACTED' instead of 'password=REDACTED'. Only
fall back to REDACTED when the rendered form did not change either,
which is where interpolation can mangle the key pattern the scrub
would otherwise see.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): install python deps so the logger bridge test runs

The end-to-end bridge test skipped silently when litellm's Python deps
were absent. uv sync --no-install-project installs them without a
maturin build, and PYTHONPATH makes them visible to the embedded
interpreter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 18:44:15 -07:00
devin-ai-integration[bot]
327515a3ba
fix(mcp): preserve credential authority in DCR bridge authentication (#42563)
* fix(mcp): admit dcr_bridge envelope alongside an explicit litellm credential and mint under jwt principals

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(mcp): suppress LIT002 on concrete dict header payloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): mint bridge envelope for jwt mapped to a key without a user_id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): mint and admit bridge envelopes under the master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): bind mapped JWT envelopes to stored key tokens

* fix(mcp): preserve master envelope scope enforcement

* fix(mcp): reject bridge minting that loses JWT restrictions

---------

Co-authored-by: joshua <joshua@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-22 18:09:52 -07:00
moe-berri
2157351004
feat(ui): simplify auto-router setup and clarify feature limits (#42625)
* feat(ui): simplify auto-router setup and clarify feature limits

* fix(ui): validate auto-router drafts before saving

* fix: keep auto-router allowances consistent after deletes and refreshes
2026-09-22 18:03:32 -07:00
devin-ai-integration[bot]
944f44d82b
fix(utils): isolate callback errors in async_post_call_success_deployment_hook (#42535)
* fix(utils): isolate callback errors in async_post_call_success_deployment_hook

A callback that raises inside async_post_call_success_deployment_hook no longer
fails the completed request. The exception is logged with the callback class and
call_type, the response stays as it was, and later callbacks still run. Guardrail
callbacks are exempt because raising is how a post-call guardrail blocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): drop unrelated ruff autofixes from test_utils

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): drop fastapi import from guardrail propagation regression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): cover every success deployment hook call type with a raising hook

Parametrize the unit regression over video, embedding, responses, image, rerank,
transcription, chat and anthropic messages responses and assert the failure log
names the callback and call type. Run the integration test through a real proxy
for /v1/chat/completions, /v1/embeddings, /v1/responses and /v1/videos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): move raising success hook cases into the existing callback delivery file

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 17:59:50 -07:00
devin-ai-integration[bot]
a80379baf8
fix(proxy): keep the in-flight daily spend batch when shutdown cancels the flush (#42593)
* fix(proxy): keep the in-flight daily spend batch when shutdown cancels the flush

A daily spend batch drained from the in-memory queue was dropped for good when
the scheduler tick was cancelled by shutdown, because asyncio.CancelledError
bypasses the except Exception requeue. The flush now requeues the drained rows
on cancellation and re-raises, and each daily batch upsert runs in an
interactive transaction so a statement that already reached Postgres is rolled
back with the cancel instead of committing behind the requeue

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): requeue the cancelled daily spend batch before its rollback returns

Behind a lock the rollback of the cancelled interactive transaction only
returns once the blocked statement does, which is after the shutdown flush
has already run. The commit now runs as a shielded task so the cancelled
tick requeues the batch at once and lets the rollback finish in the
background. The final flush then finds the rows and writes them exactly once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): give the recording db a transaction seam for the bulk upsert tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): route the mocked daily tag spend upsert through the transaction seam

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): restore the drained Redis tag batch when shutdown cancels its commit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 17:58:08 -07:00
devin-ai-integration[bot]
8ee6bab529
fix(bedrock): treat blank AWS_S3_* env vars as unset for batch jobs (#42528)
* test(e2e): pin bedrock batch create with blank S3 env vars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): treat blank S3 env vars as unset for batch jobs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): trim blank S3 env gateway config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): register blank_s3_env capability and clean gateway tempdir

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): move blank S3 env batch test to its own module

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 17:55:22 -07:00
mubashir1osmani
b4ccb5b747
fix(s3): replace colons in generated log filenames (#40452)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(s3): replace colons in generated log filenames

Bedrock and Vertex AI batch file uploads use s3:// and gs:// URIs as
response ids. The shared filename sanitizer replaced slashes but kept
the scheme colon, producing log object keys that Hadoop-style consumers
reject as a relative path in an absolute URI.

Fixes #40234

* test(s3): drop docstrings flagged by review

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 17:33:44 -07:00
devin-ai-integration[bot]
7688f56256
fix(pricing): drop the duplicate cache_read_input_token_cost_batches key from 23 entries (#42623)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 17:28:48 -07:00
devin-ai-integration[bot]
ca95fc2bd4
fix: answer get_api_base for github_copilot and chatgpt without running the login flow (#42602)
* fix: answer get_api_base for github_copilot and chatgpt without running the login flow

* refactor(get_api_base): dispatch the provider helpers with if-chains

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 17:23:44 -07:00
Pawan Shahane
9bad2c35e5
fix(ollama): send PNG and JPEG images without requiring Pillow. (#41979)
* fix(ollama): send PNG and JPEG images without requiring Pillow

The ollama/ completion transport imported Pillow before it looked at the image, so every image request failed with a 500 on installs without Pillow. That includes the Docker image, where Pillow is only a CI dependency

Detect PNG and JPEG from their leading bytes and pass them through untouched. Pillow is now imported only when another format has to be re-encoded as JPEG, and that case still raises the same install hint

* fix(ollama): address Greptile findings on image conversion

Catch all exceptions on Pillow import, not just ImportError, so the helpful
install hint always appears. Break a line that exceeded 120 characters
2026-09-22 17:09:39 -07:00
devin-ai-integration[bot]
da82ea8e94
fix(ui): let the Create Key user picker find users by user_id, not just email (#41687)
* feat(ui): search users by id or email when assigning a key owner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): label users without an email by user id in key owner picker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): freeze merged user-filter where, format create key test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): suppress module-global patch findings in ui_view_users search test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep merged user-filter where as a plain dict for prisma serialization

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): forward search param from userFilterUICall to /user/filter/ui

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): mention user ID in the Create Key user picker helper text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 16:30:29 -07:00
devin-ai-integration[bot]
392e807172
feat(logging): add normalized_error cluster key to error_information (#41715)
* feat(logging): add normalized_error cluster key to error_information

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): stop classifying parameter length errors as context window errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): assert failure spend rows share normalized_error across provider wording

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): map agent model access denials and ignore non-string proxy error types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover budget exceeded errors with custom wording

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): cluster router no-healthy and provider-budget wording correctly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): let the exception class win over router fallback wording

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): cluster peer closed connection errors as provider connection errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): cluster tag routing denials as 403_MODEL_ACCESS_DENIED

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-22 15:56:50 -07:00
devin-ai-integration[bot]
58a05a9eae
fix(anthropic): return 400 instead of 500 when a content list holds a bare string (#42420)
* fix(anthropic): skip non-dict content items in beta-header and file-id helpers so malformed content lists return 400 instead of 500

Fixes #42094
Supersedes #42101

Co-authored-by: Pawan-Shahane <shahanepawan511@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): spawn the DB-less regression proxy with -P so the cwd cannot shadow the pinned checkout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): launch the DB-less proxy via -I -c with an explicit sys.path so python 3.10 works, drop DIRECT_URL, remove restating docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the self-booted DB-less proxy behind the owned_gateway opt-in the Buildkite container cannot satisfy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): move the bare string content item repro to tests/integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): wrap the anthropic bare string wire test to the 120 column limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): wrap anthropic common_utils test literals to the 120 column limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: fix ruff findings in touched test files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Pawan-Shahane <shahanepawan511@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:44:30 -07:00
devin-ai-integration[bot]
3db94b932e
fix(spend): return 400 from /spend/calculate for a model with no pricing row (#42497)
* fix(spend): return 400 from /spend/calculate for a model with no pricing row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): assert error type and param for unpriced /spend/calculate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): move the repro to tests/integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: alias ModelNotMappedError re-export to satisfy F401

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(utils): raise ModelNotMappedError only when the pricing row is missing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:44:00 -07:00
devin-ai-integration[bot]
238f434153
fix(otel): record the GenAI exception event through the Logs API on both OpenTelemetry lines (#42431)
* fix(otel): record the GenAI exception event without the removed Events API

OpenTelemetry removed opentelemetry._events in 1.44.0, so the three imports of
it broke 7 modules under litellm.integrations.otel, including the entry point.
Two things then failed quietly: with LITELLM_OTEL_V2 set the otel callback
resolved to None and nothing was exported, and with it unset the newrelic
callback was dropped as well, because that branch imports the v2 logger
ungated

Build and emit the event through the Logs API, which both lines carry. The
event name keeps riding the event.name attribute: the event_name log record
field that replaces it only exists from 1.44.0, and this package pins 1.28.0,
so the attribute is the only form both can write. It is also what the Events
API wrote, so exported events keep their shape

Emitting a plain record drops the default the Events SDK applied, so the
timestamp now falls back to time_ns() here

* style(otel): trim the event name key and regression test prose

Keep only the constraint a reader cannot infer from the code, that the
event_name record field does not exist on the pinned OpenTelemetry line

* fix(otel): export the GenAI exception event on both OpenTelemetry 1.28 and 1.44 lines

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): drop the record selection comment

Co-authored-by: Pawan-Shahane <shahanepawan511@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): collapse the record selection conditional for ruff format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): build the record fields with a dict literal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): set the native event_name on the 1.44 log record

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): spell out the record kwargs so the type gate sees each call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): suppress the version-gated kwargs for the type gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): drop the version-window prose and correct the event_name suppression reason

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): wrap the compat test docstring to the line limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Pawan-Shahane <shahanepawan511@gmail.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:16:46 -07:00
berriai-litellm-provider-info-sync[bot]
95d4613bd3
chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps] (#42591)
* chore(prices): sync xAI prices: 3 models, 3 new [3 with gaps]

xai/grok-code-fast: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema
xai/grok-code-fast-1-0825: supports_vision, supports_prompt_caching, input_cost_per_token, output_cost_per_token, input_cost_per_image_token, cache_read_input_token_cost, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, supports_function_calling, supports_tool_choice, supports_response_schema

* fix(prices): add context limits and reasoning flag to xai grok-code-fast aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prices): keep the xai sync diff limited to the alias fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 15:14:48 -07:00
devin-ai-integration[bot]
a173657dfb
fix(caching): keep embedding cache hits aligned with request inputs (#42571)
* fix(caching): keep embedding cache hits aligned with request inputs

Partial hits now send only the uncached inputs to the provider and merge
fresh vectors back into their original positions. Responses whose item
count differs from the input count (one input scoring many documents)
are no longer written to the per-input cache, since a later hit would
return a single item.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): drop mutable collection builds flagged by the type discipline gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): bypass embedding cache entries written before the per input cardinality check

Embedding cache entries now carry format_version and readers treat entries without it as
misses, so entries that only hold the first row of a multi row response are refetched instead
of served until their TTL expires. The provider call also receives a copy of the request kwargs
with the uncached inputs rather than mutating the caller's mapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): assert a partial embedding cache hit becomes a full hit on repeat

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): await pending embedding cache writes before asserting on cache hits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): validate cached embeddings without mutating responses or request kwargs

Validate cache rows through a frozen pydantic model so import does not depend on
TypeAdapter support for ReadOnly TypedDicts, accept string embeddings, build the
merged partial hit response instead of mutating the cached one, and hand the
provider request mapping to post call hooks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep cache_hit and response_ms on merged partial embedding hits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 17:14:44 -05:00
devin-ai-integration[bot]
d7c27cdc08
feat(proxy): configurable key_alias_pattern for key generate, update, and regenerate (#42553)
* feat(proxy): configurable key_alias_pattern for key generate, update, and regenerate

Adds litellm_settings.key_alias_pattern, a regex every key_alias sent to
/key/generate, /key/service-account/generate, /key/update, and
/key/{key}/regenerate has to fully match. A non-matching alias gets a 400
that names the setting and the pattern. When set, it replaces the built-in
rule enable_key_alias_format_validation turns on, and the baseline
unsafe-name check still runs first. An invalid regex fails config load.

* fix(proxy): cap key_alias length under key_alias_pattern and type the test fixtures

* style(proxy): declare key_alias_pattern with a PEP 604 union

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 14:52:01 -07:00
devin-ai-integration[bot]
d31e8aac6d
feat(cost-map): add Claude Opus 5.5 for Vertex AI and Azure AI (#42599)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:48:07 +00:00
devin-ai-integration[bot]
9082f8e5d0
feat(bedrock): add Claude Opus 5.5 pricing and capabilities (#42588)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:37:12 +00:00
devin-ai-integration[bot]
ecce7cdd9c
fix(proxy_cli): import proxy_server once on script-style boot (#42584)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 14:29:50 -07:00
joshua-berri
b277be0867
fix(mcp): preserve discovery attribution and sanitize logging headers (#42541)
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-22 14:26:48 -07:00
devin-ai-integration[bot]
6764868861
fix(otel): keep text completion choice fields beside the synthesized message (#42537)
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 14:24:32 -07:00
devin-ai-integration[bot]
075536eca1
chore(cost-map): remove models past their deprecation date (#42435)
* chore(cost-map): remove models past their deprecation date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop the empty parametrize left behind by the gemini web search removal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): drop merge base block left by conflict resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop gemini image cost tests pinned on removed model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:19:26 +00:00