Commit graph

53239 commits

Author SHA1 Message Date
devin-ai-integration[bot]
6f5ad78a1f
fix(cost-map): registry audit 2026-09-28, openai deep-research shutdown dates, azure deepseek v4.1 flash direct price, vertex gemini 3.8 live avatar video price, drop azure_ai/muse-spark-1.3 (#43566)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: KaiyiQuan <KaiyiQuan@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 09:46:37 -07:00
devin-ai-integration[bot]
f191e08d67
fix(proxy): log key owner identity on expired key auth failures (#43105)
* fix(proxy): log key owner identity on expired key auth failures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): escape control characters in logged key identity fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): gate key identity in auth failure logs behind log_auth_failure_key_identity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply log_auth_failure_key_identity from DB config reloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 08:18:48 -07:00
devin-ai-integration[bot]
c60c714278
test(rust): enforce shared upstream error contract in wheel checks (#43520)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 08:15:41 -07:00
devin-ai-integration[bot]
2c9b0e00ac
refactor(types): replace Any with proven types in 8 files (#43551)
* refactor(types): replace Any with proven types in 8 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep email logger untyped where its alert types differ

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 04:26:41 -07:00
devin-ai-integration[bot]
90e4962c81
refactor: clean up fresh tech debt from 2026-09-27 (#43538)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 01:28:40 -07:00
berriai-litellm-provider-info-sync[bot]
b79fc9f1b0
feat(azure): add mistral ocr pricing and azure max output limits (#43530) 2026-09-27 23:07:45 -07:00
devin-ai-integration[bot]
74cad08997
refactor(rust): remove delivery routing abstraction (#43514)
* wip

* refactor: finish removing delivery routing abstraction

* refactor(rust): remove chat completion decline admission

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 03:31:22 +00:00
devin-ai-integration[bot]
f184ace25b
feat(rust): add the MCP gateway (#43470)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
6e0926edde
feat(rust): add gateway UI login and sessions (#43469)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
ed43556e92
feat(rust): add virtual key storage contracts (#43468)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
5f637a2b11
feat(rust): separate gateway authentication and authorization (#43467)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
876539e1b3
feat(rust): add structured route lifecycle tracing (#43466)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:18:58 -07:00
devin-ai-integration[bot]
8ab124309f
feat(rust): connect Python inference bindings to shared routes (#43465)
* feat(rust): support the HTTP Responses API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): enable native Python inference opt-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust_bridge): cover only python-only routes in the native load guard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep Python inference rollout disabled

* fix(rust): preserve inference defaults and continuation IDs

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:05:39 -07:00
devin-ai-integration[bot]
9b08112ed2
feat(rust): add litellm-db and litellm-db-testing workspace scaffolding (#43504)
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 18:48:45 -07:00
berriai-litellm-provider-info-sync[bot]
cc30d93984
fix(cost-map): set together_ai gpt-oss-20b and gemma-4-31B-it deprecation_date to 2026-09-15 (#43509)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:49:49 -07:00
berriai-litellm-provider-info-sync[bot]
2566b33cea
fix(cost-map): sync openrouter prices from the models API (#43506)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:38:47 -07:00
devin-ai-integration[bot]
0a21f24d50
feat(rust): support the HTTP Responses API (#43464)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 17:38:15 -07:00
berriai-litellm-provider-info-sync[bot]
0d6b8b5ab4
fix(cost-map): add deprecation_date to together_ai Salesforce/Llama-Rank-V1 (#43507)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:25:56 -07:00
devin-ai-integration[bot]
1ceeefbf84
refactor(rust): use shared execution in gateway inference (#43463)
* feat(rust): add the HTTP host driver

* refactor(rust): use shared execution in gateway inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:50:05 -07:00
devin-ai-integration[bot]
438bffc26e
build(rust): package the gateway container (#43471)
* build(rust): package the gateway container

* ci: exempt the gateway Dockerfile from the CI coverage gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:49:39 -07:00
devin-ai-integration[bot]
18933c8a21
feat(rust): add the HTTP host driver (#43462)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 16:18:10 -07:00
devin-ai-integration[bot]
36784e3b79
refactor(rust): share call lifecycle across route-owned inference (#43461)
* feat(rust): expand gateway configuration parsing

* refactor(rust): unify core calls and host lifecycle

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): unify core calls and host lifecycle

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): read environment secrets when litellm is not importable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(core): drop the duplicate rstest attribute

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): make shared route dispatch route-owned

* fix(rust): satisfy Clippy in messages regression test

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:05:30 -07:00
devin-ai-integration[bot]
b94f5bdbed
feat(rust): expand gateway configuration parsing (#43460)
* feat(rust): expand gateway configuration parsing

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:54:04 -07:00
devin-ai-integration[bot]
268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
berriai-litellm-provider-info-sync[bot]
22b36cbcf6
chore(cost-map): update azure_ai/grok-4.6 input price and add azure_ai/MAI-Cyber-1-Flash (#43446)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 09:14:33 -07:00
berriai-litellm-provider-info-sync[bot]
ff462f7a77
chore(cost-map): update azure_ai/grok-4.6 input price from Azure pricing page (#43440) 2026-09-27 08:56:50 -07:00
devin-ai-integration[bot]
f4308bc124
refactor(types): replace Any with proven types in 5 files (#43304)
* refactor(types): replace Any with proven types in 6 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): keep enterprise email import inside try-except for unsafe-import check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): keep email_logging_instance annotation as Any pending a guarded alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert iterator override typing in proxy utils

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 01:28:02 -07:00
devin-ai-integration[bot]
8e6d99d74a
fix(token_counter): count Gemini function_declarations tools (#43417)
* fix(token_counter): count Gemini function_declarations tools

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(token_counter): skip non-dict tools when formatting definitions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 22:28:31 -07:00
Anmol Jaiswal
4274bdda44
fix(vertex_ai): stop importing the vertexai SDK in partner-model completion (#42274)
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
2026-09-26 22:14:53 -07:00
devin-ai-integration[bot]
b831e9b4ac
fix(bedrock): keep the provider status code on unprocessable image errors (#43416)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: dbalintx <dbalintx@amazon.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:10:24 +00:00
devin-ai-integration[bot]
c1f761eba5
test(vertex_ai): move stray Gemma streaming tests to tests/unit so CI coverage passes (#43422)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:02:44 +00:00
devin-ai-integration[bot]
9a0ff249d5
fix(anthropic): forward the per-turn-control beta to Azure AI Foundry (#43415)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:02:35 +00:00
devin-ai-integration[bot]
491d454826
fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking (#43414)
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking

Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.

Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.

(cherry picked from commit bc9b6f8a5c)

* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard

PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 04:51:56 +00:00
Stewart Park
2101c860c2
fix(vertex_ai): make Gemma fake streams work with traced Responses (#43147)
* test(vertex_ai): reproduce traced Gemma Responses stream failure

* fix(vertex_ai): wrap Gemma fake streams for Responses tracing

* test(vertex_ai): cover Gemma traced streams and usage options

* test(vertex_ai): inject gemma test deps and assert hidden usage accounting

Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.

Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.

* test(vertex_ai): drop explanatory comment from usage-option assertions
2026-09-26 21:26:38 -07:00
Thippaluri Yaseen Basha
829cba1bf1
fix(gemini): forward seed to the Gemini API instead of rejecting it (#43197)
* fix(gemini): forward seed to the Gemini API instead of rejecting it

The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough

* test(gemini): assert the forwarded seed without mutating shared state
2026-09-26 21:18:33 -07:00
Techboy bebop
21055e3fd8
fix(tools): salvage concatenated JSON tool call arguments (#43260)
* fix(tools): salvage concatenated JSON tool call arguments

* fix(tools): harden concatenated tool-call salvage for review findings

Skip non-dict JSON during split so salvage cannot emit empty tool calls.
Collapse srvtoolu_ expansions to the first object so server results stay paired.
Allocate __concat_n ids that cannot collide with sibling tool call ids.
Propagate cache_control onto every expanded Anthropic tool_use block.
Rename the XML invoke loop variable so the key-leak gate no longer flags {args}

* test(tools): cover concat id bump and srvtoolu array keep

Only collapse srvtoolu_ when concatenated salvage expanded; a valid JSON
array argument stays one server tool input

* revert(anthropic): drop concat expansion from pass-through adapter

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

* revert(tools): keep concat salvage out of request-side tool converters

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

* fix(tools): expand strictly salvaged concatenated tool arguments in normalized tool calls

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

* fix(tools): retain at most the salvage cap while validating concatenated arguments

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

* test(tools): assert concat sibling ids unique after sanitization

A sibling id that only collides after colon-to-underscore sanitization must force the next concat suffix

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

* refactor(tools): drop unused strict mode from split_concatenated_json_objects

Strict mode had no production caller. Rejection cases now sit on salvage, and split matches upstream main

Co-authored-by: Techboy bebop  <kumarpriyanshu09@users.noreply.github.com>

---------

Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
2026-09-26 21:15:12 -07:00
yuneng-jiang
303434d573
test(e2e): report batch cleanup leftovers as a plain UserWarning (#43405)
The leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. Same change as #43391 on rc/1.103.0
2026-09-27 03:26:23 +00:00
Jeremy Schoemaker
ba2d1c2785
fix(anthropic): drop thinking blocks with empty thinking text, not just missing signature 🧠🚫 (#38049)
_is_unsignable_thinking_block() only checked block["signature"], so a
thinking block with a valid-looking signature but empty (or
whitespace-only) thinking text sailed through _drop_unsignable_thinking_blocks
and into anthropic_messages_pt(). Anthropic rejects that with:

  400 messages.N.content.M.thinking: each thinking block must contain thinking

This is reachable whenever a thinking_blocks history item gets replayed
through this Anthropic-shaped request path (e.g. a non-Anthropic reasoning
turn with no summary text), the same class of bug PR #36033 fixed on the
Responses adapter's own separate code path.

Now the signature check runs first (unsigned blocks are still dropped, same
as before), then an additional check drops the block if `thinking` is
missing, not a string, or strips to empty. redacted_thinking blocks are
untouched since they don't have type == "thinking".
2026-09-26 20:17:44 -07:00
agustin18
9ba552d527
fix(vertex_ai): consider tools when validating context caching min tokens (#43319)
* fix(vertex_ai): consider tools when validating context caching min tokens

Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.

Fixes #42804

* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
2026-09-26 19:51:45 -07:00
yuneng-jiang
be35b22dfc
fix(streaming): keep litellm Usage on text-completion usage chunks (#43047)
* fix(streaming): keep litellm Usage on text-completion usage chunks

* fix(streaming): convert provider usage to litellm Usage instead of dropping it
2026-09-27 01:53:12 +00:00
tin-berri
013d5fa015
feat(cli): reuse saved agent setup and add reconfigure (#43392) 2026-09-26 18:46:36 -07:00
devin-ai-integration[bot]
7aba77197d
feat(otel): add SigNoz preset for OpenTelemetry v2 (#43296)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
* feat(otel): add SigNoz preset for OpenTelemetry v2

Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite

Absorbs the work from https://github.com/BerriAI/litellm/pull/38206

Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): drop explanatory comments from the SigNoz preset

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(types): keep signoz dynamic param lines within ruff format width

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(signoz): assert the missing-endpoint boot path directly instead of in an except block

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for the signoz health service

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 18:15:45 -07:00
tin-berri
de06c93767
feat(router): opt in to prompt-cache cost routing (#43232)
* feat(router): opt in to prompt-cache cost routing

* fix(router): address prompt-cache routing review

* ci: include cache-routing regressions in coverage
2026-09-26 18:13:13 -07:00
devin-ai-integration[bot]
501ef23f4a
feat(rust): add the openai_like chat config foundation (#43379)
* feat(rust): add the openai_like chat config foundation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): let max_completion_tokens outrank max_tokens and decline refusal responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 18:07:09 -07:00
devin-ai-integration[bot]
b396b0b724
feat(e2e): read management routes back from the control plane replicas (#43373)
* feat(e2e): read management routes back from the control plane replicas

The suite's management read-backs (/key/info, /team/info and friends)
polled the same replica list as the data plane. On a componentized stack
whose LITELLM_PROXY_REPLICA_URLS names the gateway pods directly, that
list answers those routes 404, since a gateway pod trims the management
routes at startup. A new LITELLM_CONTROL_PLANE_REPLICA_URLS names the
addresses a management read-back polls instead: an exported list wins,
and when it is unset the old rule stands, the data-plane replicas while
the control plane shares the suite's base URL and the control-plane base
alone once it is split.

build_proxy_client takes the list as control_replica_urls and
read_back_everywhere picks its replicas per path, the way the rest of the
client already does.

* fix(e2e): derive the control replicas of a client built for another proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-26 18:05:55 -07:00
devin-ai-integration[bot]
eea1d0f269
fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507)
* fix(responses): stream guardrail pre-call block as SSE with a typed output item

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): import blocked usage helper from the guardrail utils module

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop narrating docstrings and poll without rebinding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover pre-call guardrail block on /v1/responses stream and json

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for responses guardrail block contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): observe upstream on the recorded chat route for responses denial cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tidy responses denial audit cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for worker count to recover after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): require a replacement worker after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the blocked response test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-26 18:05:21 -07:00
berriai-litellm-provider-info-sync[bot]
e73abe6c72
chore(cost-map): drop stale cache hit field from openrouter deepseek-v4-pro-0813 (#43389)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-26 17:36:34 -07:00
berriai-litellm-provider-info-sync[bot]
4c2458a0b1
chore(cost-map): sync openrouter prices for deepseek, minimax, qwen and glm rows (#43384)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-26 17:18:23 -07:00
tin-berri
5d777c16d9
fix(mcp): align hub publication status and controls (#43241)
* fix(mcp): align hub publication status and controls

* refactor(mcp): keep hub visibility guard outside table rendering

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-26 16:53:02 -07:00
tin-berri
7244040908
fix(mcp): report reachability without stored credentials (#43240) 2026-09-26 16:44:49 -07:00