Commit graph

360 commits

Author SHA1 Message Date
devin-ai-integration[bot]
efb93e62f8
fix(completion_extras): forward non-enum reasoning_effort through the Responses bridge instead of dropping it (#42452) 2026-09-23 11:04:07 -07:00
joshua-berri
b277be0867
fix(mcp): preserve discovery attribution and sanitize logging headers (#42541)
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-22 14:26:48 -07:00
devin-ai-integration[bot]
075536eca1
chore(cost-map): remove models past their deprecation date (#42435)
* chore(cost-map): remove models past their deprecation date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop the empty parametrize left behind by the gemini web search removal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): drop merge base block left by conflict resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost-calc): drop gemini image cost tests pinned on removed model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 21:19:26 +00:00
devin-ai-integration[bot]
7210403e23
fix(mcp): apply post-call rewrites without stale structured output (#41530)
* fix(mcp): apply async_post_mcp_tool_call_hook content changes to the tool result

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): drop structuredContent when a post-call hook rewrites tool content

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): satisfy type discipline and result contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): document internal logging patch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): avoid Final assignments inside callback loops

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): run every post-call hook and chain the rewritten content

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(mcp): credit the original fix from #33403

Co-authored-by: eric <mitrecx@163.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover post-call logging fallback paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover proxy hook logging context

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): preserve native structured guardrail replacements

* fix(mcp): invalidate stale structure after direct content edits

* fix(mcp): reconcile direct edits after callback exceptions

* fix(mcp): preserve successful in-place callback rewrites

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: eric <mitrecx@163.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-22 13:44:25 -07:00
togear
fc0055497c
feat: add configurable provider affinity header mapping (#41033)
* feat: add configurable provider affinity header mapping

* fix: sync provider affinity API types

* fix: harden provider affinity header mapping

* fix: avoid provider affinity import cycle

* fix: preserve input callback header mutations

* fix: address provider affinity code scanning findings

* fix: satisfy provider affinity type discipline gate

* test: cover omitted pre-call argument isolation

* fix: resolve remaining provider affinity codeql alerts

* fix: redact provider affinity headers after calls

* refactor: drop provider affinity header log redaction

* fix: reject control characters in affinity session ids as a bad request

* chore: regenerate the openapi snapshot on python 3.12 and reuse the session marker constant

* fix(responses): read the affinity session from the named metadata argument

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-22 13:31:33 -07:00
devin-ai-integration[bot]
3106d9c573
feat(rust-bridge): add cache and secret migration foundations (#42328)
* docs(rust): plan Python interop foundation

* fix(rust): preserve Python settings coercion at the native boundary

* chore(rust): drop interop planning note

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): resolve OCR provider secrets through an async SecretSource before transformation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): project the Python secret manager into the bridge and resolve OCR secrets through it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): drop premium_user from the secret manager snapshot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust-bridge): read the private key management globals once in the settings snapshot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): bound the bridge secret manager state cache to the active snapshot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): inline coercion unit tests

* fix(rust): preserve Python secret manager bindings

* refactor(rust-bridge): let settings projectors own their contract specs

Each settings group now declares its SettingSpec rows next to the projector
that reads them, and the manifest test derives python_settings.json from those
tables instead of a hand-copied duplicate. Field carries (group, name) instead
of a dotted path, and coercion gains the dict-item reader plus the Redis
Boolean, certificate-requirement, non-empty string, and numeric adapters that
the cache configuration projection adopts next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(rust-bridge): capture the secret manager binding in one settings read

The secret_manager accessor now carries the live client and settings objects,
so the bridge classifies the binding from a single snapshot instead of
re-reading litellm globals. The unreachable native arm and the service alias
go away, the binding-to-state mapping moves next to the snapshot, and the
Python callback precomputes its key_manager name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(rust-bridge): execute typed settings field declarations

* refactor(rust-bridge): compare cache backends by identity behind one exact trait

cache-response gains an object-safe ExactResponseCache so every exact-match
backend sits behind one pointer; WriteBuffer flushes through it. The bridge's
NativeResponseCache shrinks from nine variants and fifteen per-backend
accessors to an exact service plus the three semantic backends, and facade
mismatch detection compares BackendIdentity values instead of matching on
each backend type. Request projections move next to NativeRequest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(rust-bridge): drive both Python-embedded semantic caches through one execution

Redis-semantic and Valkey-semantic operations now share one SemanticExecution
body: await the Python embedder, seed the task-local vector, run the native
backend, repeat per batch entry. Valkey drops its with_embedder path in favor
of the same seeded embedder, and each backend keeps its own embedding-failure
policy. PythonEmbedder exposes one call shape. Redis-semantic thresholds are
compared at the backend's f32 width, which un-breaks the redis-stack parity
tests that a 0.8 facade threshold failed before this branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* wip

* feat(rust-bridge): complete response cache runtime surface

* fix(rust-bridge): preserve secret manager callback exceptions

* refactor(rust-bridge): unify route cache and secret rollout catalog

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 03:41:04 +00:00
devin-ai-integration[bot]
5cf17f9ce8
fix(responses): drop client_metadata and merge system messages for Databricks chat-only models (#42390)
* fix(responses): drop client_metadata before bridging to chat completions

Codex CLI sends client_metadata on every /v1/responses call. For a
provider with no native Responses config the chat-completions bridge
forwarded the raw kwargs, so client_metadata reached the provider as a
chat body field and Databricks rejected the request with an unknown
field 400. The bridge now drops the Responses-only request fields
before calling completion while still passing every other kwarg
through, so deployment-level params such as chat_template_kwargs keep
reaching providers without a native config.

* fix(databricks): merge consecutive system messages for chat-template models

Codex sends instructions plus a leading developer item, which the Responses
bridge and the developer-to-system translation turn into two consecutive
system messages that Databricks chat-template models reject with "System
message must be at the beginning". Each run of consecutive system messages
is now merged into one before the request is built for non-Claude models.

Also keep client_metadata out of the bridged chat request even when
allowed_openai_params names it, so both bridge branches drop the same set.

* fix(databricks): skip empty system messages when merging consecutive ones

Databricks drops empty content before the merge, so a system message in a
run could carry no content key and the merge iterated None. Those messages
are now skipped; a run with no content at all keeps its first message.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-21 20:09:54 -07:00
Mateo Wang
055b7314e0
Merge pull request #40121 from Atharva-Kanherkar/fix/mcp-responses-stream-single-lifecycle
fix(responses): stream one lifecycle across MCP auto-execute rounds
2026-09-21 15:11:57 -07:00
mateo-berri
e7557ada57 fix(responses): list executed MCP calls as completed mcp_call items in the final output 2026-09-21 14:56:02 -07:00
mateo-berri
239bff3315 test(responses): type the MCP lifecycle test helpers 2026-09-21 13:57:50 -07:00
Devin AI
83223885e6 fix(responses): forward safety_identifier through the chat completion bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 04:09:59 +00:00
Mateo Wang
40f51df90f
Merge pull request #41953 from BerriAI/litellm_bridge_drop_tool_search
fix(responses): drop tool_search and local_shell in the chat completions bridge
2026-09-19 04:42:28 -07:00
Mateo Wang
c39ec34553
Merge pull request #41234 from BerriAI/litellm_invalid_tool_choice_400
fix(utils): reject an untranslatable tool_choice with a 400 instead of a 500
2026-09-19 04:02:04 -07:00
mateo-berri
5477dbe74c fix(responses): drop tool_search and local_shell in the chat completions bridge
Hosted Responses API tools with no Chat Completions equivalent were forwarded
verbatim, so Codex 0.140+ got a 400 from the provider on every turn. The bridge
now drops tool_search and local_shell the same way it drops computer_use,
image_generation, and shell, and also drops parallel_tool_calls when no chat
tools remain, since chat completions only accepts it alongside tools
2026-09-19 03:55:23 -07:00
mateo-berri
b8c2787bcf Merge remote-tracking branch 'origin/main' into litellm_internal_copy_36887 2026-09-19 02:54:13 -07:00
mateo-berri
2c2aa5df11 Merge remote-tracking branch 'origin/main' into litellm_invalid_tool_choice_400
# Conflicts:
#	tests/test_litellm/responses/litellm_completion_transformation/test_litellm_completion_responses.py
#	tests/test_litellm/test_main.py
2026-09-19 02:35:29 -07:00
Mateo Wang
b7f07469bc
Merge pull request #41564 from BerriAI/litellm_responses_bridge_message_item_lit4622
fix(responses): announce message item before text events in the chat completions bridge
2026-09-18 22:36:28 -07:00
mateo-berri
4468c9fdcb fix(responses): close the reasoning item before announcing the message item 2026-09-18 21:37:01 -07:00
Mateo Wang
b46612cfeb
Merge pull request #41893 from BerriAI/litellm_fix_responses_ws_encrypted_content_affinity
fix(responses): restore encrypted_content and apply affinity on the native WebSocket relay
2026-09-18 21:09:17 -07:00
mateo-berri
9662b2a35c refactor(responses): type the websocket test parameters and suppress the error-frame send explicitly 2026-09-18 18:35:48 -07:00
mateo-berri
fcc7efa4db fix(responses): forward the routed input and report routing rejections on the websocket 2026-09-18 18:10:26 -07:00
mateo-berri
f213855558 refactor: drop the docstrings from the websocket relay and its tests 2026-09-18 17:44:32 -07:00
Mateo Wang
cda022ca68
Merge pull request #40243 from zoroyihan7/fix-responses-stream-error-events
fix(responses): emit typed streaming failure events
2026-09-18 17:29:57 -07:00
mateo-berri
5030214671 Merge origin/main into litellm_fix_responses_ws_encrypted_content_affinity
Resolves the conflicts with the WebSocket request defaults from main (PR #41881):
the relay keeps both custom_llm_provider and request_defaults, and a masked
response.create frame is re-serialized when the defaults changed it.

Keeps the first-frame routing hints (input, previous_response_id) out of the
deployment request defaults so they never get injected into later frames on the
same connection, with a regression test.
2026-09-18 17:23:59 -07:00
mateo-berri
febe9aec65 fix(responses): book a rejected WebSocket connection as a failed request 2026-09-18 17:10:01 -07:00
Mateo Wang
ff7dc86947
Merge pull request #41892 from BerriAI/litellm_gemini_contentless_candidate_finish_reason
Some checks failed
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Issue label sync / sync-issue-labels-tests (push) Has been cancelled
Issue label sync / sync-issue-labels (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
VS Code Extension / vscode-extension (push) Has been cancelled
fix(gemini): preserve candidates with finishReason and no content (#40477)
2026-09-18 16:34:46 -07:00
Mateo Wang
3490754e65
Merge pull request #41871 from BerriAI/litellm_bedrock_eager_input_streaming
feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools
2026-09-18 16:05:10 -07:00
Mateo Wang
ec9435cbf4
Merge pull request #41881 from BerriAI/litellm_responses_ws_deployment_defaults
fix(responses): merge deployment litellm_params into native websocket response.create frames
2026-09-18 16:04:39 -07:00
mateo-berri
1c15d9f291 fix(responses): restore encrypted_content and apply affinity on the native WebSocket relay 2026-09-18 15:44:06 -07:00
mateo-berri
417a88daed fix(responses): carry dict-valued reasoning_effort and keep the frame type on websocket defaults
A deployment whose reasoning_effort is an object is copied through as
reasoning the way the HTTP mapper does it instead of being dropped, and
the relay re-asserts the response.create frame type after merging
extra_body so a type key inside it can never replace it. The lazy
OpenAPI snapshot goes back to main: the earlier regeneration came from a
Python 3.14 interpreter dedenting docstrings, which CI on 3.12 rejects
2026-09-18 15:13:07 -07:00
mateo-berri
cf05466a27 fix(gemini): map every documented finishReason and reset per-candidate state
A content-less candidate is now kept as a choice whenever it carries a
finishReason, with the raw value on the choice's provider_specific_fields.
NO_IMAGE, IMAGE_RECITATION, IMAGE_OTHER and ESCALATION map to content_filter;
UNEXPECTED_TOOL_CALL and MISSING_THOUGHT_SIGNATURE map to stop. The
/v1/responses bridge reports content_filter and refusal as incomplete with
incomplete_details, and tool calls and reasoning no longer leak from one
candidate into the next.
2026-09-18 14:54:18 -07:00
mateo-berri
4a951847bb fix(responses): merge deployment litellm_params into native websocket response.create frames 2026-09-18 14:15:39 -07:00
mateo-berri
9d39d7efaa chore: merge main into fix/gemini-candidate-finish-reason-no-content 2026-09-18 14:05:25 -07:00
mateo-berri
df1b3c849b fix(responses): keep upstream error details in response.failed
RateLimitError and InternalServerError now carry the provider body, so
the OpenAI exception mapper keeps upstream codes like cyber_policy and
the upstream message instead of a generic mapped one

The proxy's response.failed event prefers the upstream body's code,
message, and type over the mapped exception's, and numeric error codes
in an error event map to their own HTTP status
2026-09-18 13:58:53 -07:00
mateo-berri
e5fd2bec83 fix: forward eager_input_streaming through the Responses API tool bridge 2026-09-18 13:34:29 -07:00
mateo-berri
c984c329fe Merge remote-tracking branch 'origin/main' into litellm_fix_responses_ws_litellm_params_leak 2026-09-18 13:32:23 -07:00
Mateo Wang
db756b9393
Merge pull request #40730 from BerriAI/litellm_responses_nested_additional_drop_params
fix(responses): honor nested additional_drop_params paths
2026-09-18 11:04:38 -07:00
mateo-berri
79029d89f9 test(responses): drop the history docstrings from the bridge regression tests 2026-09-17 15:51:35 -07:00
mateo-berri
ca91751d5b fix(responses): keep the addressed response id off bridged provider requests
The Responses id security hook keeps the id a client addressed under
`_litellm_addressed_response_id` in the request body so internal retries can
re-authorize it. On a model without a native Responses config that body is
bridged into `completion()` kwargs, the key was treated as a provider param,
and providers rejected it, so every follow-up turn carrying
`previous_response_id` returned 400.

Register the key in `all_litellm_params` so it is dropped before any provider
request, and share one constant between the hook and the param list.
2026-09-17 15:35:11 -07:00
joshua-berri
1e7b03a6ed
Merge pull request #41619 from BerriAI/litellm_fix_mcp_guardrail_context_4889
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 19:04:04 +00:00
Yujong Lee
b26935416a Merge remote-tracking branch 'github/main' into litellm_rust_bridge_declarative_route_catalog
# Conflicts:
#	tests/e2e/access_control/test_model_access_group_e2e.py
2026-09-17 11:08:08 -07:00
Yujong Lee
b170d61b8d route stuff through dispatch no direct main 2026-09-17 11:06:46 -07:00
Joshua Valluru
f4918e69f4 test(mcp): use the shared guardrail exception in regression 2026-09-17 10:54:28 -07:00
Joshua Valluru
743684bdbe fix(mcp): preserve request-selected guardrails during tool execution 2026-09-17 09:57:08 -07:00
Devin AI
5861240946 test(responses): type new streaming bridge test parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:43:14 +00:00
Devin AI
42c4c81633 fix(responses): allocate the message output index from the shared item allocator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:15:07 +00:00
Devin AI
d4e54a0f34 fix(responses): keep sync text deltas and give the message item its own output index
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:05:54 +00:00
Devin AI
441021fc96 fix(responses): announce message item before text events in the chat completions bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 04:46:26 +00:00
Yujong Lee
c635399f6e merge: origin/main into litellm_rust_bridge_declarative_route_catalog
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 22:54:51 +00:00
Yujong Lee
9617312ab2 add PublicDispatch 2026-09-16 15:44:46 -07:00