Commit graph

261 commits

Author SHA1 Message Date
mateo
134d1f3899 fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:17:14 +00:00
tin-berri
13837d319d
fix(openai): drop tool schema regex patterns OpenAI's validator cannot compile (#40485)
* fix(openai): drop tool schema regex patterns OpenAI's validator cannot compile

OpenAI validates function tool parameters with jsonschema's format checker,
which compiles every pattern with Python re. Claude Code's Artifact tool ships
an ECMA-262 pattern with \p{..} Unicode property escapes, so any OpenAI target
behind /v1/messages, /v1/responses or /v1/chat/completions 400s with
"Invalid schema for function 'Artifact': '...' is not a 'regex'" for every
model family. Drop only the patterns Python re rejects, keep the rest, at the
same seams that already flatten top-level combinators.

* fix(openai): walk only schema positions, iteratively, and drop regexes for every openai deployment

Review round: the regex sanitizer now walks JSON Schema applicator positions
only (properties, items, prefixItems, combinators, $defs, additionalProperties
and the rest), so a pattern key inside default, examples, const or a vendor
extension is data and stays. It also drops patternProperties keys Python re
cannot compile, which OpenAI checks the same way. The walk is level-order and
rebuilt deepest level first instead of recursive, so the code-quality recursion
gate passes and there is no depth cap below what a JSON parser admits. On the
chat wire an openai deployment with a custom api_base now drops such regexes
too, since that base is usually a proxy in front of the same validator, while
the lossier combinator flattening stays limited to api.openai.com hosts.
2026-09-09 22:08:30 -07:00
Mateo Wang
53f3e70f02
Merge pull request #40461 from BerriAI/litellm_lit_7373_responses_custom_tool_call_guardrail
fix(guardrails): scan and rewrite Responses custom_tool_call output items
2026-09-09 18:13:00 -07:00
mateo-berri
a248ff2c6e fix(guardrails): fail open on Responses tool-call rewrites in another shape and keep item names
Guardrails that hand back tool_calls in their own shape (vendor JSON, user code output) raised a KeyError on the non-stream Responses write-back. Returned tool calls are now validated before comparison; a shape or count that does not line up leaves every tool-call item unchanged and logs a warning naming the guardrail. A tool call's name is written back only when the guardrail changed it, so a nameless custom_tool_call no longer picks up the custom_tool placeholder.
2026-09-09 16:14:50 -07:00
mateo-berri
ae6893bc64 refactor(guardrails): narrow the Responses tool-call item mapping types 2026-09-09 15:38:55 -07:00
mateo-berri
0a053d2c81 test(guardrails): cover streamed tool-call name rewrites on chat and Messages
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Skipping the name write-back in either handler left every test green; a
guardrail that renames a tool call now has a regression test on both the
chat chunk path and the Anthropic SSE path
2026-09-09 14:21:40 -07:00
mateo-berri
a9cce4f1ed fix(guardrails): scan and rewrite Responses custom_tool_call output items
Post-call guardrails on /v1/responses only treated function_call output
items as tool calls, so a custom_tool_call item (Codex's exec shell tool
on GPT-5.6 models) was never scanned or masked, non-streaming and
streaming alike. Both item types now flow through the shared
tool_call_dict_from_output_item helper, ended-stream delivery syncs the
custom_tool_call_input delta/done events and the item's input field, and
the completed-response scan key fingerprints both kinds of item.

Non-streaming Responses tool-call MASK rewrites were also never written
back to the output item even for function_call; they are now.
2026-09-09 13:38:51 -07:00
mateo-berri
1133507565 refactor: type the buffered stream rewrite helpers without Any 2026-09-08 17:46:55 -07:00
mateo-berri
89e11949c8 fix(guardrails): key Responses stream tool-call rewrites by call_id
Bridged Responses streams give reasoning and message items output_index 0
and start function calls at 1, so keying rewrites by output_index rewrote
the wrong items. Rewrites now follow each function call's call_id through
the buffered item and argument events, refuse when an event cannot be
resolved to a rewritten call, and the refusal branches on all three
handlers get regression tests
2026-09-08 13:22:30 -07:00
mateo-berri
e65c56876f fix(guardrails): write tool-call rewrites into typed Responses envelopes
The response.completed envelope carries its function_call output items as
SDK objects without a get shim, so the write-back skipped them and the
envelope still showed the original arguments after every stream event had
been rewritten. Write the item whenever one is present, and cover the typed
event shape the live proxy carries in the handler test.
2026-09-08 11:42:51 -07:00
mateo-berri
cbe340a31c feat(guardrails): deliver tool-call rewrites into buffered streams
A post_call pipeline guardrail that rewrites a streamed tool call (its
arguments or its name) now has that rewrite written back across the buffered
chunks on chat, Responses, and Messages streams, so the client receives the
rewritten tool call instead of the original. The chat handler rewrites the
first fragment of each tool-call index and blanks the rest, the Responses
handler syncs the function_call output items and their argument events, and
the Messages handler rewrites the tool_use content_block_start and
input_json_delta events in both dict and SSE-bytes chunks.

The delivers_ended_stream_text_rewrites flag becomes
delivers_ended_stream_rewrites, since the write-back now covers both text and
tool calls, and the executor only discards a tool-call rewrite on translations
without write-back or on a shape the translation refuses.
2026-09-08 11:38:18 -07:00
mateo-berri
0d5ea553da Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline 2026-09-07 17:24:10 -07:00
Mateo Wang
9eaf15bcf9
Merge pull request #38842 from BerriAI/litellm_fix_responses_reasoning_drop_params
fix(responses): drop unsupported reasoning param for openai non-reasoning models
2026-09-07 15:14:13 -07:00
mateo-berri
d748cf40b7 fix(responses): only drop reasoning when it carries an effort and let an explicit map flag win
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
LiteLLM Rust / release wheel (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-09-05 18:53:37 -07:00
mateo-berri
f62130a479 fix(responses): floor reasoning support on the bundled cost map and resolve fine-tuned ids
A live cost map older than this release, or a proxy whose map fetch lags, could
strip `reasoning` from a model this release knows accepts it. The bundled map is
now the floor: any OpenAI entry it flags as reasoning keeps the param whatever
the live map says. Fine-tuned ids with an empty suffix (`ft:gpt-4o-2024-08-06:org::id`)
now resolve to their base entry instead of failing open, `chat-latest` carries
the flag, and the schema test keeps every codex, deep-research, and chat-latest
entry flagged. The none-effort check goes through a public wrapper so the
responses config stops importing a private helper.
2026-09-05 16:24:36 -07:00
devin-ai-integration[bot]
4df284e16d
fix(guardrails): record guardrail information for undecorated custom apply_guardrail overrides (#39727)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 11:39:24 -07:00
mateo-berri
d992937900 fix(responses): read reasoning support from the cost map instead of model-name rules 2026-09-05 03:17:11 -07:00
mateo-berri
832a05458b Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_responses_reasoning_drop_params 2026-09-05 02:40:00 -07:00
tin-berri
78ad88f52c
fix(responses): decode JSON-string tool schemas before sending to the provider (#39844)
* fix(responses): decode JSON-string tool schemas before sending to the provider

A caller that hands a tool schema over already JSON-encoded reached the
Responses API with a string `parameters`, and the provider rejected the
request with a 400 naming the routed model instead of the offending tool.
Decode it at the one place every Responses request converges, and refuse
anything that is neither an object nor a string encoding one.

Collapses the duplicated input/tool sanitization block shared by the
request and compact-request builders into a single owner, so the decode
cannot be wired into one path and not the other.

* test(responses): pin null tool schemas as accepted, and type the parametrized cases

The Responses API serves `parameters: null` and an omitted schema alike, so
neither may raise. Pin both against a future tightening, annotate the
parametrized inputs, and trim the docstrings back to what the code does not
already say.
2026-09-05 05:14:44 +00:00
mateo-berri
19da217167 fix(openai): mint workload identity tokens for PrivateLink and regional api.openai.com hosts 2026-09-03 14:38:47 -07:00
Mateo Wang
025a3ca42f
Merge pull request #39631 from BerriAI/litellm_gpt_6_astra_detection
fix: treat gpt-6 names as the gpt-5 request family in OpenAI and Azure configs
2026-09-03 14:00:59 -07:00
mateo-berri
ab515dbc90 fix: treat gpt-6 names as the gpt-5 request family in OpenAI and Azure configs 2026-09-03 13:20:13 -07:00
Mateo Wang
80250807db
Merge pull request #38808 from BerriAI/litellm_headroom_ccr_streaming_responses
fix(headroom): resolve CCR retrieval on streaming /v1/responses
2026-09-03 13:13:18 -07:00
Mateo Wang
1e2d6abc18
Merge pull request #39614 from BerriAI/litellm_fix_stream_usage_default_openai_hosts
fix(openai): default stream usage on PrivateLink and regional api.openai.com hosts
2026-09-03 13:13:00 -07:00
mateo-berri
1e75668a25 fix(openai): default stream usage on PrivateLink and regional api.openai.com hosts 2026-09-03 12:37:17 -07:00
mateo-berri
ec2e35b679 fix(image_gen): keep the provider's echoed size, quality, and output_format on gpt-image responses 2026-09-03 11:07:08 -07:00
mateo-berri
4d3c1998af fix(image_gen): report the requested output_format on gpt-image responses 2026-09-03 10:50:23 -07:00
Mateo Wang
4f7b20ec10
fix(guardrails): skip streaming guardrail rounds that re-scan cleared output (#39386)
* fix(guardrails): skip streaming guardrail rounds that re-scan cleared output

Streaming guardrails scanned the finished answer twice at end of stream
whenever the chunk count landed on a multiple of the sampling rate, ran
sampled rounds whose payload was identical to the previous one, and on
/v1/messages could scan an empty text before the first content chunk.
Every redundant round is a paid guardrail provider call.

Each endpoint handler now exposes a scan key describing what a round
would hand to apply_guardrail (the text so far, plus tool calls once the
stream has ended), and the unified streaming hook skips a sampled or
end-of-stream round whose key equals the last scanned one or carries
nothing to scan yet. Rounds that carry tool calls are never skipped.

* test(guardrails): expect one end-of-stream scan when the terminal chunk is sampled

Update sampled cadence expectations and use tuple-backed scan state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 18:25:28 -07:00
yassin
6d5a3ab42c merge: litellm_internal_staging into litellm_headroom_ccr_streaming_responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 00:24:17 +00:00
mateo-berri
7cbc74399d fix(responses): accept pydantic tool objects returned by guardrails 2026-09-02 15:24:43 -07:00
mateo-berri
5561b8438c fix(responses): keep a namespace's non-function members when every function member is dropped 2026-09-02 14:18:16 -07:00
mateo-berri
dc12e4c2b4 fix(responses): match guardrail tools by ordinal in one pass
Sort the chat-tool keys once and number duplicates with groupby instead of
rescanning every preceding key per position, so the guardrail merge stays
O(n log n) on client-supplied tool lists. Drop the comment that restated the
unsupported-tool warning in the Responses-to-chat transformation.
2026-09-02 12:57:01 -07:00
mateo-berri
d7ee215c57 fix(responses): keep namespace tools intact when a guardrail returns them unchanged
Any pre_call guardrail on /v1/responses flattened Codex namespace tools
into ns__member functions and wrote the flattened list back to the
request, so the model called mcp__server__tool with no namespace and
Codex rejected the call as unsupported.

The handler now keeps the client's original tools, hands the guardrail a
deep copy of the flattened ones, and rebuilds data["tools"] by matching
the guardrail's output to the originals by type and name. Unchanged
tools go back as the original objects, a dropped or edited namespace
member changes only that member, and tools the guardrail injects are
still appended.

Fixes #39183
2026-09-02 11:32:04 -07:00
mateo-berri
4a5d0b8163 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_headroom_ccr_streaming_responses
# Conflicts:
#	litellm/llms/openai/responses/guardrail_translation/handler.py
#	tests/test_litellm/llms/openai/responses/test_openai_responses_guardrail_handler.py
2026-09-01 22:25:59 -07:00
mateo-berri
c09db7c7a3 Fail closed on rewrites for buffers that never reached their terminal event
An Anthropic buffer without a stop_reason only ran the flat text scan, so a
rewrite there was dropped while the executor trusted the translation to have
delivered it. A Responses buffer ending at response.output_item.done returned
after the tool-call scan without ever checking the text. Both now reach the
flat scan and raise UndeliverableStreamRewrite when a caller expects the
rewrite delivered, matching the existing Responses no-envelope fallback.
2026-09-01 22:01:04 -07:00
mateo-berri
2d7697569a Merge branch 'litellm_fix_post_call_policy_pipeline' into litellm_post_call_pipeline_stream_rewrite
Brings the stack base up to origin/litellm_internal_staging. Staging's
get_streaming_string_so_far now joins delta-only Responses parts itself, so
the handler's separate delta joiner and the flag-gated delta scan go away
and the fallback keeps failing closed on a rewrite it cannot deliver. The
two tests that pinned the old flag-gated behaviour are dropped in favour of
staging's delta-assembly tests, and _process_ended_stream takes the same
UserAPIKeyAuth | None the typed process_output_response now expects
2026-09-01 19:04:22 -07:00
mateo-berri
c9435b5ff3 fix(guardrails): key stream rewrites by choice index and scan delta-only responses buffers
Chat streaming write-backs now match chunks by the choice's index field
instead of its list position, delivering rewrites to the right choice on
n>1 streams; an ended-stream rewrite on a multi-choice buffer fails
closed since stream_chunk_builder collapses the choices. The Responses
fallback joins output_text.delta events when delivery is expected, so a
delta-only buffer is guardrail-checked instead of released raw.
2026-09-01 18:00:42 -07:00
mateo-berri
8c7fe00d80 fix: compare stream event types by equality so typed completed events keep their usage 2026-09-01 17:37:07 -07:00
mateo-berri
4fbe4ce2e2 fix(guardrail_translation): deliver stream rewrites on incomplete and failed responses terminals 2026-09-01 17:29:52 -07:00
mateo-berri
ec677e5e74 Merge remote-tracking branch 'origin/litellm_fix_post_call_policy_pipeline' into litellm_post_call_pipeline_stream_rewrite
# Conflicts:
#	litellm/llms/anthropic/chat/guardrail_translation/handler.py
#	litellm/llms/openai/chat/guardrail_translation/handler.py
2026-09-01 17:01:42 -07:00
mateo-berri
a38dfecd96 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks 2026-09-01 14:49:06 -07:00
Mateo Wang
deb67ce6e2 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream
# Conflicts:
#	tests/test_litellm/proxy/guardrails/guardrail_hooks/test_bedrock_guardrails.py
2026-09-01 14:11:29 -07:00
Mateo Wang
286a754999
Merge pull request #38839 from BerriAI/litellm_fix_chat_anyof_tool_schema
fix(openai): flatten top-level tool schema combinators on chat completions
2026-09-01 13:25:03 -07:00
Mateo Wang
8ae072b501
Merge pull request #39147 from BerriAI/litellm_openai_drop_toolless_tool_choice
fix(openai): drop tool_choice when request has no tools on chat completions
2026-09-01 12:50:38 -07:00
mateo-berri
4567fc784c fix: close non-message open items as incomplete when a stream is blocked 2026-09-01 12:45:23 -07:00
mateo-berri
4c7dd0b522 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks
# Conflicts:
#	type-discipline-budget.json
2026-09-01 12:36:48 -07:00
mateo-berri
c01ef712a2 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_headroom_ccr_streaming_responses
# Conflicts:
#	tests/test_litellm/responses/litellm_completion_transformation/test_streaming_iterator_transformation.py
2026-09-01 11:27:03 -07:00
mateo-berri
93a03a9ffd fix(openai): drop tool_choice when request has no tools on chat completions 2026-09-01 11:16:56 -07:00
mateo-berri
ab2c9aed0f fix(responses): normalize tool call id shapes across the anthropic bridge and openai replay
The chat-completions bridge emitted Responses output items whose item ids
were raw Anthropic tool ids (toolu_/srvtoolu_), which OpenAI rejects on
replay with "Expected an ID that begins with 'fc'", breaking router
fallback conversations from gpt-5 to claude models.

Four fixes, composable and independently useful:
- emission: bridge output items get fc_/ctc_-prefixed item ids while
  call_id stays raw so tool_result pairing keeps working (streaming and
  non-streaming share the same helpers)
- openai replay: request transformation drops tool call item ids that do
  not match OpenAI's own shapes instead of forwarding them, gated to
  OpenAI and Azure, since the API accepts the items with no id at all
- anthropic replay: a replayed srvtoolu_ call whose paired server tool
  result is unavailable degrades to a plain client tool_use instead of a
  dangling server_tool_use that 400s the client's tool_result
- tool-only turns no longer emit a message output item with output_text
  text null, matching native OpenAI output
2026-09-01 11:12:24 -07:00
mateo-berri
0a9676bd4f fix(openai): scope unknown-model reasoning_effort forwarding to the plain openai provider 2026-08-31 21:42:31 -07:00