Commit graph

16473 commits

Author SHA1 Message Date
mateo-berri
6bd3699d43 fix(responses): keep guardrailed input items and bridge stream usage intact
- _write_back_structured_messages now patches only the rewritten rows back
  into the original input items, so reasoning items (encrypted_content),
  function_call ids, and web_search_call items survive a guardrail rewrite
  verbatim; rewrites that cannot be row-mapped fall back to the previous
  full conversion
- the responses bridge stream snapshot restores usage hidden in
  _hidden_params when stream_options is unset, so converted fake streams
  report real input_tokens instead of 0
2026-08-29 16:39:14 -07:00
mateo-berri
608603ee63 fix(speech): honor pcm response_format for Gemini TTS and reject unsupported containers 2026-08-29 16:36:34 -07:00
ryan-crabbe-berri
f82708f41d Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 16:36:09 -07:00
mateo-berri
4ef5db7c91 fix(responses): drop unsupported reasoning param for openai non-reasoning models 2026-08-29 16:36:05 -07:00
ryan-crabbe-berri
ce96db5a61 test(proxy): pass the window spend args in the access group requeue test 2026-08-29 16:36:04 -07:00
ryan-crabbe-berri
0777e37849 feat(ui): set a model access group's shared budget from the dashboard
Model access group budgets shipped API-only, so the only way to give a group a
budget was a curl. Adds an Access Group Budgets tab under Models & Endpoints
that lists every group with the spend drawn against its shared pool, and a
modal to set, edit or clear the budget.

/access_group/list now carries each group's budget and spend inline, so the
table renders from one read instead of one follow-up request per row.
2026-08-29 16:32:28 -07:00
mateo-berri
36c036cd2d test(reasoning-effort-grid): expect capped thinking on messages-route budget models 2026-08-29 16:28:49 -07:00
mateo-berri
a85e16f731 Merge remote-tracking branch 'origin/litellm_fix_audio_speech_content_type' into litellm_fix_gemini_tts_container 2026-08-29 16:27:27 -07:00
mateo-berri
d804b9d4fe fix(vertex_ai): skip non-dict property values in set_schema_property_ordering
The typed rewrite made the properties recursion call .get on every child,
so a malformed schema with a string or list property value raised
AttributeError where it previously passed through untouched.
2026-08-29 16:26:47 -07:00
mateo-berri
b0f4e19fcd Merge remote-tracking branch 'origin/litellm_fix_gemini_tts_response_format' into litellm_fix_gemini_tts_container 2026-08-29 16:26:34 -07:00
mateo-berri
855f56fa94 fix(openai): flatten top-level tool schema combinators on chat completions 2026-08-29 16:25:28 -07:00
ryan-crabbe-berri
3cc2f615da Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 16:24:07 -07:00
yuneng-jiang
e4ae1c1f2e
Merge pull request #38835 from BerriAI/litellm_38816_classifier_cost_savings
fix(proxy): count auto-router classifier cost in savings and benchmarks
2026-08-29 16:23:50 -07:00
ryan-crabbe-berri
140950f52d Merge branch 'litellm_window_spend_schema' into litellm_window_spend_writer 2026-08-29 16:23:45 -07:00
mateo-berri
af186eaaf3 fix(azure): flatten top-level tool schema combinators for Azure Responses GPT-4-family deployments 2026-08-29 16:23:01 -07:00
Mateo Wang
0c8321efc3
Merge pull request #38741 from BerriAI/litellm_fix_anthropic_messages_dict_detail_error
fix(anthropic_endpoints): serialize dict-detail HTTPExceptions on /v1/messages like sibling surfaces
2026-08-29 16:20:08 -07:00
Mateo Wang
6bc8dafa99
Merge pull request #38740 from BerriAI/litellm_vertex_gemini_35_transcribe
feat(vertex_ai): support gemini-3.5-transcribe on /v1/audio/transcriptions
2026-08-29 16:19:56 -07:00
Yuneng Jiang
fd4b540a44
test(e2e): require spent completion tokens before accepting empty fallback content
The first pass accepted any empty completion whose finish_reason was "length",
which also swallowed a fallback that produced nothing at all. Require the
response to have billed completion tokens as well, so empty content is
accepted only when the budget was demonstrably spent on non-visible reasoning.

Asserts on completion_tokens rather than reasoning_tokens because the latter
is provider-optional; with empty content and a refusal of null, consumed
completion tokens are reasoning by elimination, since visible text would be
content. Both counts are reported in the failure message.

Folds the three body accessors onto one _parsed helper instead of re-parsing
per call, and adds completion_tokens_of / reasoning_tokens_of alongside.
2026-08-29 16:19:41 -07:00
Tin Chi Lo
d3db7cebca fix(proxy): count auto-router classifier cost in savings and benchmarks
The LLM classifier's cost was recorded on the routing decision but never
reached any savings surface: per-request autorouter_savings stayed gross
and the session rollup recorded only the served request's spend, so
/auto_router/benchmarks overstated savings and understated routed spend.

Net the classifier cost into the savings figure at its one computation
owner and fold it into the rollup turn's spend, keeping
baseline_spend = spend + saved_spend. The response header's numeric
guard now shares the same reader.

Fixes #38816
2026-08-29 16:09:35 -07:00
Yuneng Jiang
6e341b79d8
test(e2e): stop the fallback tests flaking on gpt-5.5's reasoning budget
max_tokens=64 caps reasoning plus visible output on gpt-5.5, so the fallback
target can legitimately return finish_reason="length" with empty content.
litellm-e2e build 90 hit exactly that: the response cost of $0.002005 backs
out to 64 completion tokens at gpt-5.5's $3e-05/token, i.e. the whole budget
spent reasoning about "say hi" with none left to answer. The fallback itself
worked -- 200, served by gpt-5.5-2026-04-23, x-litellm-attempted-fallbacks
present -- so the only thing that failed was an assertion about OpenAI's token
budgeting rather than about routing.

Raise the reliability helper's budget to 512 and accept empty content only
when finish_reason is "length". Empty content under any other finish_reason
still fails, so the tests keep catching a fallback that returns nothing for a
reason we do control.

The relaxed assertion lives in the helper both reliability fallback tests
share, so test_timeout_routes_to_fallback is covered too; it has the same
shape and had not tripped yet.
2026-08-29 16:08:29 -07:00
ryan-crabbe-berri
ec934c490b
Merge pull request #38784 from BerriAI/litellm_model_access_group_budgets
feat(budgets): enforce shared budgets on model access groups
2026-08-29 16:06:54 -07:00
devin-ai-integration[bot]
3e2999f29f
fix(proxy): run SMTP send_email off the event loop with a connection timeout (#38473)
* fix(proxy): run SMTP send_email off the event loop with a connection timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): format utils.py and update _create_smtp_connection tests for timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep malformed SMTP_TIMEOUT inside the email error boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: retrigger ci

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: exclude misaligned circleci coverage flag from merged codecov report

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: retrigger ci for codecov and benchmarks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: disable carryforward for the circleci codecov flag

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: exclude carried-forward coverage from the codecov patch status

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: stop carrying forward the dead circleci codecov flag

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 16:05:57 -07:00
ryan-crabbe-berri
faad94af94 Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 16:02:46 -07:00
ryan-crabbe-berri
b2a08100e1 Merge branch 'litellm_window_spend_schema' into litellm_window_spend_writer 2026-08-29 16:02:23 -07:00
ryan-crabbe-berri
e263c09e4f test(e2e): cover model access group budgets against a live proxy
Four cases in tests/e2e/quota_management/budgets, driving real OpenAI calls
through a group whose shared pool is drained to exhaustion: the spender key
stays blocked, a key that spent nothing of its own is blocked by the same
pool, a sibling group with no budget keeps serving, and the budget read
reports the spend drawn against the group.

Adds set/get/delete access group budget to BudgetClient and the four
matching rows to the coverage registry.
2026-08-29 15:46:49 -07:00
mateo-berri
2247fbc66d fix(policy_engine): withhold streams when a pipeline guardrail rewrites output at runtime 2026-08-29 15:44:08 -07:00
mateo-berri
af179be681 fix(together_ai): stop writing context_length as max_output_tokens in the serverless sync
The Together catalog exposes only context_length, so the sync was recording
every chat model's context window as its output ceiling. New entries now carry
max_input_tokens and the legacy max_tokens from the catalog and get an output
ceiling only from a reviewed capability rule. GLM-5.2 and GLM-5.3-Flash rules
carry the documented 128K ceiling, and the 26 other inflated together_ai chat
entries drop max_output_tokens in both registry copies.
2026-08-29 15:26:44 -07:00
mateo-berri
f269b52a15 refactor(speech): type the bridge voice param as a mapping and wrap a long test line 2026-08-29 15:26:05 -07:00
mateo-berri
71a951691a fix(anthropic): cap reasoning_effort thinking budget below max_tokens on /v1/messages
A deployment carrying reasoning_effort in its litellm_params on the
/v1/messages passthrough mapped the effort to a legacy thinking block
whose budget_tokens was forwarded as is, so any request whose max_tokens
sat at or below that budget was rejected upstream with a 400. The mapped
budget now runs through the same cap the adaptive-to-legacy branch and
the chat path already use: it is clamped to max_tokens - 1, and dropped
with a warning when even the minimum budget cannot fit.

The cap helper becomes public since three call sites outside
AnthropicConfig use it.
2026-08-29 15:24:05 -07:00
ryan-crabbe-berri
40a7fe9221 fix(budgets): make the model access group ceiling exclusive
A pool whose recorded spend has reached max_budget has nothing left to give, so
the next request is refused rather than admitted. This departs from the tag
check it otherwise mirrors and matches where keys and organizations already
draw the line.

A non-positive budget now means no budget here too, so the read-time check and
the reservation path agree on what counts as unbudgeted.
2026-08-29 15:20:10 -07:00
mateo-berri
0a31ab38b9 fix(speech): stop forwarding response_format as a chat param for Gemini TTS 2026-08-29 15:17:22 -07:00
tin-berri
36ea28b092
fix(anthropic): emit signature-only thinking blocks on the /v1/messages bridge (#38809) 2026-08-29 15:04:22 -07:00
mateo-berri
1695b7f7b1 fix(headroom): fake-stream converted sync /v1/responses calls too
The sync response_api_handler agentic branch returned the completed ResponsesAPIResponse for a request the Headroom guardrail had converted from streaming, so litellm.responses(stream=True) handed callers a non-iterable object. Wrap it in the same fake stream the async path uses
2026-08-29 15:02:20 -07:00
ryan-crabbe-berri
d7c0bc1e6d test(budgets): clear the test-quality violations this branch added
The four model access group callback tests now share one helper, so nine
patches of proxy_server internals become three, and both mock-echo assertions
go with them. The delete_access_group tests share a context manager for the
same reason.

test_group_exactly_at_its_max_budget_passes gained the assertion it was
missing: it now proves the group reached the spend comparison, which a group
skipped for a missing budget row would not. The route-allowed patch beside it
was dead, so it is gone.

What is left is suppressed with the collaborator each one cannot inject.
2026-08-29 14:44:00 -07:00
Mateo Wang
20cfccaf5f
Merge pull request #37208 from BerriAI/litellm_managed_batches_observability
fix(batches): aggregate reasoning tokens and per-line pass/fail counts
2026-08-29 14:43:03 -07:00
mateo-berri
9448293903 fix(openai): flatten tool schema unions only for models whose validator rejects them
GPT-5 and later accept a top-level anyOf natively and call tools better with it intact, so the flattening now runs only for the gpt-4, gpt-3.5, chatgpt-4o, o1, o3, and o4 families. Non-dict tool entries pass through untouched, a typeless root that carries properties counts as an object, and the bounded $ref walker is listed in the recursion detector allowlist.
2026-08-29 14:38:08 -07:00
mateo-berri
ef96af5121 fix(headroom): resolve CCR retrieval on streaming /v1/responses
Streaming /v1/responses requests that carried Headroom's retrieve tool were
sent upstream as streams, so the model's headroom_retrieve function_call was
streamed straight back to a client that never declared the tool and the
retrieval never resolved. Chat completions already avoid this by converting
the request to non-stream in the pre-call deployment hook, letting the
agentic loop resolve the retrieve, and fake-streaming the final answer.

The hook now converts responses call types too, the responses handler wraps
the resolved result as a fake stream whenever any interception converted the
stream (shared converted_stream_requested helper instead of per-integration
key checks), and the follow-up request filter drops every non-code-interpreter
interception key through is_interception_internal_key.

Resolves LIT-6481
2026-08-29 14:34:42 -07:00
mateo-berri
90c8031dd7 fix(policy_engine): fail closed on content filter category MASK steps for streaming pipelines 2026-08-29 14:32:35 -07:00
Mateo Wang
c62c2afa09
Merge pull request #38234 from BerriAI/litellm_request_timeouts
fix(proxy): give every `requests` call a timeout so a silent server cannot hang the caller
2026-08-29 14:30:46 -07:00
mateo-berri
318b6a4b36 fix(mcp): forward staged credentials on /mcp-rest/test/connection like /test/tools/list
The connection preview built its temporary MCP client without the credentials the
not-yet-saved server config carries: the Authorization bearer an OAuth2 authorization_code
server had just been granted, the auth_value of an api_key, bearer_token, basic, or
authorization server, and the stored credentials of a saved server being edited. The tools
preview forwarded all three, so the same request succeeded there and failed on the connection
test with the generic "Failed to connect to MCP server" message

Both previews now resolve those credentials through one shared staging step, so they cannot
drift apart again, and the Authorization header is only forwarded upstream when the primary
x-litellm-api-key header carried admission, since otherwise it is the caller's LiteLLM key
2026-08-29 14:26:48 -07:00
ryan-crabbe-berri
acf3ed7d9b fix(budgets): narrow model access group spend counters to the served deployment
The database writer already intersects the auth-matched groups with the ones
the served deployment declares, but the live spend counters got the unnarrowed
set. A caller granted two pools that both cover a model group debited both
counters while only one row moved, so the in-memory ceiling could block a pool
its persisted spend never touched.

Narrow once at the callback so both consumers read the same set.
2026-08-29 14:25:56 -07:00
mateo-berri
429ad06972 fix(guardrails): write structured_messages rewrites back into /v1/responses input
Message-rewriting guardrails such as Headroom return their rewrite in
structured_messages and leave texts untouched. The responses guardrail
translation only mapped texts back, so compression never reached the
upstream request on /v1/responses while the retrieve tool still got
injected. Convert the returned messages back to Responses input (plus
instructions) the way the chat and Anthropic handlers already do, and
keep developer messages as input_text in the chat-to-responses bridge.

Resolves LIT-6494
2026-08-29 14:23:57 -07:00
mateo-berri
4cd11e8b19 fix(audio_utils): validate MPEG frame headers before labeling sniffed audio 2026-08-29 14:20:40 -07:00
mateo-berri
0c1f33dff7 fix(policy_engine): fail closed on content filter MASK steps for streaming pipelines
A litellm_content_filter step with a MASK action masks chat streams through
its own iterator hook under guardrails.add, which pipeline-managed guardrails
skip, so the pipeline path released the stream unmasked. CustomGuardrail now
declares rewrites_streamed_output (mask_response_content by default, any MASK
action for the content filter) and the upfront streaming check names such
steps in the same 400 it gives mask_response_content and incremental_diff
2026-08-29 14:19:21 -07:00
mateo-berri
c32eb41aad feat(openai_like): let a passthrough deployment keep cache_control ttl via model_info.cache_control_ttl
The supported_endpoints passthrough had no way to keep ttl for an upstream
that honors it, so the deployment now opts in with
model_info.cache_control_ttl: true, injected into the config the same way
the providers.json constraint is for JSON providers
2026-08-29 14:17:13 -07:00
mateo-berri
c251d6d609 fix(vertex_ai): label TTS audio bytes with their real content-type 2026-08-29 14:11:51 -07:00
ryan-crabbe-berri
6b2e7f8a1f refactor(budgets): declare route dependencies with Annotated instead of argument defaults 2026-08-29 14:10:58 -07:00
Mateo Wang
2963b47cda test: patch the Logging handler instead of the class in the poller error-file test 2026-08-29 14:09:33 -07:00
mateo-berri
b777947364 fix(openai): bound $ref expansion and keep root plus branch required when flattening tool schemas 2026-08-29 14:09:08 -07:00
mateo-berri
d51198fdeb test(policy_engine): cover streaming pipeline gate branches
Adds regression tests for the modify_response block on the Anthropic route,
the gate with no iterator overrides, and the per-chunk hook skipping
pipeline-managed guardrails. Corrects the gate docstring: an allow releases
the chunks as the endpoint translation left them, not verbatim
2026-08-29 14:07:34 -07:00