Commit graph

16900 commits

Author SHA1 Message Date
samtsai15
0e7562dbc6 test(guardrails): cover every Anthropic image source shape in the extractor's own suite
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.

Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].

The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.

Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:37:51 +08:00
mateo-berri
de1f38820a fix(passthrough): flush interrupted streams on client disconnect and reuse cached gigachat http clients 2026-08-30 13:36:51 -07:00
mateo-berri
a5fa8ebfa7 fix(passthrough): keep upstream error body readable for streaming error status mapping 2026-08-30 12:59:18 -07:00
mateo-berri
db1e0717f9 fix(guardrail_translation): assemble responses stream text from delta events for terminal-failure scans 2026-08-30 12:52:03 -07:00
mateo-berri
fb9ec79d7c Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:49:29 -07:00
mateo-berri
4261198b2f Merge branch 'litellm_internal_staging' into litellm_headroom_ccr_streaming_responses 2026-08-30 12:47:33 -07:00
mateo-berri
b8e11a75fa test: use local model cost map in import-isolation subprocess 2026-08-30 12:47:15 -07:00
mateo-berri
99a6dd02af fix(proxy): narrow audio_speech response before reading upstream content-type 2026-08-30 12:46:11 -07:00
mateo-berri
1f702f50ad Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Staging moved again mid-recovery; only the ANN201 ratchet conflicted and this branch's tighter limit stands.
2026-08-30 12:43:40 -07:00
mateo-berri
95af93af37 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline
# Conflicts:
#	litellm/proxy/guardrails/guardrail_hooks/unified_guardrail/unified_guardrail.py
2026-08-30 12:38:34 -07:00
mateo-berri
24c5846c75 Merge branch 'litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream 2026-08-30 12:35:39 -07:00
mateo-berri
611750cd11 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_master_key_rotation_blocked
# Conflicts:
#	litellm/proxy/management_endpoints/key_management_endpoints.py
2026-08-30 12:35:21 -07:00
mateo-berri
f0a2a23127 fix(proxy): register SkillsInjectionHook at proxy startup instead of import time 2026-08-30 12:32:55 -07:00
mateo-berri
11e0502239 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Resolves budget-ratchet conflicts by taking staging's tighter limits and reworks the embedding raw-response helpers so the branch stays net-negative on the LIT001/LIT002 ceilings staging lowered: the request methods now return the LegacyAPIResponse and each caller keeps a single dict(headers) conversion.
2026-08-30 12:32:49 -07:00
mateo-berri
f04bfa457a Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema 2026-08-30 12:31:11 -07:00
mateo-berri
4675acf02c Merge remote-tracking branch 'origin/litellm_fix_post_call_policy_pipeline' into litellm_post_call_pipeline_stream_rewrite
# Conflicts:
#	litellm/proxy/policy_engine/pipeline_executor.py
2026-08-30 12:28:54 -07:00
mateo-berri
ce52e39052 fix(gigachat): honor ssl_verify config on router passthrough and type the request body 2026-08-30 12:28:41 -07:00
mateo-berri
739f61df7d test(model_management): drop docstring that restates the serialization path 2026-08-30 12:26:06 -07:00
mateo-berri
c236bcf241 Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:16:13 -07:00
mateo-berri
2420e3f202 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gemini_tts_container 2026-08-30 10:02:43 -07:00
mateo-berri
77ca1a3b31 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_responses_reasoning_drop_params
# Conflicts:
#	tests/test_litellm/llms/openai/responses/test_openai_responses_transformation.py
2026-08-30 10:01:43 -07:00
Michael van den Berg
0ca434a789 Merge remote-tracking branch 'upstream/litellm_internal_staging' into litellm_presidio_new_entities 2026-08-30 14:29:32 +02:00
Devin AI
5c7e6b80c9 test: isolate global MCP registry and pin savings tests to bundled cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 10:18:15 +00:00
siyoon
3822ecc8ff chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertion without adding value.
2026-08-30 14:29:42 +09:00
siyoon
e9f1af8473 chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertions without adding value.
2026-08-30 14:29:28 +09:00
siyoon
e7bfe99cd3 feat(friendli): add zai-org/GLM-5.3 model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $1.40 input / $4.40 output / $0.26 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- text-only (no vision), flagship GLM model
2026-08-30 14:23:36 +09:00
siyoon
3ea4b715ba feat(friendli): add zai-org/GLM-5.3-Flash model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $0.15 input / $0.50 output / $0.03 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- image + video input (native multimodal)
2026-08-30 14:22:41 +09:00
mateo-berri
673d1743a6 fix(policy_engine): apply post_call pipeline text rewrites on streams
Buffered streams governed by post_call policy pipelines now deliver text
rewrites back into the stream per surface (chat SSE, responses SSE,
anthropic messages SSE) instead of rejecting the request with a 400
upfront. Rewrites chain across pipeline steps; tool-call rewrites and
translations without stream write-back still withhold the stream.
2026-08-29 22:12:59 -07:00
mateo-berri
a928c1429e fix(proxy): preserve model table columns on master key rotation 2026-08-29 22:12:35 -07:00
mateo-berri
05e4d2f946 fix(guardrails): scan generateContent systemInstruction text and drop fastapi import from handler tests 2026-08-29 22:11:47 -07:00
mateo-berri
eac5dc10f3 fix(guardrails): apply PUT /guardrails/{id} to the serving worker immediately and reject invalid configs with 422 2026-08-29 22:11:36 -07:00
mateo-berri
b0ce17c755 fix(gigachat): generic env-credential passthrough fallback plus type hardening
- forward unrouted /gigachat/* requests with env credentials like other passthrough providers (the old fallback returned 400 on any request without a routed model, /gigachat/models included)
- fix basedpyright budget breaches across the gigachat provider, common_request_processing, and llm_passthrough_endpoints with real narrowing, no new suppressions
- add regression tests for the fallback target, auth header, and model-less endpoints
2026-08-29 22:08:54 -07:00
mateo-berri
70e2f4e68f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
# Conflicts:
#	litellm/llms/gigachat/chat/transformation.py
2026-08-29 22:08:54 -07:00
mateo-berri
574010a2ce fix(responses): make guardrail input provenance O(n) and guard non-list structured_messages
_input_item_provenance converted every input prefix, so an n-item request paid
for n+1 full conversions. It now converts each item once, glues consecutive
function_call items (plus their trailing-assistant context) into units so the
transform's tool_call merging is reproduced inside the unit conversion, and
verifies the unit concatenation against one full conversion, bailing to the
full-conversion fallback on any mismatch. Messages from multi-item units are
tainted, which keeps parallel tool calls patchable exactly like the old prefix
pass while unpredicted merges fall back safely.

A guardrail handing back a non-list structured_messages payload (the
HiddenLayer v2 evaluation dict) previously fell through the length-mismatch
fallback and 500ed converting the dict's keys as messages. The write-back is
now skipped for non-list payloads, restoring the previous no-write-back
behavior on the Responses surface.

Also refreshes the compresr texts-mirror docstring, which still claimed the
Responses translation cannot round-trip structured_messages.
2026-08-29 22:04:26 -07:00
mateo-berri
bcee01a7a7 fix(policy_engine): merge guardrail metadata writes back on block and modify_response so failure spend records keep guardrail cost and status 2026-08-29 21:52:52 -07:00
mateo-berri
60296cb540 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit 2026-08-29 21:49:13 -07:00
Mateo Wang
5e4b3838aa
Merge pull request #37778 from BerriAI/litellm_decrease_anys_opus5
chore(typing): clear Any seams across 47 files, ratchet basedpyright ceilings -3,302
2026-08-29 21:48:11 -07:00
mateo-berri
43c838f4b9 Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5 2026-08-29 21:41:04 -07:00
mateo-berri
fd72ae830c test(model_management): drive /model/block and /model/unblock through response serialization
The route-level regression test returns a real prisma row from a mocked
update and asserts both routes serialize it to a 200 with the toggled
blocked flag, which is exactly the path that raised AttributeError before
the validator guard. Also binds the loop variable in the e2e poll lambda
(ruff B023).
2026-08-29 21:39:07 -07:00
mateo-berri
a099be02fd fix(guardrails): resolve generateContent routes and async-first passthrough call types
API_ROUTE_TO_CALL_TYPES listed the sync llm_passthrough_route first, so every
call_types[0] consumer resolved /llm_passthrough to a call type with no
guardrail translation handler, and the {model}:generateContent patterns never
matched a concrete route because the placeholder segment carries a literal
suffix the matcher treated as an exact segment. Reorder the passthrough
entries async-first, teach the matcher placeholder-with-suffix segments plus
suffixed multi-segment tails (mirroring FastAPI's {model_name:path}), add the
missing /v1beta generateContent entries, and register a Google GenAI
guardrail translation handler so guardrails actually scan generateContent
requests, responses, and streams.
2026-08-29 21:35:19 -07:00
mateo-berri
badefa395c fix(policy_engine): snapshot guardrail inputs before apply_guardrail so in-place stream rewrites are withheld 2026-08-29 21:30:31 -07:00
mateo-berri
b418ccd738 fix(azure): flatten top-level tool schema combinators on Azure chat completions
Azure's chat completions validator rejects tool parameters carrying a
top-level anyOf/oneOf/allOf for every model family. AzureOpenAIConfig and
the o-series config now flatten them via the shared helper moved to
prompt_templates common_utils. Requests bridged to the Responses API for
gpt-5.4+ with reasoning active keep the union, which that surface accepts
2026-08-29 21:27:57 -07:00
mateo-berri
a23f0fc3c3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit
# Conflicts:
#	litellm/proxy/common_request_processing.py
2026-08-29 21:23:35 -07:00
samzong
26e71ddc54 fix(proxy): serialize model block responses
Signed-off-by: samzong <samzong.lu@gmail.com>
2026-08-29 21:22:17 -07:00
mateo-berri
98a9a7e525 fix(streaming): carry hidden usage on the async fake-stream final chunk
The sync __next__ exhaustion branch stores calculate_total_usage() in the
final chunk's _hidden_params when stream_options is None, but the async
__anext__ sibling branch never did. Converted (fake) streams, like the ones
the Headroom guardrail produces by flipping streaming /v1/responses calls to
non-streaming, are consumed async, so their real usage never reached the
completion-to-responses bridge and it token-counted from scratch, reporting
input_tokens=0. Mirror the sync branch's hidden-usage block into the async
exhaustion branch and add a regression test that async-iterates a
CustomStreamWrapper over a MockResponseIterator and asserts the final chunk
carries the mock response's usage.
2026-08-29 21:12:41 -07:00
mateo-berri
35a375e26f fix(speech): stop vertex gemini tts from dropping response_format in cloud tts param mapping 2026-08-29 21:10:15 -07:00
mateo-berri
b67b44bdaa fix(proxy): map audio_speech errors to their status codes instead of a blanket 500 2026-08-29 21:10:14 -07:00
mateo-berri
45a6b1de23 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline
# Conflicts:
#	litellm/proxy/policy_engine/pipeline_executor.py
2026-08-29 21:01:45 -07:00
mateo-berri
0e27e09fae Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema 2026-08-29 20:57:35 -07:00
Tin Chi Lo
f62aa1b3a8 fix(tests): derive the no-cache-read-rate savings baseline from the model map 2026-08-29 19:35:10 -07:00