Commit graph

49299 commits

Author SHA1 Message Date
samtsai15
1f80f93750 fix(guardrails): carry Anthropic url and file image sources through to guardrails
_image_sources returned source["data"] only. An Anthropic image block has three
shapes (types/llms/anthropic.py:259) and only the base64 one carries "data", so
{"type": "url", "url": ...} yielded nothing and the image never reached any
guardrail at all.

This is not Bedrock-specific. Five guardrails consume
GenericGuardrailAPIInputs["images"] (vigil_guard, custom_code, deepkeep, straiker,
generic_guardrail_api) and every one of them was blind to url sources on
/v1/messages.

base64 now returns a data URI rather than the bare payload. A consumer otherwise
has no way to recover media_type, and an API like Bedrock's ApplyGuardrail needs
the format to build its request.

The file shape stays unresolvable here: the bytes live behind the Files API and
this extractor has no client to fetch them. Documented rather than silently
dropped, so a consumer treating a missing entry as "no image to scan" is a known
gap and not a surprise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:37:51 +08:00
mateo-berri
de1f38820a fix(passthrough): flush interrupted streams on client disconnect and reuse cached gigachat http clients 2026-08-30 13:36:51 -07:00
mateo-berri
14f392bb9b fix(docker): bump wolfi-base for glibc 2.44 and pin apk python to 3.13 2026-08-30 13:11:32 -07:00
mateo-berri
58c5223ab4 refactor(responses): rename write-back helper out of the method's name
The recursion detector in code-quality reads the staticmethod
_written_back_request_fields calling the module-level function of the
same name as a recursive call. Renaming the module-level helper to
_patch_or_convert_request_fields removes the shadowing and describes
what it does: patch changed rows in place, else fall back to full
conversion.
2026-08-30 12:59:58 -07:00
mateo-berri
a5fa8ebfa7 fix(passthrough): keep upstream error body readable for streaming error status mapping 2026-08-30 12:59:18 -07:00
mateo-berri
db1e0717f9 fix(guardrail_translation): assemble responses stream text from delta events for terminal-failure scans 2026-08-30 12:52:03 -07:00
mateo-berri
fb9ec79d7c Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:49:29 -07:00
mateo-berri
4261198b2f Merge branch 'litellm_internal_staging' into litellm_headroom_ccr_streaming_responses 2026-08-30 12:47:33 -07:00
mateo-berri
b8e11a75fa test: use local model cost map in import-isolation subprocess 2026-08-30 12:47:15 -07:00
mateo-berri
99a6dd02af fix(proxy): narrow audio_speech response before reading upstream content-type 2026-08-30 12:46:11 -07:00
mateo-berri
1f702f50ad Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Staging moved again mid-recovery; only the ANN201 ratchet conflicted and this branch's tighter limit stands.
2026-08-30 12:43:40 -07:00
mateo-berri
3004b12e90 refactor(responses): unify guardrail input processing and drop mutable request params
The staging merge tightened the ruff-strict and type-discipline budgets, so the
two `data: dict` parameters the write-back helpers introduced (LIT001) and the
17-branch `process_input_messages` (C901) no longer fit.

Fold the duplicated string/list guardrail flow into one path: a pure
`_extract_guardrail_inputs` builds the guardrail payload, the write-back
helpers become pure functions returning `_RequestFields` (patched input items
plus the resulting instructions value), and the request dict is only mutated
in `process_input_messages` itself. `_apply_guardrail_responses_to_input`
takes Sequence views since it only reads. A non-list `structured_messages`
payload now falls through to the plain texts write-back, matching the
pre-write-back behavior for guardrails that never touch structured messages.

Handler file deltas vs the merge base: LIT001 57 -> 53, LIT002 42 -> 40,
LIT010 28 -> 18, C901 3 -> 3.
2026-08-30 12:40:40 -07:00
mateo-berri
95af93af37 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline
# Conflicts:
#	litellm/proxy/guardrails/guardrail_hooks/unified_guardrail/unified_guardrail.py
2026-08-30 12:38:34 -07:00
mateo-berri
24c5846c75 Merge branch 'litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream 2026-08-30 12:35:39 -07:00
mateo-berri
611750cd11 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_master_key_rotation_blocked
# Conflicts:
#	litellm/proxy/management_endpoints/key_management_endpoints.py
2026-08-30 12:35:21 -07:00
mateo-berri
f0a2a23127 fix(proxy): register SkillsInjectionHook at proxy startup instead of import time 2026-08-30 12:32:55 -07:00
mateo-berri
11e0502239 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Resolves budget-ratchet conflicts by taking staging's tighter limits and reworks the embedding raw-response helpers so the branch stays net-negative on the LIT001/LIT002 ceilings staging lowered: the request methods now return the LegacyAPIResponse and each caller keeps a single dict(headers) conversion.
2026-08-30 12:32:49 -07:00
mateo-berri
f04bfa457a Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema 2026-08-30 12:31:11 -07:00
mateo-berri
4675acf02c Merge remote-tracking branch 'origin/litellm_fix_post_call_policy_pipeline' into litellm_post_call_pipeline_stream_rewrite
# Conflicts:
#	litellm/proxy/policy_engine/pipeline_executor.py
2026-08-30 12:28:54 -07:00
mateo-berri
ce52e39052 fix(gigachat): honor ssl_verify config on router passthrough and type the request body 2026-08-30 12:28:41 -07:00
mateo-berri
739f61df7d test(model_management): drop docstring that restates the serialization path 2026-08-30 12:26:06 -07:00
mateo-berri
c236bcf241 Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:16:13 -07:00
Mateo Wang
4ba8517134
Merge pull request #38722 from BerriAI/litellm_bedrock_guardrail_stream_audit
feat(bedrock): honor streaming buffer/sampling config for unbuffered post_call scans
2026-08-30 10:17:11 -07:00
Mateo Wang
8a156ed42d
Merge pull request #36722 from BerriAI/litellm_decrease_anys_fable6
chore(typing): clear 1.2k basedpyright Any errors across 16 hotspot files
2026-08-30 10:09:02 -07:00
mateo-berri
2420e3f202 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gemini_tts_container 2026-08-30 10:02:43 -07:00
mateo-berri
fbf7644676 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_fable6
# Conflicts:
#	basedpyright-code-budget.json
#	litellm/integrations/websearch_interception/handler.py
#	litellm/proxy/response_polling/background_streaming.py
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-30 10:02:06 -07:00
mateo-berri
b6bd749c02 fix(ui): withhold team-scoped model writes from view-only sessions too
The route-level RBAC in litellm/proxy/auth/route_checks.py 403s
/model/new, /model/update, and /model/delete for proxy_admin_viewer on
the session role alone, before ModelManagementAuthChecks' team-admin
carve-out can run. A view-only session therefore gets no model write
affordance, team admin or not.
2026-08-30 10:01:44 -07:00
mateo-berri
77ca1a3b31 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_responses_reasoning_drop_params
# Conflicts:
#	tests/test_litellm/llms/openai/responses/test_openai_responses_transformation.py
2026-08-30 10:01:43 -07:00
Michael van den Berg
0ca434a789 Merge remote-tracking branch 'upstream/litellm_internal_staging' into litellm_presidio_new_entities 2026-08-30 14:29:32 +02:00
Devin AI
5c7e6b80c9 test: isolate global MCP registry and pin savings tests to bundled cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 10:18:15 +00:00
Devin AI
693279afb4 chore(techdebt): clear fresh debt from the 2026-08-29 window
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 07:57:16 +00:00
Devin AI
36c53e1288 chore(ui): regenerate schema.d.ts for updated endpoint descriptions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 07:46:45 +00:00
Devin AI
4fb1440747 docs(proxy): clarify spend semantics on /v2/user/info and /user/daily/activity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 07:40:56 +00:00
siyoon
b346dd414b fix(scripts): use PEP 604 union syntax for Optional[str]
Ruff UP045 flags Optional[str] as legacy typing; the repo lint baseline
expects the modern str | None form for new/changed annotations.
2026-08-30 14:52:35 +09:00
siyoon
2157d04f8e Merge remote-tracking branch 'upstream/litellm_internal_staging' into feat/friendli-model-metadata-sync 2026-08-30 14:51:42 +09:00
siyoon
3822ecc8ff chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertion without adding value.
2026-08-30 14:29:42 +09:00
siyoon
e9f1af8473 chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertions without adding value.
2026-08-30 14:29:28 +09:00
siyoon
e7bfe99cd3 feat(friendli): add zai-org/GLM-5.3 model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $1.40 input / $4.40 output / $0.26 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- text-only (no vision), flagship GLM model
2026-08-30 14:23:36 +09:00
siyoon
3ea4b715ba feat(friendli): add zai-org/GLM-5.3-Flash model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $0.15 input / $0.50 output / $0.03 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- image + video input (native multimodal)
2026-08-30 14:22:41 +09:00
mateo-berri
673d1743a6 fix(policy_engine): apply post_call pipeline text rewrites on streams
Buffered streams governed by post_call policy pipelines now deliver text
rewrites back into the stream per surface (chat SSE, responses SSE,
anthropic messages SSE) instead of rejecting the request with a 400
upfront. Rewrites chain across pipeline steps; tool-call rewrites and
translations without stream write-back still withhold the stream.
2026-08-29 22:12:59 -07:00
mateo-berri
a928c1429e fix(proxy): preserve model table columns on master key rotation 2026-08-29 22:12:35 -07:00
mateo-berri
05e4d2f946 fix(guardrails): scan generateContent systemInstruction text and drop fastapi import from handler tests 2026-08-29 22:11:47 -07:00
mateo-berri
eac5dc10f3 fix(guardrails): apply PUT /guardrails/{id} to the serving worker immediately and reject invalid configs with 422 2026-08-29 22:11:36 -07:00
mateo-berri
b0ce17c755 fix(gigachat): generic env-credential passthrough fallback plus type hardening
- forward unrouted /gigachat/* requests with env credentials like other passthrough providers (the old fallback returned 400 on any request without a routed model, /gigachat/models included)
- fix basedpyright budget breaches across the gigachat provider, common_request_processing, and llm_passthrough_endpoints with real narrowing, no new suppressions
- add regression tests for the fallback target, auth header, and model-less endpoints
2026-08-29 22:08:54 -07:00
mateo-berri
70e2f4e68f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
# Conflicts:
#	litellm/llms/gigachat/chat/transformation.py
2026-08-29 22:08:54 -07:00
mateo-berri
574010a2ce fix(responses): make guardrail input provenance O(n) and guard non-list structured_messages
_input_item_provenance converted every input prefix, so an n-item request paid
for n+1 full conversions. It now converts each item once, glues consecutive
function_call items (plus their trailing-assistant context) into units so the
transform's tool_call merging is reproduced inside the unit conversion, and
verifies the unit concatenation against one full conversion, bailing to the
full-conversion fallback on any mismatch. Messages from multi-item units are
tainted, which keeps parallel tool calls patchable exactly like the old prefix
pass while unpredicted merges fall back safely.

A guardrail handing back a non-list structured_messages payload (the
HiddenLayer v2 evaluation dict) previously fell through the length-mismatch
fallback and 500ed converting the dict's keys as messages. The write-back is
now skipped for non-list payloads, restoring the previous no-write-back
behavior on the Responses surface.

Also refreshes the compresr texts-mirror docstring, which still claimed the
Responses translation cannot round-trip structured_messages.
2026-08-29 22:04:26 -07:00
mateo-berri
bcee01a7a7 fix(policy_engine): merge guardrail metadata writes back on block and modify_response so failure spend records keep guardrail cost and status 2026-08-29 21:52:52 -07:00
mateo-berri
60296cb540 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit 2026-08-29 21:49:13 -07:00
Mateo Wang
5e4b3838aa
Merge pull request #37778 from BerriAI/litellm_decrease_anys_opus5
chore(typing): clear Any seams across 47 files, ratchet basedpyright ceilings -3,302
2026-08-29 21:48:11 -07:00
mateo-berri
667f761f1d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_fable6 2026-08-29 21:43:56 -07:00