Commit graph

15348 commits

Author SHA1 Message Date
Devin AI
dd031f1036 fix(ci): parse paginated gh api output without splitting on unicode line breaks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 09:42:07 +00:00
Devin AI
b81bca2696 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260830 2026-08-31 09:18:31 +00:00
Cursor Agent
43c64e2471
fix(cli): keep an existing ENABLE_TOOL_SEARCH value
Default remains true so lite claude turns tool search back on through
a proxy. An explicit false or auto in the env or settings is left alone

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-08-31 06:00:45 +00:00
Cursor Agent
cafdfda8ba
feat(cli): set ENABLE_TOOL_SEARCH=true for lite claude
Claude Code turns tool search off when ANTHROPIC_BASE_URL is a proxy.
lite claude, lite up, login --config-claude, and autoroute now force
ENABLE_TOOL_SEARCH=true so MCP tools stay deferred through the proxy

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-08-31 05:52:17 +00:00
feng.tsai
bb51c121cf docs: reference the source union by type instead of a line number
The line number went stale when the base moved.
2026-08-31 12:21:24 +08:00
samtsai15
0e7562dbc6 test(guardrails): cover every Anthropic image source shape in the extractor's own suite
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.

Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].

The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.

Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:37:51 +08:00
mateo-berri
de1f38820a fix(passthrough): flush interrupted streams on client disconnect and reuse cached gigachat http clients 2026-08-30 13:36:51 -07:00
mateo-berri
a5fa8ebfa7 fix(passthrough): keep upstream error body readable for streaming error status mapping 2026-08-30 12:59:18 -07:00
mateo-berri
db1e0717f9 fix(guardrail_translation): assemble responses stream text from delta events for terminal-failure scans 2026-08-30 12:52:03 -07:00
mateo-berri
fb9ec79d7c Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:49:29 -07:00
mateo-berri
4261198b2f Merge branch 'litellm_internal_staging' into litellm_headroom_ccr_streaming_responses 2026-08-30 12:47:33 -07:00
mateo-berri
99a6dd02af fix(proxy): narrow audio_speech response before reading upstream content-type 2026-08-30 12:46:11 -07:00
mateo-berri
1f702f50ad Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Staging moved again mid-recovery; only the ANN201 ratchet conflicted and this branch's tighter limit stands.
2026-08-30 12:43:40 -07:00
mateo-berri
24c5846c75 Merge branch 'litellm_internal_staging' into litellm_fix_bedrock_buffered_responses_stream 2026-08-30 12:35:39 -07:00
mateo-berri
611750cd11 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_master_key_rotation_blocked
# Conflicts:
#	litellm/proxy/management_endpoints/key_management_endpoints.py
2026-08-30 12:35:21 -07:00
mateo-berri
11e0502239 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_openai_embedding_encoding_format_omit
Resolves budget-ratchet conflicts by taking staging's tighter limits and reworks the embedding raw-response helpers so the branch stays net-negative on the LIT001/LIT002 ceilings staging lowered: the request methods now return the LegacyAPIResponse and each caller keeps a single dict(headers) conversion.
2026-08-30 12:32:49 -07:00
mateo-berri
f04bfa457a Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema 2026-08-30 12:31:11 -07:00
mateo-berri
ce52e39052 fix(gigachat): honor ssl_verify config on router passthrough and type the request body 2026-08-30 12:28:41 -07:00
mateo-berri
739f61df7d test(model_management): drop docstring that restates the serialization path 2026-08-30 12:26:06 -07:00
mateo-berri
c236bcf241 Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5_r2 2026-08-30 12:16:13 -07:00
mateo-berri
2420e3f202 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gemini_tts_container 2026-08-30 10:02:43 -07:00
Devin AI
5c7e6b80c9 test: isolate global MCP registry and pin savings tests to bundled cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 10:18:15 +00:00
siyoon
3822ecc8ff chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertion without adding value.
2026-08-30 14:29:42 +09:00
siyoon
e9f1af8473 chore(tests): drop redundant capability comment
Per greptile review + CLAUDE.md comment policy: the comment restated
the immediately following assertions without adding value.
2026-08-30 14:29:28 +09:00
siyoon
e7bfe99cd3 feat(friendli): add zai-org/GLM-5.3 model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $1.40 input / $4.40 output / $0.26 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- text-only (no vision), flagship GLM model
2026-08-30 14:23:36 +09:00
siyoon
3ea4b715ba feat(friendli): add zai-org/GLM-5.3-Flash model pricing
Per https://api.friendli.ai/serverless/v1/models:
- $0.15 input / $0.50 output / $0.03 cached input per MTok
- 1M context, 1M max output, reasoning with effort low/high/max
  (per HF chat_template.jinja: low/high honored, anything else -> max)
- tool calling, parallel tool calls, structured output, prompt caching
- image + video input (native multimodal)
2026-08-30 14:22:41 +09:00
mateo-berri
a928c1429e fix(proxy): preserve model table columns on master key rotation 2026-08-29 22:12:35 -07:00
mateo-berri
eac5dc10f3 fix(guardrails): apply PUT /guardrails/{id} to the serving worker immediately and reject invalid configs with 422 2026-08-29 22:11:36 -07:00
mateo-berri
b0ce17c755 fix(gigachat): generic env-credential passthrough fallback plus type hardening
- forward unrouted /gigachat/* requests with env credentials like other passthrough providers (the old fallback returned 400 on any request without a routed model, /gigachat/models included)
- fix basedpyright budget breaches across the gigachat provider, common_request_processing, and llm_passthrough_endpoints with real narrowing, no new suppressions
- add regression tests for the fallback target, auth header, and model-less endpoints
2026-08-29 22:08:54 -07:00
mateo-berri
70e2f4e68f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
# Conflicts:
#	litellm/llms/gigachat/chat/transformation.py
2026-08-29 22:08:54 -07:00
mateo-berri
574010a2ce fix(responses): make guardrail input provenance O(n) and guard non-list structured_messages
_input_item_provenance converted every input prefix, so an n-item request paid
for n+1 full conversions. It now converts each item once, glues consecutive
function_call items (plus their trailing-assistant context) into units so the
transform's tool_call merging is reproduced inside the unit conversion, and
verifies the unit concatenation against one full conversion, bailing to the
full-conversion fallback on any mismatch. Messages from multi-item units are
tainted, which keeps parallel tool calls patchable exactly like the old prefix
pass while unpredicted merges fall back safely.

A guardrail handing back a non-list structured_messages payload (the
HiddenLayer v2 evaluation dict) previously fell through the length-mismatch
fallback and 500ed converting the dict's keys as messages. The write-back is
now skipped for non-list payloads, restoring the previous no-write-back
behavior on the Responses surface.

Also refreshes the compresr texts-mirror docstring, which still claimed the
Responses translation cannot round-trip structured_messages.
2026-08-29 22:04:26 -07:00
mateo-berri
60296cb540 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit 2026-08-29 21:49:13 -07:00
Mateo Wang
5e4b3838aa
Merge pull request #37778 from BerriAI/litellm_decrease_anys_opus5
chore(typing): clear Any seams across 47 files, ratchet basedpyright ceilings -3,302
2026-08-29 21:48:11 -07:00
mateo-berri
43c838f4b9 Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5 2026-08-29 21:41:04 -07:00
mateo-berri
fd72ae830c test(model_management): drive /model/block and /model/unblock through response serialization
The route-level regression test returns a real prisma row from a mocked
update and asserts both routes serialize it to a 200 with the toggled
blocked flag, which is exactly the path that raised AttributeError before
the validator guard. Also binds the loop variable in the e2e poll lambda
(ruff B023).
2026-08-29 21:39:07 -07:00
mateo-berri
b418ccd738 fix(azure): flatten top-level tool schema combinators on Azure chat completions
Azure's chat completions validator rejects tool parameters carrying a
top-level anyOf/oneOf/allOf for every model family. AzureOpenAIConfig and
the o-series config now flatten them via the shared helper moved to
prompt_templates common_utils. Requests bridged to the Responses API for
gpt-5.4+ with reasoning active keep the union, which that surface accepts
2026-08-29 21:27:57 -07:00
mateo-berri
a23f0fc3c3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit
# Conflicts:
#	litellm/proxy/common_request_processing.py
2026-08-29 21:23:35 -07:00
samzong
26e71ddc54 fix(proxy): serialize model block responses
Signed-off-by: samzong <samzong.lu@gmail.com>
2026-08-29 21:22:17 -07:00
mateo-berri
98a9a7e525 fix(streaming): carry hidden usage on the async fake-stream final chunk
The sync __next__ exhaustion branch stores calculate_total_usage() in the
final chunk's _hidden_params when stream_options is None, but the async
__anext__ sibling branch never did. Converted (fake) streams, like the ones
the Headroom guardrail produces by flipping streaming /v1/responses calls to
non-streaming, are consumed async, so their real usage never reached the
completion-to-responses bridge and it token-counted from scratch, reporting
input_tokens=0. Mirror the sync branch's hidden-usage block into the async
exhaustion branch and add a regression test that async-iterates a
CustomStreamWrapper over a MockResponseIterator and asserts the final chunk
carries the mock response's usage.
2026-08-29 21:12:41 -07:00
mateo-berri
35a375e26f fix(speech): stop vertex gemini tts from dropping response_format in cloud tts param mapping 2026-08-29 21:10:15 -07:00
mateo-berri
b67b44bdaa fix(proxy): map audio_speech errors to their status codes instead of a blanket 500 2026-08-29 21:10:14 -07:00
mateo-berri
0e27e09fae Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema 2026-08-29 20:57:35 -07:00
Tin Chi Lo
f62aa1b3a8 fix(tests): derive the no-cache-read-rate savings baseline from the model map 2026-08-29 19:35:10 -07:00
yucheng-berri
d44d281d1d
fix(proxy): emit timing headers and overhead for /v1/messages and /v1/responses (#38840) 2026-08-29 18:11:58 -07:00
yassin
2f958f2187 fix(managed_files): match provider-format ids against model_object_id and flat_model_file_ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-30 00:51:06 +00:00
yuneng-jiang
6b33d17563
Merge pull request #38850 from BerriAI/litellm_e2e_retry_transient_upstream
test(e2e): retry upstream-saturation failures in the claude CLI driver
2026-08-29 17:48:04 -07:00
yuneng-jiang
df848d85ff
Merge pull request #38833 from BerriAI/litellm_deflake_reliability_fallbacks
test(e2e): stop the reliability fallback tests flaking on gpt-5.5's reasoning budget
2026-08-29 17:44:30 -07:00
yucheng-berri
3c2fa5fafb
fix(otel/v2): detach credential-routed tenant spans into their own trace (#38847)
* fix(otel/v2): detach credential-routed tenant spans into their own trace

Multi-tenant OTel v2 routes a team or key's LLM-call span to that tenant's
own vendor account (New Relic, Arize, Langfuse, Weave) via dynamic OTLP
credential headers, while the request-root, auth, and db spans stay on the
operator's default backend. The span was still parented into the request
trace, so the tenant account received a child whose parent it never got,
and New Relic rendered it as a fragmented trace with a missing parent.

Detach a credential-routed span the same way a project-routed (Phoenix)
span already detaches: root a fresh trace in the tenant account and link
back to the request trace for correlation. Service-name routing keeps
parenting, since it only relabels service.name on the same operator
backend where the parent is present.

Guard the detach on the callback actually owning an OTLP exporter the
credentials can reach: a callback owning only a console or in_memory
exporter has nowhere to stamp them, so the span would export to the
default backend unchanged and detaching would orphan it on the very
backend that holds its parent. In that case warn once and keep the
default tracer.

* fix(otel/v2): derive tenant-route routability from resolved exporter transport

A denylist classified an owned exporter as routable whenever its kind was
not console/in_memory, so a typo'd or unavailable kind (e.g. "otlp",
"grcp") passed the check while _exporter_from_spec falls it back to a
header-ignoring console exporter. Detaching such a span would root a fresh
trace that only ever reaches the operator console, never the tenant
backend, orphaning it on both sides.

Route on a shared exporter_transport() predicate that resolves the kind the
same way _exporter_from_spec builds it (registered factories + otlp_http
aliases -> http, otlp_grpc aliases -> grpc, else headerless), so an
unresolvable kind is headerless and stays parented. Fixes the same latent
gap in project routability.
2026-08-29 17:41:25 -07:00
ryan-crabbe-berri
9da0b30888
Merge pull request #38843 from BerriAI/litellm_mag_budget_ui
feat(ui): set a model access group's shared budget from the dashboard
2026-08-29 17:27:17 -07:00
Mateo Wang
ff2f06e37f
Merge pull request #38444 from BerriAI/litellm_mcp_connector_bulk_import
feat(mcp): bulk-import Anthropic MCP connectors via API and admin UI
2026-08-29 17:21:57 -07:00