Commit graph

48347 commits

Author SHA1 Message Date
mateo-berri
6966a33150 test(vector-stores): type the pre-call hook regression tests without Any 2026-09-02 22:09:58 -07:00
mateo-berri
06e60e08d2 ci(ui): run the UI build check through the image's ui-builder stage
The build-ui check compiled the dashboard from a full checkout, so any
import reaching above ui/litellm-dashboard/ resolved there and only broke
inside the images, where the stage copies the dashboard tree alone.
Building the stage itself puts the check on the same file boundary the
shipped images use.
2026-09-02 22:09:55 -07:00
mateo-berri
59f9d3a799 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_fix_failing_request_slowdown
# Conflicts:
#	basedpyright-code-budget.json
#	type-discipline-budget.json
2026-09-02 22:00:30 -07:00
mateo-berri
cf958c0e6f ci(rust): build and test the ai-gateway server feature
litellm-ai-gateway's server feature is off by default and nothing in the workspace turns it on, so the workspace clippy and test steps never compiled src/auth, src/routes, src/state, src/realtime or the gateway binary. 43 tests ran instead of 57.

Adds the two steps CLAUDE.md already documents as the local gate, and fixes the three collapsible_if violations that had accumulated behind the flag.
2026-09-02 21:56:23 -07:00
mateo-berri
3ea61c23c7 fix(vector-stores): survive a failing vector store search in the chat completions hook
One unreachable vector store used to wipe out every store's context on a
chat completion carrying vector_store_ids: the search raised, the blanket
handler returned the original messages, and the request answered with no
retrieved context at all. Each store's search now has its own handler that
warns with the vector store id and moves on to the next store.

The same loop appended every store's results to the original messages
instead of the running copy, so with two healthy stores only the last one
reached the model. It now chains through modified_messages.

The Router is injected through a ProxyRuntime protocol instead of an
in-function litellm.proxy.proxy_server import, so the hook's routing can
be driven in tests without touching proxy globals.
2026-09-02 21:55:16 -07:00
mateo-berri
af15f87c5a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_search_results_with_guardrails 2026-09-02 21:51:47 -07:00
mateo-berri
f62e87d28c fix(ui): read the preset catalog through the shared mock in the lib test 2026-09-02 21:42:47 -07:00
mateo-berri
7c87451ead chore(router): document the breadcrumb write-back and ratchet lint budgets 2026-09-02 21:40:25 -07:00
mateo-berri
a581399027 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
# Conflicts:
#	basedpyright-code-budget.json
2026-09-02 21:34:20 -07:00
mateo-berri
7f7e0d5517 fix(vector-store): embed through the SDK when the Router does not serve the query embedding model
The Router executor only routed a query embedding when the vector store
carried extra embedding configuration, so a store registered with no
embedding model at all always went to the Router and 500'd on the
s3_vectors default text-embedding-3-small when no deployment served it.
Route on whether the Router serves the model, which is the rule the
executor had before, and keep the request metadata on the SDK fallback so
the embedding stays attributed either way.
2026-09-02 21:32:00 -07:00
mateo-berri
cd296814be Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_containers_error_passthrough_pagination 2026-09-02 21:26:40 -07:00
mateo-berri
425c8d37fc fix(containers): page upstream until a non-admin container list fills its limit
Forwarding limit to OpenAI made the ownership filter cut the page down after
the fact, so a key that owned an older container got an empty first page and
its cursor never moved. Non-admin lists now walk upstream pages of 100 until
they have enough owned containers (or five pages), trim to the requested
limit, and report first_id, last_id and has_more off what the caller keeps.

Also assigns tests/test_litellm/proxy/container_endpoints to a CI shard.
2026-09-02 21:26:34 -07:00
mateo-berri
99e62d5fb6 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_bedrock_bearer_skip_sigv4_chain 2026-09-02 21:14:06 -07:00
mateo-berri
f9f32b49a6 fix(tests): qualify traced function names on Python 3.10 2026-09-02 21:13:15 -07:00
Mateo Wang
ff17e8b987
Merge pull request #39472 from BerriAI/litellm_fix_kb_hook_test_embedding_executor
test(vector-store): accept embedding_executor in the Bedrock KB hook fake handler
2026-09-02 21:11:24 -07:00
mateo-berri
32f71950ec Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci 2026-09-02 21:06:02 -07:00
mateo-berri
9f4fe3144a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_prisma_timeout_killpg 2026-09-02 21:01:11 -07:00
mateo-berri
7ea862d020 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_rag_query_store_credentials 2026-09-02 20:55:50 -07:00
mateo-berri
7bc2d0b06e fix(router): keep retry breadcrumbs per request and out of the request snapshot
Retry breadcrumbs were appended to one list owned by the Router and shared by
every request, and each breadcrumb copied the whole kwargs including the proxy's
snapshot of the inbound request. That snapshot's body aliases the live request
metadata, breadcrumbs included, so every new breadcrumb nested all the earlier
ones inside itself. Memory stayed small because these are shared references, but
under --detailed_debug the repr of that structure expands, so one debug line grew
from 10k to 219M characters over 14 failing requests and the proxy stopped
answering.

Breadcrumbs now accumulate in the metadata of the request that produced them, the
request snapshot is excluded from a breadcrumb, and the cap of the last 4 failed
attempts applies per request.
2026-09-02 20:55:41 -07:00
ishaan-berri
e058aa68c4
test: add mistral ocr transformation parity coverage (#39482)
* test: cover mistral ocr transformation parity

Co-Authored-By: Claude Code <noreply@anthropic.com>

* test: map mistral ocr parity contracts

Co-Authored-By: Claude Code <noreply@anthropic.com>

---------

Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-02 20:40:00 -07:00
devin-ai-integration[bot]
291d02f8aa
fix(mcp): never exchange the LiteLLM virtual key as the upstream subject token (#39446) 2026-09-02 20:32:43 -07:00
Tin Chi Lo
fcc9b813af fix(ui): resolve the preset catalog relative to the mock, not cwd
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HqEPCNLDrssxsuezhAaL4j
2026-09-02 20:01:35 -07:00
Tin Chi Lo
099d26204f fix(ui): read the preset catalog at runtime in the vitest mock
The autoRouterPresets mock imported litellm/proxy/public_endpoints/autorouter_presets.json
as a module. That path sits outside ui/litellm-dashboard, the only directory the UI
Dockerfile copies, so `next build` type-checking inside the image failed with
"Cannot find module" and the ui-image job went red on every PR that touched an
image-scan path. Read the file with fs at runtime instead; vitest still derives
expectations from the real bundled catalog.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HqEPCNLDrssxsuezhAaL4j
2026-09-02 19:54:09 -07:00
yassin
9ecc953618 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_agent_mcp_grants
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 02:51:45 +00:00
ishaan-berri
bcd3e2d94d
feat(rust-python-harness): wire existing e2e SDK tests into the matrix (#39463)
Adds chat_completions and transcription as SDK function columns, backed
by the existing rust_bridge test files. Adds a fourth strategy folder,
existing_e2e_test_sdk, that points at already-existing live-API SDK
tests (tests/ocr_tests/ as a whole folder, plus chat completion and
Whisper transcription tests) instead of writing new parity tests.
Extends selector_matches_node with trailing-slash folder selectors so
a whole test folder can back one matrix cell.
2026-09-03 02:50:08 +00:00
yucheng-berri
291e84e565
feat(datadog_llm_obs): cost tag dimensions, router decision fields, reasoning token metric, redaction gating (#39402)
* feat(datadog_llm_obs): cost tag dimensions, router decision fields, reasoning token metric, redaction gating

* test(datadog_llm_obs): satisfy test quality gate

* fix: forward integer parent_id as its string form

* fix(datadog): sanitize redacted message roles

* fix(datadog): keep the A2A agent role on redacted spans

* fix(datadog): merge current staging budget

* style(datadog): format redaction tests

* fix(datadog): handle malformed redacted roles

* test(datadog): put the test quality suppression on the reported line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 19:46:09 -07:00
tin-berri
64e45a069d
feat(complexity_router): opt-in modality override of a kept session-affinity pin (#39454)
A session pinned to a text-only model failed every image turn with a provider 400, because the modality gate exempts a kept pin by cause. Add modality_pin_override so that exemption is conditional: the image turn is re-placed on a capable model for that request only, reported as cause modality_pin_override, and the stored pin is left untouched so the next text turn replays it.

The pin write on the replay path already happens upstream of the gate and stores the session's own model, so pin survival is structural rather than bookkeeping. The new cause joins the non-pinnable set. Default off at every layer.
2026-09-02 19:36:09 -07:00
mateo-berri
f77b3b2b52 refactor(s3_vectors): embed search queries through the shared vector store executor
S3 Vectors now subclasses BaseQueryEmbeddingVectorStoreConfig, so its query
embedding runs through the Router executor with the request metadata instead
of a private router lookup. embedding_model stays accepted as an alias of
litellm_embedding_model. The router kwarg is gone from the search handler and
every provider transform now that nothing but the executor fallback read it.
2026-09-02 19:26:31 -07:00
ishaan-berri
a5639b8e2a
test: add OCR python-to-rust test parity ledger (WIP) (#39434)
* test: add OCR python-to-rust test parity ledger

* feat: add ledger loader for OCR test parity data

* feat: add drift audit for OCR test parity ledger (WIP, untested)

* fix: correct drift in OCR test-parity ledger

Two entries referenced a typo'd Python test name, four duplicated entries already tracked under TestProxySecurityGuard, five real Python tests in test_rust_bridge.py were untracked, and three real Rust custom_logger tests were missing from rust_only_tests. Found by running validate_ledger.py's audit against the live repo.

* test: add regression coverage for the OCR ledger and audit script

Covers schema validation, AST/regex test enumeration, drift detection on both the Python and Rust sides, LedgerDriftError content, and a live-repo clean-audit guard against future drift.

* test: simplify ledger test to one drift-guard assertion

Replace the ledger-internals unit tests with a single test that runs the real audit against the live repo and asserts every OCR test is accounted for (mapped, unmapped-with-reason, or rust_only), printing the exact diff on failure.

* chore: move OCR test-parity ledger to core/ocr

validate_sub_methods/ mixes strategy-catalog metadata with the ledger. Ledger data belongs under a per-function core/<function>/ path instead.

* fix: point LEDGER_PATH at the new core/ocr location
2026-09-02 19:22:52 -07:00
mateo-berri
f81928f7ae test(vector-store): accept embedding_executor in the Bedrock KB hook fake handler 2026-09-02 19:22:12 -07:00
mateo-berri
ad0f0af192 fix(helm): normalize ingress.extraPaths path types for ingress-nginx too
A dotted extraPaths entry kept its requested Exact or Prefix type under
ingress.controller=nginx, so the admission webhook rejected the render the
option exists to avoid, and the duplicate check compared the raw type against
the built-in paths' normalized one, letting a repeated /favicon.ico through.
Extra paths now go through the same controller normalization before both the
duplicate check and the render.
2026-09-02 19:15:41 -07:00
yucheng-berri
7a81ae98e6
fix(model_armor): handle Anthropic Messages and Responses streams in post_call (#39181)
* fix(model_armor): handle Anthropic Messages and Responses streams in post_call

The post_call streaming hook buffered every chunk and fed it to
stream_chunk_builder, which only understands chat-completion deltas.
/v1/messages streams raw Anthropic SSE bytes and /v1/responses streams
typed Responses events, so both raised litellm.APIError and surfaced to
the client as a 500 on every streamed request.

Assemble each surface with its own reader, frame guardrail failures as
terminal items in that surface's wire format, and pass the stream
through unscanned when it cannot be assembled instead of raising.

* fix(model_armor): classify the stream surface and fail closed when it cannot be assembled

Decide the wire format explicitly instead of inferring it from a boolean pair, so an
opaque raw SSE stream (the Google :streamGenerateContent route) is never refused in
Anthropic framing, and a stream that cannot be assembled is blocked rather than
released unscanned unless fail_on_error is disabled.

Also scan Responses tool-call arguments, read the body only off a terminal Responses
event, and record the applied guardrail on the fail-closed path.

* test(model_armor): pin the error-only stream predicate against content-carrying streams

is_sse_error_stream decides whether a buffered stream is forwarded to the client
untouched, so a stream that still carries content must not qualify: the frames-only
join drops typed chunks, an empty stream is not a refusal, and a content event may
carry an empty error field.

* fix(model_armor): let a streamed de-identify match mask instead of blocking

A de-identify template reports MATCH_FOUND for every redaction it makes. The
streaming block check omitted allow_sanitization, so with mask_response_content
enabled that match read as a refusal and the client got a 400 where the
non-streaming sibling returned the redacted text. Pass the flag through, as the
non-streaming hook already does, and stamp the logged status from the same
decision so the spend row agrees with what the client received.

Also drop Any from the chat-completion assembler's parameter; stream_chunk_builder
takes a bare list, so list[object] carries the mutability requirement without
erasing the element type.

* fix(model_armor): fail closed when a streamed de-identify match cannot be applied

Allowing sanitization past the streaming block check is a promise to apply the
redaction Model Armor asked for. Two paths broke that promise and released the
buffered original instead: a match that comes back with no sanitized text, and a
surface with no assembled body to rewrite.

The outcome is now resolved once, before it is recorded, so the status stamped on
request metadata agrees with what the client receives rather than reporting the
success the block check alone would have implied.

* fix: scan the deltas when a Responses stream ends without a body

response.failed and response.incomplete are terminal events like
response.completed, but a turn that broke mid-generation reports an empty
output while the deltas ahead of it already spelled the answer out to the
client. Reading only the terminal body found nothing to scan there, and the
empty-content shortcut then forwarded every buffered delta past the guardrail.

Fall back to the text the delta events carry whenever a Responses stream
assembles to nothing.

* fix: read the Responses delta event types off the event enum

The hand-listed set left out response.mcp_call_arguments.delta, so a turn that
streamed only MCP tool arguments and then reported an empty body still took the
no-content shortcut and forwarded those chunks unscanned.

Deriving the set from ResponsesAPIStreamEvents keeps it complete as the enum
grows, and the str guard in the reader already covers any event whose delta is
not text.

* fix(model_armor): scan responses deltas alongside the terminal body

A /v1/responses stream spells out reasoning summaries and tool-call arguments in
delta events that its terminal body never repeats, so scanning the body alone
handed every summary delta to the client unscanned whenever the body carried text.

* fix(model_armor): scan responses delta fields apart from each other

A Responses turn spells out its reasoning summary, its visible answer and its tool-call
arguments in separate delta events. Joining every delta into one string let a finding form
across the boundary between two fields that each carry nothing to find, so a safe stream
could be blocked. Group the deltas by the field they belong to, join a field's own deltas
as they streamed, and keep the fields apart.

* fix(model_armor): scan each responses field once, not twice

Separating delta fields stopped the terminal body from matching the delta text, so a turn
with two visible fields sent Model Armor both copies. Only the delta fields the body does not
already carry are appended now.

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-02 19:15:14 -07:00
devin-ai-integration[bot]
8065ede40b
test(guardrails): expect the deduped end-of-stream scan in crowdstrike cadence test (#39467)
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 02:11:08 +00:00
mateo-berri
f0f5f1ec78 test(proxy-extras): wrap an over-long monkeypatch line 2026-09-02 19:03:17 -07:00
mateo-berri
c15f4e066f fix(rag): let the managed store's params win over caller kwargs on the search call 2026-09-02 18:59:28 -07:00
yucheng-berri
e0e249225b
feat(azure): support credential chain for storage (#39229)
* feat(azure): support credential chain for storage

* test(azure): clarify credential seam suppressions

* fix(azure): read chain tokens in a worker thread

The credential chain walk (IMDS probe, CLI subprocess) is blocking I/O,
so reading the provider inline in async set_valid_azure_ad_token stalls
every request on the worker's event loop
2026-09-02 18:55:22 -07:00
mateo-berri
e03eb6961c fix(deps): bump uvloop to 0.22.1 so the proxy boots on Python 3.14 2026-09-02 18:53:56 -07:00
mateo-berri
872e115295 fix(containers): pass upstream error status through and forward list pagination params
The container retrieve, list, delete, create and file routes validated the
provider's error body against the success model, so a deleted or unknown
container and a rejected API key surfaced as 500 pydantic errors instead of
the upstream 404 or 401. The handlers now raise the provider error class with
the upstream status and message before transforming the response.

GET /v1/containers dropped after, limit and order before calling the
provider, and GET /v1/containers/{id}/files dropped the same three, so
paginated list calls ignored their pagination arguments. Both routes now
forward their declared query params.
2026-09-02 18:46:00 -07:00
mateo-berri
b8680120ce fix(helm): render ingress-nginx compatible path types via ingress.controller
ingress-nginx's admission webhook (strict-validate-path-type, on by default
from v1.12.0 until v1.12.6 / v1.13.2 allowed dots again) rejects the chart's
/favicon.ico Exact and /eu.assemblyai Prefix rules, so helm install fails on
any cluster it fronts. A new ingress.controller value (alb, the default, or
nginx) renders dotted built-in paths as ImplementationSpecific under nginx,
which serves them as plain prefix locations, and drops the ALB-only /*.txt
wildcard rule there. The default render is unchanged
2026-09-02 18:45:18 -07:00
mateo-berri
45988143bc Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_fix_agent_mcp_grants
# Conflicts:
#	type-discipline-budget.json
2026-09-02 18:42:33 -07:00
mateo-berri
eba721139f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_bearer_skip_sigv4_chain 2026-09-02 18:36:10 -07:00
mateo-berri
7a32ef131f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4
# Conflicts:
#	litellm/router_utils/fallback_event_handlers.py
2026-09-03 01:34:31 +00:00
mateo-berri
b6c10d31e8 chore(lint): ratchet down Any and strict-typing budgets 2026-09-03 01:33:11 +00:00
tin-berri
78ff5ac9cd
feat(router): arm safeguard-refusal fallback on generic chains when no content-policy list exists (#39274) 2026-09-02 18:31:25 -07:00
mateo-berri
85961201e7 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci 2026-09-02 18:30:13 -07:00
mateo-berri
105f99cea2 fix(proxy-extras): kill the whole Prisma process group when a command times out
Every Prisma CLI call now goes through one runner that starts the command in
its own session and SIGKILLs the process group on timeout, so the Node process
and the Rust schema engine die together with the Python wrapper instead of
being reparented to pid 1, where they kept applying migrations after the proxy
had given up and held the Prisma advisory lock against every retry and every
later boot. Tests that faked subprocess.run now fake the runner, and the fake
Prisma CLI in the migration tests forks a grandchild that must not outlive a
timed-out migrate deploy.
2026-09-02 18:29:55 -07:00
devin-ai-integration[bot]
8441dd6e8c
fix(proxy): keep SpendLogs and callback session ids in sync when the request has none (#39450)
* fix(proxy): keep SpendLogs and callback session ids in sync when the request has none

Add general_settings.missing_session_id (generate | reject). In generate mode one id is
stamped into litellm_session_id, litellm_trace_id and metadata.session_id before callbacks
run, so LiteLLM_SpendLogs.session_id and the Langfuse session id match. In reject mode such
requests get a 400. Unset keeps the legacy behavior. MCP routes are not affected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy): regenerate schema.d.ts and shorten mutable-ok comment for ruff format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): mark generated session ids so affinity consumers do not pin on them

Fireworks x-session-affinity, the router session_affinity pre-call check and the
complexity router session pin all read metadata.session_id as a caller-chosen
stable key. A missing_session_id: generate id is fresh per request, so it now
carries metadata.litellm_session_id_generated and those consumers skip it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 18:28:06 -07:00
mateo-berri
f29fdee625 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4
# Conflicts:
#	basedpyright-code-budget.json
#	litellm/llms/azure_ai/vector_stores/transformation.py
#	litellm/llms/milvus/vector_stores/transformation.py
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-09-03 01:27:33 +00:00
devin-ai-integration[bot]
92edcb90db
fix: keep litellm importable on Python 3.10 and guard 3.11-only typing imports in CI (#39448)
* ci: guard against Python 3.10-incompatible typing imports

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: address Python 3.10 typing guard review

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): honor version-guard direction and scan litellm-proxy-extras in py310 typing check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 18:27:19 -07:00
Yuneng Jiang
63482cfdbd
chore(deps): lower the pymongo floor for the mongodb extra to 4.9
4.17 was picked on the belief that dnspython only became a core pymongo
dependency there, which is wrong: pymongo has declared dnspython>=1.16.0,<3.0.0
as a core requirement since well before that, so mongodb+srv:// URIs resolve at
4.9 too. The real floor is 4.9, the release AsyncMongoClient landed in, and 4.8
has no AsyncMongoClient at all.

Verified against live Atlas on 4.9: sync and async search, list_search_indexes,
same top hit and score as 4.17. Resolution is unchanged, pymongo 4.17.0 either
way, so this only widens what an existing environment is allowed to bring.
2026-09-02 18:25:51 -07:00