Review turned up two real problems in the TTS path.
Router.aspeech forwarded voice=None whenever the caller omitted it, which overwrote a
voice set in the deployment's litellm_params, so a configured fallback voice was
ignored on voice-less requests. It now leaves the key alone when no voice is passed.
get_complete_url also fell back to MISTRAL_API_BASE, but speech() always receives a
non-null api_base from get_llm_provider, whose mistral branch only reads
MISTRAL_AZURE_API_BASE and otherwise hardcodes the public host. That branch could
never run, and its unit test asserted a behavior the real path does not have. The
working override is api_base on the deployment, now pinned by an end-to-end test
POST /v1/messages returns an Anthropic-shaped body whose `id` is the only
request id the caller ever sees, but the spend row was written with a
`chatcmpl-<uuid>` (non-streaming) or the bare `litellm_call_id` (streaming and
the /anthropic/v1/messages passthrough), so
GET /spend/logs?request_id=msg_... returned [].
The logging conversion now carries the provider's response id through:
_handle_anthropic_messages_response_logging seeds the ModelResponse it builds
with the Anthropic id, and the passthrough logging handler prefers the id it
read off the response body or the message_start chunk over litellm_call_id.
get_spend_logs_id already prefers response_obj["id"], so the spend row and
standard_logging_object["id"] now both carry the id the client holds.
Every field of a multipart form arrives as a string, so `n` reached the
provider as "2" and Bedrock Nova Canvas rejected the request with
"expected type: Number, found: String". Restore the type the request
schema declares at the boundary where the form is parsed, driven by the
schema's own type hints so the helper covers any int- or float-typed
field on any multipart endpoint.
Resolve type-discipline-budget.json by taking the lower limit per rule so no
ceiling ratchets back up.
Staging's 66a3d24b3f left a duplicate embedding_executor parameter in the
Bedrock KB fake handler, which makes ruff fail on the whole tests tree. Drop
the duplicate here so this branch compiles; #39502 makes the same change on
staging.
Resolves two conflicts:
- tests/test_litellm/vector_stores/test_main.py: staging moved search() to a
RouterVectorStoreEmbeddingExecutor while this branch parametrized the same
test over query; keep both the executor assertions and the parametrize.
- tests/logging_callback_tests/test_bedrock_knowledgebase_hook.py: staging
carries a duplicate embedding_executor kwarg that makes the file a
SyntaxError; drop the trailing duplicate.
Resolves the tests/test_litellm/test_main.py collision, where both sides appended a
new test at the end of the file, by keeping both.
Also carries the one-line fix from #39502: staging arrived with a duplicate
embedding_executor kwarg in the Bedrock KB fake handler, which ruff rejects as a
syntax error, so every commit here would otherwise fail lint. The change is byte
identical to #39502, so that PR merges cleanly once it lands.
* Fix hide-secrets guardrail: playground redaction, UI dropdown entry, spend-log telemetry
The hide-secrets guardrail never implemented apply_guardrail, so the UI test
playground echoed secrets verbatim; it was missing from the Add Guardrail
dropdown; and it recorded no guardrail_information, so Spend Logs could not
distinguish a redacted request from a clean one.
- implement apply_guardrail (unified interface) with use_native_lifecycle_hooks
so proxied traffic stays on async_pre_call_hook (per-key opt-out and
data["prompt"] handling live only there)
- record standard_logging_guardrail_information (allow/mask + masked_entity_count)
via _process_response/_process_error; opted-out keys and legacy nameless
callback instances record nothing
- advertise hide-secrets in /guardrails/ui/add_guardrail_settings (pre_call only)
and /guardrails/ui/provider_specific_params with a config model
Resolves LIT-3548
* Fix hide-secrets passthrough telemetry and JSON config input
* fix(guardrails): validate hide-secrets object config before submit
- apply_guardrail treats empty-string-only texts as no input, so no
false allow is recorded
- the UI object field keeps raw text while editing and blocks submission
until it parses to a JSON object, instead of posting a string to an
object-only API
- supported_modes_by_provider keeps its dict[str, list[str]] value type
* fix(guardrails): record no hide-secrets telemetry when nothing was inspected
walk_user_text and the prompt redaction now report how many non-empty
strings they visited; when neither inspected anything (image-only
content, empty strings), the run records no guardrail entry instead of
an 'allow' row that counts a check which never saw any text.
The v2 migration resolver gave `prisma migrate deploy` four attempts, and
every recovery path ended in a bare `continue`, so each one burned an attempt.
A database first brought up with `--use_prisma_db_push` has a full schema and
no migrations ledger, so the baseline spent attempt one and the first three
migrations whose objects already existed spent the rest. The proxy then exited
before binding its port, and that database could never be moved onto the
resolver.
The retry budget now counts only attempts that got nowhere. Creating the
baseline, and each migration newly marked applied, leaves the budget alone, so
a push-created database works through its pre-existing objects one pass at a
time. Timeouts, deadlock rollbacks, advisory-lock waits, and a repeat of a
recovery that already ran still spend an attempt, so a run that stops making
progress gives up exactly as before.
Two branches independently added embedding_executor to the same fake
search handler in this file, #39472 in the middle of the signature and
#39474 at the end. Neither conflicted with the other, so both edits
merged and the function ended up declaring the parameter twice.
Python rejects that at compile time, so the whole module fails to
import and every test in the file is uncollectable, taking the
logging_testing job down on staging.
Keep the earlier of the two, which sits where the real handler declares
the parameter.
* fix(sso): resolve multi-valued role claims to the highest privilege role
A role claim carrying several roles used to resolve to whichever one the IdP
listed first, so a user holding both proxy_admin_viewer and internal_user lost
org-level spend visibility depending on claim ordering alone.
get_litellm_user_role now picks the highest privilege role out of a list-valued
claim, and the Entra app_roles path shares that same resolution instead of
keeping its own copy of the hierarchy. SAML assertions carrying several role
values go through the same path rather than taking the first value.
* test(sso): lock ranked-over-unranked role resolution for mixed claims
org_admin, team and customer sit outside the privilege ladder. Pin the
resolution for a claim that mixes one of them with a ranked role so the
asymmetry is covered rather than implicit.
* fix(sso): label the claim-sequence cast for the type-discipline gate
* fix(sso): resolve claim entries without recursing
The repo's recursive-function gate rejects self-recursion here, and a role
claim is flat anyway. Pull the single-value lookup into its own helper so the
list branch maps over it instead of calling back into itself.
One unreachable vector store used to wipe out every store's context on a
chat completion carrying vector_store_ids: the search raised, the blanket
handler returned the original messages, and the request answered with no
retrieved context at all. Each store's search now has its own handler that
warns with the vector store id and moves on to the next store.
The same loop appended every store's results to the original messages
instead of the running copy, so with two healthy stores only the last one
reached the model. It now chains through modified_messages.
The Router is injected through a ProxyRuntime protocol instead of an
in-function litellm.proxy.proxy_server import, so the hook's routing can
be driven in tests without touching proxy globals.
The Router executor only routed a query embedding when the vector store
carried extra embedding configuration, so a store registered with no
embedding model at all always went to the Router and 500'd on the
s3_vectors default text-embedding-3-small when no deployment served it.
Route on whether the Router serves the model, which is the rule the
executor had before, and keep the request metadata on the SDK fallback so
the embedding stays attributed either way.
Forwarding limit to OpenAI made the ownership filter cut the page down after
the fact, so a key that owned an older container got an empty first page and
its cursor never moved. Non-admin lists now walk upstream pages of 100 until
they have enough owned containers (or five pages), trim to the requested
limit, and report first_id, last_id and has_more off what the caller keeps.
Also assigns tests/test_litellm/proxy/container_endpoints to a CI shard.
Retry breadcrumbs were appended to one list owned by the Router and shared by
every request, and each breadcrumb copied the whole kwargs including the proxy's
snapshot of the inbound request. That snapshot's body aliases the live request
metadata, breadcrumbs included, so every new breadcrumb nested all the earlier
ones inside itself. Memory stayed small because these are shared references, but
under --detailed_debug the repr of that structure expands, so one debug line grew
from 10k to 219M characters over 14 failing requests and the proxy stopped
answering.
Breadcrumbs now accumulate in the metadata of the request that produced them, the
request snapshot is excluded from a breadcrumb, and the cap of the last 4 failed
attempts applies per request.
Adds chat_completions and transcription as SDK function columns, backed
by the existing rust_bridge test files. Adds a fourth strategy folder,
existing_e2e_test_sdk, that points at already-existing live-API SDK
tests (tests/ocr_tests/ as a whole folder, plus chat completion and
Whisper transcription tests) instead of writing new parity tests.
Extends selector_matches_node with trailing-slash folder selectors so
a whole test folder can back one matrix cell.
* feat(datadog_llm_obs): cost tag dimensions, router decision fields, reasoning token metric, redaction gating
* test(datadog_llm_obs): satisfy test quality gate
* fix: forward integer parent_id as its string form
* fix(datadog): sanitize redacted message roles
* fix(datadog): keep the A2A agent role on redacted spans
* fix(datadog): merge current staging budget
* style(datadog): format redaction tests
* fix(datadog): handle malformed redacted roles
* test(datadog): put the test quality suppression on the reported line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A session pinned to a text-only model failed every image turn with a provider 400, because the modality gate exempts a kept pin by cause. Add modality_pin_override so that exemption is conditional: the image turn is re-placed on a capable model for that request only, reported as cause modality_pin_override, and the stored pin is left untouched so the next text turn replays it.
The pin write on the replay path already happens upstream of the gate and stores the session's own model, so pin survival is structural rather than bookkeeping. The new cause joins the non-pinnable set. Default off at every layer.
S3 Vectors now subclasses BaseQueryEmbeddingVectorStoreConfig, so its query
embedding runs through the Router executor with the request metadata instead
of a private router lookup. embedding_model stays accepted as an alias of
litellm_embedding_model. The router kwarg is gone from the search handler and
every provider transform now that nothing but the executor fallback read it.
* test: add OCR python-to-rust test parity ledger
* feat: add ledger loader for OCR test parity data
* feat: add drift audit for OCR test parity ledger (WIP, untested)
* fix: correct drift in OCR test-parity ledger
Two entries referenced a typo'd Python test name, four duplicated entries already tracked under TestProxySecurityGuard, five real Python tests in test_rust_bridge.py were untracked, and three real Rust custom_logger tests were missing from rust_only_tests. Found by running validate_ledger.py's audit against the live repo.
* test: add regression coverage for the OCR ledger and audit script
Covers schema validation, AST/regex test enumeration, drift detection on both the Python and Rust sides, LedgerDriftError content, and a live-repo clean-audit guard against future drift.
* test: simplify ledger test to one drift-guard assertion
Replace the ledger-internals unit tests with a single test that runs the real audit against the live repo and asserts every OCR test is accounted for (mapped, unmapped-with-reason, or rust_only), printing the exact diff on failure.
* chore: move OCR test-parity ledger to core/ocr
validate_sub_methods/ mixes strategy-catalog metadata with the ledger. Ledger data belongs under a per-function core/<function>/ path instead.
* fix: point LEDGER_PATH at the new core/ocr location
* fix(model_armor): handle Anthropic Messages and Responses streams in post_call
The post_call streaming hook buffered every chunk and fed it to
stream_chunk_builder, which only understands chat-completion deltas.
/v1/messages streams raw Anthropic SSE bytes and /v1/responses streams
typed Responses events, so both raised litellm.APIError and surfaced to
the client as a 500 on every streamed request.
Assemble each surface with its own reader, frame guardrail failures as
terminal items in that surface's wire format, and pass the stream
through unscanned when it cannot be assembled instead of raising.
* fix(model_armor): classify the stream surface and fail closed when it cannot be assembled
Decide the wire format explicitly instead of inferring it from a boolean pair, so an
opaque raw SSE stream (the Google :streamGenerateContent route) is never refused in
Anthropic framing, and a stream that cannot be assembled is blocked rather than
released unscanned unless fail_on_error is disabled.
Also scan Responses tool-call arguments, read the body only off a terminal Responses
event, and record the applied guardrail on the fail-closed path.
* test(model_armor): pin the error-only stream predicate against content-carrying streams
is_sse_error_stream decides whether a buffered stream is forwarded to the client
untouched, so a stream that still carries content must not qualify: the frames-only
join drops typed chunks, an empty stream is not a refusal, and a content event may
carry an empty error field.
* fix(model_armor): let a streamed de-identify match mask instead of blocking
A de-identify template reports MATCH_FOUND for every redaction it makes. The
streaming block check omitted allow_sanitization, so with mask_response_content
enabled that match read as a refusal and the client got a 400 where the
non-streaming sibling returned the redacted text. Pass the flag through, as the
non-streaming hook already does, and stamp the logged status from the same
decision so the spend row agrees with what the client received.
Also drop Any from the chat-completion assembler's parameter; stream_chunk_builder
takes a bare list, so list[object] carries the mutability requirement without
erasing the element type.
* fix(model_armor): fail closed when a streamed de-identify match cannot be applied
Allowing sanitization past the streaming block check is a promise to apply the
redaction Model Armor asked for. Two paths broke that promise and released the
buffered original instead: a match that comes back with no sanitized text, and a
surface with no assembled body to rewrite.
The outcome is now resolved once, before it is recorded, so the status stamped on
request metadata agrees with what the client receives rather than reporting the
success the block check alone would have implied.
* fix: scan the deltas when a Responses stream ends without a body
response.failed and response.incomplete are terminal events like
response.completed, but a turn that broke mid-generation reports an empty
output while the deltas ahead of it already spelled the answer out to the
client. Reading only the terminal body found nothing to scan there, and the
empty-content shortcut then forwarded every buffered delta past the guardrail.
Fall back to the text the delta events carry whenever a Responses stream
assembles to nothing.
* fix: read the Responses delta event types off the event enum
The hand-listed set left out response.mcp_call_arguments.delta, so a turn that
streamed only MCP tool arguments and then reported an empty body still took the
no-content shortcut and forwarded those chunks unscanned.
Deriving the set from ResponsesAPIStreamEvents keeps it complete as the enum
grows, and the str guard in the reader already covers any event whose delta is
not text.
* fix(model_armor): scan responses deltas alongside the terminal body
A /v1/responses stream spells out reasoning summaries and tool-call arguments in
delta events that its terminal body never repeats, so scanning the body alone
handed every summary delta to the client unscanned whenever the body carried text.
* fix(model_armor): scan responses delta fields apart from each other
A Responses turn spells out its reasoning summary, its visible answer and its tool-call
arguments in separate delta events. Joining every delta into one string let a finding form
across the boundary between two fields that each carry nothing to find, so a safe stream
could be blocked. Group the deltas by the field they belong to, join a field's own deltas
as they streamed, and keep the fields apart.
* fix(model_armor): scan each responses field once, not twice
Separating delta fields stopped the terminal body from matching the delta text, so a turn
with two visible fields sent Model Armor both copies. Only the delta fields the body does not
already carry are appended now.
---------
Co-authored-by: yassin <yassin@berri.ai>
* feat(azure): support credential chain for storage
* test(azure): clarify credential seam suppressions
* fix(azure): read chain tokens in a worker thread
The credential chain walk (IMDS probe, CLI subprocess) is blocking I/O,
so reading the provider inline in async set_valid_azure_ad_token stalls
every request on the worker's event loop
The container retrieve, list, delete, create and file routes validated the
provider's error body against the success model, so a deleted or unknown
container and a rejected API key surfaced as 500 pydantic errors instead of
the upstream 404 or 401. The handlers now raise the provider error class with
the upstream status and message before transforming the response.
GET /v1/containers dropped after, limit and order before calling the
provider, and GET /v1/containers/{id}/files dropped the same three, so
paginated list calls ignored their pagination arguments. Both routes now
forward their declared query params.