Adds run_only_on_call_types (allowlist) and skip_call_types (denylist) so an
operator can keep calls such as embeddings away from the guardrail endpoint.
A filtered call is passed through untouched on both the request and response
hooks, nothing is sent, and a not_run guardrail entry records why. The
allowlist wins when both are set. Values are CallTypes values; unknown values,
a bare string instead of a list, and call_mcp_tool (MCP tool calls reach the
guardrail without a call type) are rejected at startup. A call whose type
cannot be resolved still runs the guardrail. Both options default to off, so
behavior is unchanged unless configured
The call type comes from the authenticated user_api_key_request_route, then
the logging object. The route is stable across both hooks, while the logging
object's call type is rewritten to "responses" when a chat or messages call is
bridged to the Responses API
Batch-file scans now run each record under its own endpoint route instead of
the upload route, so a filter, and any guardrail that reads the key's route,
sees the record as the chat or embeddings call it is. A record only gets that
route when its body shape agrees with its url and the Bedrock and Vertex
record classifiers would run it as the same call type. Any other record, such
as a chat body labeled /v1/messages, gets no route and is always scanned
The batch scan also strips a litellm_logging_obj key from each record body
before the guardrails run. On main a truthy caller value there made the
generic guardrail fail, and fail_on_error: false then shipped the record
unscanned
Adds get_primary_call_type_for_route and uses it in the unified guardrail so
both resolve a route to a call type the same way. GenericGuardrailAPI also
accepts an injected async_handler, used by the tests
Pass-through guardrails lost their inbound headers once the caller's body
copies were stripped, so an operator's extra_headers allowlist forwarded
nothing to the vendor. Both the pre-call and the post-call guardrail hooks
now get the real request headers in proxy_server_request, built once the
way the chat path builds them: clean_headers drops the proxy's auth header
and a custom litellm_key_header_name, MCP upstream credential headers are
dropped, and redact_credential_headers masks cookies and other credentials
On the success path the headers are attached after the logging object is
created, so they do not land in the logged request. When a pre-call
guardrail blocks, the failure log now receives these cleaned and redacted
headers, which matches what the chat path logs. The pass-through guardrail
leaves proxy_server_request out of the text it scans, and the upstream body
is unchanged because proxy_server_request is popped with the other litellm
params before the send
Moves the imports this PR added inside functions to the top of their
modules: UserAPIKeyAuth and BaseTranslation in realtime_streaming,
BaseTranslation, StreamingScanKey and LiteLLMProxyRequestSetup in
unified_guardrail (so its annotations no longer need quotes), and the
test-only imports in the unified guardrail, guardrail endpoint, pass-through
and realtime tests. None of them creates a cycle, and import litellm still
does not load the proxy request layer or fastapi
One import stays inside its function: base_translation's
LiteLLMProxyRequestSetup. import litellm loads base_translation before
litellm.Router exists, and litellm_pre_call_utils imports Router, so hoisting
it makes import litellm fail with "cannot import name 'Router' from
'litellm'". It carries a one-line comment saying so
Guardrails read user_api_key_alias, user_api_key_team_id, the key hash and the request route from request metadata, and the generic guardrail API forwards them to the vendor. The chat path strips caller copies of those fields and writes the real ones, but three other paths did not, so a caller could claim another key, team or request route (which call-type lookups key on)
/guardrails/apply_guardrail passed the body metadata straight to the guardrail. It now drops the fields the chat path treats as untrusted, plus the bare user_api_key and the caller's headers (a slightly wider strip than the chat path, on purpose). Then it adds the authenticated identity and the proxy's real request headers
Pass-through handed the raw body to pre_call_hook. It now runs the chat path's strip on both metadata buckets first, which also stops a body from switching off global guardrails, and drops the body's headers and proxy_server_request so it cannot pick the inbound headers a vendor sees. Those keys were already popped before the upstream send, so the forwarded body does not change. The pass-through guardrail text no longer includes litellm_metadata, which carried the key's identity dump into the scanned payload
The unified guardrail only filled litellm_metadata when it was missing, so a caller-supplied bucket won. It now overwrites the identity fields of an existing bucket and drops any user_api_key_token there, since the proxy never writes one. It leaves user_api_key_auth_metadata alone because on litellm_metadata routes the proxy has already merged team metadata into it
transform_user_api_key_dict_to_metadata used to dump every UserAPIKeyAuth field, including the raw token of non-sk keys, JWT claims, team membership, proxy config and org and project metadata with callback secrets. It now returns only the chat path's identity fields plus user_api_key_key_alias for existing readers, so MCP, pass-through, output-side handlers and realtime transcript guardrails all get that same identity allowlist and see the real user_api_key_alias. The generic guardrail uses user_api_key_token only from litellm_metadata and only when user_api_key_hash is missing, so a CLI session key's raw per-login token never reaches the vendor
The strip and the identity fields live in one place in litellm_pre_call_utils, and the chat path calls them too
* refactor(types): replace Any with proven types in 8 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep email logger untyped where its alert types differ
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build(rust): package the gateway container
* ci: exempt the gateway Dockerfile from the CI coverage gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 6 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(types): keep enterprise email import inside try-except for unsafe-import check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(types): keep email_logging_instance annotation as Any pending a guarded alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): revert iterator override typing in proxy utils
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit bc9b6f8a5c)
* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard
PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vertex_ai): reproduce traced Gemma Responses stream failure
* fix(vertex_ai): wrap Gemma fake streams for Responses tracing
* test(vertex_ai): cover Gemma traced streams and usage options
* test(vertex_ai): inject gemma test deps and assert hidden usage accounting
Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.
Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.
* test(vertex_ai): drop explanatory comment from usage-option assertions
* fix(gemini): forward seed to the Gemini API instead of rejecting it
The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough
* test(gemini): assert the forwarded seed without mutating shared state
* fix(tools): salvage concatenated JSON tool call arguments
* fix(tools): harden concatenated tool-call salvage for review findings
Skip non-dict JSON during split so salvage cannot emit empty tool calls.
Collapse srvtoolu_ expansions to the first object so server results stay paired.
Allocate __concat_n ids that cannot collide with sibling tool call ids.
Propagate cache_control onto every expanded Anthropic tool_use block.
Rename the XML invoke loop variable so the key-leak gate no longer flags {args}
* test(tools): cover concat id bump and srvtoolu array keep
Only collapse srvtoolu_ when concatenated salvage expanded; a valid JSON
array argument stays one server tool input
* revert(anthropic): drop concat expansion from pass-through adapter
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* revert(tools): keep concat salvage out of request-side tool converters
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): expand strictly salvaged concatenated tool arguments in normalized tool calls
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): retain at most the salvage cap while validating concatenated arguments
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(tools): assert concat sibling ids unique after sanitization
A sibling id that only collides after colon-to-underscore sanitization must force the next concat suffix
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* refactor(tools): drop unused strict mode from split_concatenated_json_objects
Strict mode had no production caller. Rejection cases now sit on salvage, and split matches upstream main
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
---------
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
The leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. Same change as #43391 on rc/1.103.0
_is_unsignable_thinking_block() only checked block["signature"], so a
thinking block with a valid-looking signature but empty (or
whitespace-only) thinking text sailed through _drop_unsignable_thinking_blocks
and into anthropic_messages_pt(). Anthropic rejects that with:
400 messages.N.content.M.thinking: each thinking block must contain thinking
This is reachable whenever a thinking_blocks history item gets replayed
through this Anthropic-shaped request path (e.g. a non-Anthropic reasoning
turn with no summary text), the same class of bug PR #36033 fixed on the
Responses adapter's own separate code path.
Now the signature check runs first (unsigned blocks are still dropped, same
as before), then an additional check drops the block if `thinking` is
missing, not a string, or strips to empty. redacted_thinking blocks are
untouched since they don't have type == "thinking".
* fix(vertex_ai): consider tools when validating context caching min tokens
Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.
Fixes#42804
* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
* feat(otel): add SigNoz preset for OpenTelemetry v2
Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite
Absorbs the work from https://github.com/BerriAI/litellm/pull/38206
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): drop explanatory comments from the SigNoz preset
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(types): keep signoz dynamic param lines within ruff format width
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(signoz): assert the missing-endpoint boot path directly instead of in an except block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for the signoz health service
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): add the openai_like chat config foundation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): let max_completion_tokens outrank max_tokens and decline refusal responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(e2e): read management routes back from the control plane replicas
The suite's management read-backs (/key/info, /team/info and friends)
polled the same replica list as the data plane. On a componentized stack
whose LITELLM_PROXY_REPLICA_URLS names the gateway pods directly, that
list answers those routes 404, since a gateway pod trims the management
routes at startup. A new LITELLM_CONTROL_PLANE_REPLICA_URLS names the
addresses a management read-back polls instead: an exported list wins,
and when it is unset the old rule stands, the data-plane replicas while
the control plane shares the suite's base URL and the control-plane base
alone once it is split.
build_proxy_client takes the list as control_replica_urls and
read_back_everywhere picks its replicas per path, the way the rest of the
client already does.
* fix(e2e): derive the control replicas of a client built for another proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(responses): stream guardrail pre-call block as SSE with a typed output item
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): import blocked usage helper from the guardrail utils module
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop narrating docstrings and poll without rebinding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover pre-call guardrail block on /v1/responses stream and json
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for responses guardrail block contract
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): observe upstream on the recorded chat route for responses denial cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): tidy responses denial audit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for worker count to recover after SIGKILL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): require a replacement worker after SIGKILL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the blocked response test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>