The local rate-limit fallback path introduced in #40596 restores the request from a snapshot taken before the pre-call pass, so the guardrails resolved for the requested model were dropped when another deployment was selected, and the raw request-body disable_fallbacks field gated the fallback before the key-level disable_fallbacks override had been applied
Carry the requested model's merged guardrail list onto every fallback pass and read the effective disable_fallbacks value after the pre-call pass
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Auth runs the tag budget check before pre_call_hook, so a tag that a custom guardrail adds is attributed spend but never budget checked. After the pre-call hook, budget check only the newly added tags with the same exemptions auth applied (budget-free routes, zero-cost models), keep the pre-guardrail tag baseline across fallback retries, and surface an over-budget tag as the same budget_exceeded 429 auth returns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Key metadata disable_fallbacks only lands on data during add_key_level_controls,
so the local rate-limit fallback retry now rechecks it post pre-call. Also use a
real UserAPIKeyAuth in the skip pre-call test since the path reads router_settings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The fallback retry in _pre_call_with_fallbacks re-entered common_processing_pre_call_logic with data already enriched by the first pass, so add_litellm_data_to_request deep-copied a metadata dict holding the live OTel span and the request failed with a 500 (cannot pickle '_thread.RLock') instead of the intended 429 or fallback. Capture the configured fallbacks and a snapshot of the client request before the first pass, look up the fallback chain by the normalized model group after the limiter raises, and run each fallback attempt on a fresh copy of that snapshot. Replaces the mock-heavy tests with a rig that runs the real v3 limiter and a live OTel span through the proxy_logging_obj seam
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Key and team router_settings.model_group_alias aliases were resolved only after the key/team model allowlist checks ran, so a key allowed the alias target was denied when it requested the alias. Resolve the alias during auth and rewrite the request body to the target before the allowlist checks. The alias the client sent is kept in the request scope so the response model still echoes it.
Resolves LIT-3054
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
RouterRateLimitError now carries the model group's deployment ids so it
can tell when every deployment is cooled down, and exposes that as
type=all_deployments_in_cooldown with an explicit message. A partial
cooldown keeps type=rate_limit_error. Either way the proxy no longer
reports type=internal_server_error next to code 429
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Keeps the base's rule that a non-admin id lookup matching no spend-log row answers 403, so the detail route never consults cold storage without an owner row
A team-scoped auto-router is stored under an internal
model_name_{team_id}_{uuid} with the caller-facing name in
model_info.team_public_model_name, and the four pre-routing strategy
registries key on that internal name. A team key asks for the public name,
so the strategy lookup missed, the team early-resolve exit handed back the
marker deployment itself, and every call 400'd with "Unmapped LLM provider".
The strategy lookup now resolves the requested name through the same
team-first, then global, then admin-across-teams deployment resolution the
deployment path uses, and looks the registries up under the model_name of
whatever that resolves to. Both exits of _common_checks_available_deployment
drop strategy markers through one helper, so a marker-only resolution is
rejected as uncallable on every path. The request team id has one reader.
Resolves LIT-7363
Claude-Session: https://claude.ai/code/session_01NU97S7d2FUDDvTk59k53Wp
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
The two test methods, the policy_engine fixture, and the two inner
stubs in TestBackgroundResponseRetrievalGovernance now carry full
parameter and return annotations, closing the Greptile thread that
94f9230d13 left open.
A POST /v1/responses with background: true returns a queued response, so the
post_call pipelines attached at submit time had nothing to inspect. They now
defer on queued and in_progress responses and run on GET /v1/responses/{id}
instead: the retrieval resolves the response id back to its deployment,
re-attaches the policies that governed the original model, and reports them
in the x-litellm-applied-* headers of the retrieval response.
* fix(bedrock): keep x-amzn-RequestId on chat error responses
Bedrock chat error paths built BedrockError from only a status code and a
message, so the provider response headers were gone before exception mapping
ran and the proxy had nothing to forward. AWS support needs x-amzn-RequestId
to investigate a server-side error.
- converse and invoke chat handlers pass the real headers and response when
they turn an httpx.HTTPStatusError into a BedrockError, and read the body
through error_response_text so a streamed body nobody read does not throw
- every bedrock chat get_error_class honors the headers it is already handed:
invoke, moonshot, bedrock-hosted openai, agentcore and the invoke agent
- BedrockError carries those headers into the response it synthesizes when a
caller has headers but no response, skipping values httpx cannot carry
- the bedrock 500 mapping forwards the provider response like its 4xx and 503
siblings instead of fabricating a blank one
The proxy now returns llm_provider-x-amzn-requestid on Bedrock chat errors.
* fix(bedrock): keep request-id on text-classified errors
The context-window and image branches of _map_bedrock_exception built their
litellm exception without the provider response, so a Bedrock 400 classified
by its body text lost x-amzn-RequestId while the sibling branches kept it.
Also narrows the new BedrockError types and trims its docstrings.
* chore(bedrock): drop the docstrings on the new error helpers
* fix(bedrock): keep request-id on every error path that has one
The ticket's root cause is that every BedrockError raise site under
litellm/llms/bedrock/ was built from status and message alone. The first
commits covered the chat and invoke handlers; this covers the rest.
Embeddings, rerank, image generation, image edit, count tokens, search and
the transformation layers now hand on the provider response or its headers,
and both bedrock_mantle configs return a BedrockError instead of the
OpenAI error that drops them.
Two blockers surfaced while verifying the streaming path. The trailing
`except Exception` in make_call and make_sync_call swallowed the BedrockError
raised a few lines above, relabelling a provider status as a 500, and the
non-200 branch read an unread streamed body, which throws.
The raise sites left alone have no provider response to carry: timeouts,
credential and config errors, and mid-stream event frames.
* fix(bedrock): forward provider headers from the count tokens route
The count tokens route converts BedrockError into an HTTPException, and dropped
the headers the handler had just kept, so that route still lost the request id.
get_response_headers now takes a Mapping so an httpx.Headers can be handed to it
without a copy.
* fix(bedrock): classify every bedrock surface through BedrockError
Eleven bedrock configs still inherited a provider-agnostic get_error_class
that builds a blank response, so the request id was gone before the proxy
read it. Claude platform, bedrock anthropic-messages, both image edit
configs, passthrough, realtime, vector stores and agentcore search now
return BedrockError, and a parametrized audit drives all 36 configs.
* fix(proxy): keep provider headers on the httpx status error branch
_handle_llm_api_exception forwards safe_headers on every branch except the
httpx.HTTPStatusError one, which the bedrock passthrough route reaches, so
the request id was dropped before the client saw the response.
* fix(bedrock): keep the request id on the timeout mappings
Timeout takes no response argument, so the three bedrock timeout branches
dropped the provider headers even when the upstream answered 408 or 504
with an x-amzn-RequestId. They now ride on the exception, already
llm_provider-prefixed, which is the form the proxy emits.
* fix(bedrock): keep the provider response on mapped timeouts
The previous round attached llm_provider-prefixed headers directly to the
Timeout. That shadowed the raw upstream headers for _get_response_headers,
so router cooldown and fallback cooldown stopped honouring retry-after on
bedrock 408/504 replies.
Give Timeout an optional response instead, the way every other mapped
bedrock exception already carries one. Retry logic reads the raw
retry-after off the response, and the proxy prefixes those headers on the
way out, so clients still see llm_provider-x-amzn-requestid.
* chore(bedrock): drop the explanatory comment on Timeout.response
* fix(router): keep provider response headers on streaming chat completions
The Router re-wraps a deployment's CustomStreamWrapper in FallbackStreamWrapper
(and its sync twin) so a mid-stream failure can fail over. Neither wrapper
forwarded `_response_headers`, so every streaming chat completion handed the
proxy's callbacks and its response-header builder a wrapper with no provider
headers, and a successful mid-stream fallback still published the failed
deployment's identity, `x-request-id` and rate limit counters.
Forward `_response_headers` into both wrappers, repoint the wrapper at the
deployment that served the stream once a fallback takes over, and rebuild the
proxy's response headers from that deployment while `create_response` still has
the first chunk buffered.
* fix(router): follow a nested fallback to the deployment that served the stream
A fallback the router picks is itself a fallback-aware wrapper, and it only
repoints at its own fallback once it yields, so reading its hidden params at
selection time named a deployment that produced no output. Re-read them when
the first fallback item arrives, which is still before the proxy commits
response headers.
Also addresses review feedback: the streaming header builder reads self.data
instead of taking a coarse request_data parameter, and the new test recorder
local is Final.
* test(router): cover the fallback header adoption helper directly
The router_code_coverage gate wants every router.py function named in a
router test, and this also pins the weak-reference behavior: a wrapper
collected mid-stream must not break the generator still draining it.
* refactor(proxy): take a read-only mapping for the model-id lookup
_get_model_id_from_response only reads its request payload, so a Mapping
says what it needs and the two metadata hops are narrowed instead of
assumed to be dicts.
* test: drop mutable recorder locals and routine comments from the new tests
An AsyncMock await_count and an asyncio.Event say the same thing as a
list and a dict that the test mutates.
* chore(router): justify the two rebinds in the fallback loops
Both are the one-shot re-read that follows a nested fallback, so they get
the repo's rebind-ok note like the rest of the file.
test_router.py is not ruff-formatted on staging and CI's format check only scopes
litellm/*.py, so running ruff format over the whole file rewrote ~900 lines of
unrelated code. That reflow split long single-line patch() calls into multi-line
form, which the test-quality gate counts individually, pushing TQ008 four over its
ceiling. The file is back to staging's formatting with only the compression test
class added.
test_common_request_processing.py armed a model-side guardrail name with no such
guardrail registered, which stopped working once both hops began requiring the name
to resolve to an active compression guardrail.
An auto router marker deployment can now set auto_router_routing_compression
and auto_router_model_compression in its litellm_params, naming the
compression guardrail each hop should use (or "none" for no compression on
that hop). Neither key set means the request's own compression guardrails
keep applying to both hops unchanged.
Backend: Router.async_pre_routing_hook resolves the marker's policy and
compresses a copy of the messages for the routing decision only when the
policy differs from what the model call already got; when both hops share
the same compression, it reuses what the ordinary pre-call guardrail
pipeline already produced instead of compressing twice. The proxy layer
suppresses every other compression guardrail once a policy is engaged and
arms the model-side guardrail even when it is not default_on.
UI: the auto router's Detailed Configuration gains an Advanced: Compression
section with a routing-decision selector and a same/different toggle for
the model call, matching the same/different address pattern.
A blocked guardrail (and any other HTTP error the proxy converts) came back
with "type": "None" and "param": "None", because the converters passed the
string "None" as the getattr default instead of None. OpenAI types error.type
as a required string and error.param as nullable, so type now falls back to
the type its status code stands for and param serializes as JSON null.
Covers the non-streaming body, the SSE error frame, the client-disconnect
frame, and the unclassified-exception path, so every unified LLM endpoint and
the anthropic endpoints return the same shape.
* fix(proxy): stop leaking internal exception details to clients
Public error responses could disclose internal details in two places.
A proxy-layer exception with no recognized provider status code (a bug
in a custom callback, a hook, or litellm's own code) forwarded its raw
str() text verbatim on a 5xx, including any embedded credential,
filesystem path, or internal hostname, or a full stack trace; the
same client-facing message now runs through a redaction layer built
on top of the credential redaction that already runs on log output,
so it also drops an embedded traceback and scrubs path-shaped and
hostname-shaped substrings. It intentionally never runs on server-side
logs, which must keep full detail for debugging.
exception_type(), litellm's core exception mapper, is shared by direct
SDK callers (litellm.completion()) and the proxy, and it deliberately
embeds a traceback into an unmapped exception's message as a debugging
aid for library users; a first pass at this fix stripped that
traceback inside exception_type() itself and broke that convention
(caught by tests asserting on the traceback frame). The traceback stays
in exception_type()'s own output; only the proxy's client-facing
response boundary (and the streaming response generator, which never
needs to embed one at all) strips it.
Full generic-message replacement for the unclassified-exception case
was tried first and reverted too: several routes deliberately raise a
bare exception as an informative, secret-free validation message (e.g.
the OCR endpoint's rejection of provider-native file IDs), and
replacing those wholesale broke that convention; targeted redaction
leaves them untouched.
Also stops the default uvicorn-based proxy from sending a Server
response header.
Resolves LIT-6747
* refactor(proxy): drop the unrelated error-message constant and trim redaction comments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A guardrail block or failed scan that fires after SSE chunks have been
flushed can no longer set an HTTP status, so raising HTTPException there
silently truncated the stream. _emit_streaming_http_error now routes
post-flush failures through the endpoint translation's
build_stream_error_items, emitting the surface-correct error frame on
chat completions (data: {error}), /v1/messages (event: error), and
/v1/responses (ErrorEvent with the next sequence number). Pre-flush
blocks still raise with a real HTTP status.
Successful flags-on scans also logged metadata.guardrail_information as
null: the chat handler planted litellm_metadata on a route whose bucket
is metadata, flipping the bucket for every later write, and responses
streams fired their spend log before the eos scan ran. The chat handler
now merges user_api_key metadata through get_or_create_metadata_bucket,
and deferred stream-complete logging is armed for aresponses like it
already was for anthropic_messages.
The x-litellm-response-cost header on non-streaming /v1/messages responses is
recomputed from the response body because the Anthropic TypedDict cannot carry
hidden params. That recompute ran after the body's model field had already been
restamped to the client-facing alias, so the cost calculator priced the alias
(for example together_ai/muse-glimmer-30b) instead of the deployment model that
spend logging uses. On Together AI that alias is unregistered and falls into the
parameter-size bucket, so the header overbilled cold requests by about 2.3x and
priced cache reads at zero on warm ones while recorded spend stayed correct.
Move the restamp after every cost read of the response so the header and the
spend logs price the same model, and add a regression test that pins the header
to the provider-reported model while the body still returns the alias.