Trimming the streamed buffer to the retained tail could drop a category
exception phrase that suppresses a later keyword, or the identifier word
of an unfinished sentence that a conditional category pairs with a later
block word. Refuse the cut while either would leave the buffer so the
bounded scan masks and blocks exactly like a scan of the full text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Budget is enforced once during auth, against the requested model group.
`_is_model_cost_zero` waives every budget check for a zero-cost group, and the
router then picks a fallback target afterwards, inside `run_async_fallback`,
where nothing re-checks budget. A free model with a paid fallback therefore
bills with no budget gate at all.
Add `fallback_budget_check`, the budget sibling of the existing
`fallback_access_check`: a predicate awaited per fallback target that skips
targets the caller cannot pay for. The primary attempt is untouched, so a
zero-cost model is never blocked by budget and only the paid fallback is
refused.
Counter reads pass `max_budget` so `get_current_spend` verifies against
authoritative recorded spend, matching the auth-time key and user checks; a
counter restored from an older snapshot reads as a hit rather than a clean
miss, so without it a stale-low value would keep admitting paid fallbacks.
A zero-cost fallback target is always allowed, and a team key does not inherit
the key owner's personal budget unless `apply_user_budget_to_team_keys` is set,
matching `_PROXY_MaxBudgetLimiter`.
Scope is key and user budgets. Team, team-member, end-user, org, global and
per-model budgets are not covered yet: those auth-path functions enforce rather
than report, so reusing them would fire threshold alerts and take spend
reservations for a target that is then skipped. Two limitations of that scope
are documented in the module docstring: the check reads the spend counter
rather than reserving against it, so concurrent fallbacks can cross a cap
together; and a request reaching the router without
`metadata["user_api_key_auth"]` is not restricted. Both are shared with
`fallback_model_access.py`.
Opt-in via `general_settings.enforce_fallback_budget`.
Relates to #41344
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The streaming post-call hook rescanned the whole accumulated choice buffer on every chunk, so scan cost grew quadratically with output length. Keep a bounded per-choice buffer instead: once it exceeds twice the scan context, drop the head when masking the head and tail separately yields the same output as masking the whole buffer, so no pattern, phrase or exception straddles the cut. Detections from the dropped head are kept and merged, deduplicated, into the final log row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds a persistent total_spend column to LiteLLM_VerificationToken and LiteLLM_DeletedVerificationToken, incremented in the same write as spend and left alone by budget resets. Surfaces it on /key/info, /key/list and the Admin UI Virtual Keys table and key detail view
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Key metadata disable_fallbacks only lands on data during add_key_level_controls,
so the local rate-limit fallback retry now rechecks it post pre-call. Also use a
real UserAPIKeyAuth in the skip pre-call test since the path reads router_settings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A body that serializes litellm_trace_id as null or an empty string carries no identity, so it must not
block the server span fallback. Also mark the nested metadata write as an out-param store
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
UserAPIKeyAuth.parent_otel_span is Any at runtime (opentelemetry is an optional extra), so the OTel
trace-id fallback must only format an int trace id, otherwise an object that merely quacks like a span
turns the whole request into a 500
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Pass user_api_key_dict.parent_otel_span into the outgoing W3C injection so the
legacy otel callback propagates its litellm_request span, falling back to the
otel_v2 request root span and then the ambient span. Extend the mapped unit
tests to assert the propagated trace and span ids over real captured headers
for HTTP and WebSocket passthrough with forwarding on and off.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
When the otel callback is enabled and the client sends no trace or session identity, the request now inherits the W3C trace id of the proxy's server span as litellm_trace_id and metadata.trace_id. The missing_session_id policy and SpendLogs then persist that value as session_id, so a trace in the OTel backend and its row in the Logs UI carry the same id. Explicit x-litellm-trace-id, traceparent, body metadata.trace_id and litellm_trace_id keep priority.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The fallback retry in _pre_call_with_fallbacks re-entered common_processing_pre_call_logic with data already enriched by the first pass, so add_litellm_data_to_request deep-copied a metadata dict holding the live OTel span and the request failed with a 500 (cannot pickle '_thread.RLock') instead of the intended 429 or fallback. Capture the configured fallbacks and a snapshot of the client request before the first pass, look up the fallback chain by the normalized model group after the limiter raises, and run each fallback attempt on a fresh copy of that snapshot. Replaces the mock-heavy tests with a rig that runs the real v3 limiter and a live OTel span through the proxy_logging_obj seam
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Already shaped ProxyException and HTTPException errors passing through the moderations, audio speech, Anthropic Messages, and handle_exception_on_proxy paths now answer with the x-litellm-call-id header the route logged under, without overwriting a header the exception was raised with. The GET /v1/batches failure hook receives the resolved request data so the spend log request_id matches the response header and the error log
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Removes the Singulr async_logging_hook and logging_hook overrides so logging_only runs through CustomGuardrail.async_logging_hook: the response scope reaches Singulr as an assistant message instead of a raw ModelResponse dump, a vendor timeout is recorded as guardrail_failed_to_respond, a request-scope block ends the scan, and the sync success callback thread makes no Singulr call.
Decides MCP versus LLM by the proxy logging object's call_type (then the call_type or server-only markers in request_data), never by name, arguments or mcp_tool_name keys a client can put in a chat body. REST /mcp-rest/tools/call pre-scans reach Singulr as mcp_request and a non-mapping arguments value is forwarded as tool_arguments instead of raising.
should_block is a strict bool defaulting to false so a null verdict is an invalid response that block_on_error decides; payload fields drop Any for Sequence, Mapping and AssistantMessage types; metadata carries only the keys present; docstrings and section comments removed per the repo comment policy.
Squash of BerriAI/litellm#37464 (head da298ca7) by @aniket-kardile, adopted onto main: v2 gateway payload contract with request, response, mcp_request and mcp_response scopes, typed payload models, proxy user, org and team metadata forwarded to Singulr, and the logging_only, pre_mcp_call and post_mcp_call modes.