Closes LIT-4358.
DataDogLLMObsLogger.async_send_batch posted the entire log_queue to
/api/intake/llm-obs/v1/trace/spans as one payload with no byte or
event-count cap and no failure handling. Worse, the base
CustomBatchLogger.flush_queue cleared the queue even when the send
failed, so any failed flush window silently dropped every span.
Port the _send_with_413_split strategy from the log-intake logger:
- Proactively split batches exceeding DD_MAX_BATCH_SIZE (1000) events
or DD_MAX_PAYLOAD_SIZE_BYTES (4 MB serialised) before POSTing
- On a 413 response, halve and retry; drop single un-splittable spans
- On transient errors, re-queue undelivered spans for the next flush
- Override flush_queue so the base class cannot wipe re-queued spans
- Run the proactive size check inside the send loop's try so a
serialization failure re-queues only undelivered chunks
Adds pytest cases covering proactive count and byte splits, 413
halving, single-span drop, transient and HTTP 500 re-queue,
flush_queue preservation, and the no-duplicate guarantee on size
check failure.
Streaming requests return from common_request_processing before
async_post_call_success_hook runs, so response._hidden_params.additional_headers
never gets the v3 x-ratelimit-{descriptor_key}-{remaining|limit}-{rate_limit_type}
entries. Prometheus / logging callbacks that read those values from
standard_logging_object.hidden_params.additional_headers then see nothing;
combined with the pre-existing gap that Prometheus reads from that same slot
(LIT-2577 / PR #28816), per-key remaining RPM/TPM cannot be monitored for
streaming traffic at all.
Fix in three parts:
- Stash the pre-call RateLimitResponse in the metadata channels the async
success-logging callback inherits, alongside the existing top-level entry
the non-streaming path reads.
- Add async_logging_hook to the v3 handler. It fires in a distinct earlier
loop inside async_success_handler (all callbacks' async_logging_hook
complete before any async_log_success_event starts), so mirroring the
pre-call snapshot into standard_logging_object.hidden_params.additional_headers
and response._hidden_params.additional_headers here guarantees every
downstream success callback sees the values regardless of registration
order. Non-streaming keeps the existing async_post_call_success_hook write
and this hook re-populates the same values idempotently.
- Extract the shared `_merge_ratelimit_statuses_into_additional_headers`
helper the non-streaming path already had inlined so both callsites emit
the identical key shape.
Exclude vertex_ai from pipecat tool smoke; raw-ws tool_call_round_trip
remains the Vertex source of truth. Also remove the Playwright key models
dropdown suite so stage is not blocked by that UI harness
* test(e2e): cover Langfuse logging.yaml P0 logs_spend cells
Team, user/key, and org-scoped dynamic Langfuse callbacks drive real chat
traffic and assert calculatedTotalCost matches StandardLogging response_cost
and proxy spend. Also assert tool calls and applied guardrails land on the
trace. Missing env or proxy is a hard failure, never a skip
* test(e2e): use langfuse_otel callback for Langfuse spend coverage
Team and key dynamic logging attach callback_name=langfuse_otel (OTLP to
Langfuse) instead of the classic langfuse SDK. Match generations named
litellm_request by prompt marker and user_api_key_alias
* test(e2e): require Langfuse spend assert; drop AGENTS.md
Guardrail path no longer soft-gates logs_spend. Non-stream responses must
return positive x-litellm-response-cost; remove tests/e2e/AGENTS.md
* test(e2e): fail when Langfuse spend is missing on guardrail path
Always run logs_spend assertions for tool_permission; require positive
x-litellm-response-cost on non-stream and positive /spend/logs spend
* test(e2e): do not fall back to unmatched spend log rows
poll_proxy_spend_for_key returns None when response_id or positive-spend
filters match nothing, instead of silently using rows[0]
A dynamically registered (RFC 7591) OAuth client persisted onto the MCP server row is bound to the redirect_uri it was first registered with, but that binding was never recorded. After the proxy's public origin changed, every authorize paired the reused client with the new callback and the IdP rejected it permanently.
The DCR persist now records redirect_uris alongside the client identity. The admin register path treats a positive mismatch between the recording and the current callback as stale and re-registers a replacement client; rows without a recording (pre-existing installs and admin-configured clients) are grandfathered so upgrades never re-mint client_ids or orphan refresh tokens. The persist also writes client_secret and token_endpoint_auth_method explicitly as None when absent so the credential blob merge cannot pair a re-registered public client with the previous client's secret. Public register routes and non-admin callers keep existing behavior.
Closes#32473
Exact cost-map hits resolve before fallback-generalization rules, so the
mapped Sonnet 5, Fable 5 and jp Opus 4.8 Bedrock entries bypassed the
bedrock-anthropic-claude-mid-conversation-system rule and hoisted
mid-conversation system messages, invalidating the prompt cache.
* fix(anthropic): translate adaptive thinking/effort to pre-4.6 model support
AnthropicMessagesConfig now reshapes the 4.6+ adaptive-thinking interface
(thinking:{type:adaptive} + output_config:{effort:...}) to whatever the routed
model supports. Thinking-capable non-adaptive models (e.g. Haiku 4.5, Sonnet 4.5)
get the effort translated to a legacy thinking budget_tokens. Models with no
reasoning support have thinking/effort dropped under drop_params. And because
adaptive thinking carries no budget while the legacy form must satisfy Anthropic's
max_tokens > budget_tokens rule, the translated budget is capped below max_tokens,
dropping thinking when max_tokens can't fit the minimum budget. 4.6+ models pass
through untouched.
This matters because clients like Claude Code speak native Anthropic /v1/messages
and send the adaptive interface unconditionally, regardless of the routed model.
The native passthrough previously only capability-gated the OpenAI-style
reasoning_effort alias and forwarded native output_config/adaptive thinking raw, so
a pre-4.6 model rejected it with "This model does not support the effort parameter"
and the request failed. Claude Code already gets drop_params auto-set, so its
requests now succeed.
* test(anthropic): gate undersized-max_tokens thinking drop on drop_params; add edge tests
Addresses review feedback on the max_tokens-too-small branch. Previously a
thinking-capable model whose max_tokens could not fit the minimum thinking budget
had thinking silently dropped regardless of drop_params, while a residual
output_config field in the same call still raised when drop_params was off. Gate
both consistently on drop_params: raise a clear error (naming max_tokens for the
undersized case) when drop_params is off, drop otherwise. Claude Code gets
drop_params auto-set, so it still succeeds.
Adds tests for the undersized-max_tokens raise, the residual output_config raise,
and the no-adaptive-interface passthrough on a non-adaptive model.
* fix(anthropic): make adaptive-effort translation silent to avoid breaking provider strip contracts
The previous raise-when-not-drop_params behavior broke existing bedrock and vertex
messages tests: those providers already silently strip unsupported output_config
for pre-4.6 models (issue #22797) with no drop_params required, and the shared
parent transform raising pre-empted that. It also conflicted with the goal of
keeping requests working rather than failing them.
Make the reshape silent: translate effort to legacy thinking for thinking-capable
models, drop thinking for non-reasoning models, and remove only the consumed effort
key from output_config, leaving any residual (e.g. format) for provider subclasses
(bedrock/vertex) to handle. No raise, no drop_params gating. This also resolves the
review note about inconsistent drop_params handling by making every path uniform.
Updates the tests to assert the silent behavior and residual output_config
preservation.
* fix(anthropic): handle output_config-capable but non-adaptive models (Opus 4.5)
Greptile caught a real bug: the early-return guard treated supports_output_config
as equivalent to supporting adaptive thinking. Claude Opus 4.5 advertises
supports_output_config (it accepts output_config.effort) but is not adaptive, so it
rejects thinking:{type:adaptive} with "adaptive thinking is not supported on this
model". The guard early-returned for Opus 4.5 and forwarded the adaptive thinking
block raw, reproducing the exact failure the fix is meant to prevent.
thinking:{type:adaptive} and output_config.effort are independent capabilities.
Only early-return for adaptive-thinking models. For a model that supports
output_config.effort but is not adaptive, keep the native effort and drop only the
unsupported adaptive thinking block. Verified live against Opus 4.5: the Claude Code
payload now returns 200 instead of 400.
Adds regression tests for Opus 4.5 with and without adaptive thinking.
* fix(anthropic): translate adaptive thinking for effort-capable pre-4.6 models
Claude Opus 4.5 advertises supports_output_config but not adaptive thinking,
so the early-return guard forwarded thinking.type=adaptive raw and Anthropic
rejected it. The guard now only skips true adaptive models; effort-only
requests on effort-capable models still pass through untouched. The
_map_reasoning_effort call is wrapped to surface unrecognized effort values
as a clean 400, matching _translate_reasoning_effort_to_anthropic
* fix(anthropic): fall back to legacy thinking when effort level unsupported
Opus 4.5 accepts output_config.effort but only low/medium/high; Claude Code
defaults to xhigh on newer models, so preserving that level raw gets rejected
by Anthropic. Gate the native-effort passthrough on _validate_effort_for_model
and fall through to the budget translation for unsupported levels
* fix(anthropic): keep effort-only requests untouched for provider normalization
The xhigh fall-through consumed effort-only requests on effort-capable
models, breaking bedrock invoke's own normalization which clamps xhigh to
the model's ceiling after the base transform runs
(test_bedrock_messages_normalizes_output_config_effort_for_opus). Restrict
the fall-through to requests that carry adaptive thinking; effort-only
requests pass through so provider subclasses keep owning level clamping
---------
Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com>
vertex_ai/claude-opus-4-8@default (and sibling @default models) were
misclassified as non-adaptive because _model_map_lookup_candidates only
stripped provider prefixes but never the @<suffix> portion. The lookup
produced candidates like ["vertex_ai/claude-opus-4-8@default",
"claude-opus-4-8@default"], neither of which exists in model_cost, so
_is_adaptive_thinking_model returned False. LiteLLM then sent
thinking.type=enabled to a @default Vertex AI endpoint that requires
thinking.type=adaptive, resulting in a 400.
_strip_version_suffix now removes @<suffix> from each candidate,
adding the bare model name (e.g. "claude-opus-4-8") to the lookup
chain. Also adds supports_adaptive_thinking: true to the three
@default model_cost entries that were missing it as belt-and-suspenders.
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
adds a top-level guard in _increment_remaining_budget_metrics that returns early
when all four budget gauges are NoOpMetric (excluded from prometheus_metrics_config),
and per-entity guards in each _set_*_budget_metrics_after_api_request helper for
partial disabling. eliminates four async DB/cache round-trips per successful LLM
request when budget metrics are disabled.
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(responses-api): raise APIError on in-stream error events; widen ErrorEventError.param
- BaseResponsesAPIStreamingIterator._maybe_raise_for_error_event inspects each
chunk and raises litellm.APIError for type=error and type=response.failed events
so callers see an exception instead of a benign stream chunk
- rate_limit* codes map to 429; client error codes (invalid_request_error,
context_length_exceeded, etc.) map to 400; all other codes default to 500;
raw integer codes are never used as-is as HTTP status codes
- ErrorEventError.param widened from Optional[str] to Optional[Union[str, Dict]]
to prevent Pydantic ValidationError on dict-typed param payloads silently
dropping error events before any type inspection
* test(responses-api): add streaming iterator error event tests to CI-covered path
* test(responses-api): cover response.failed, dict-error, null-error, and sync iterator paths
* test(responses-api): set completion_start_time on mock logging objects for internal staging _process_chunk
* fix(responses-api): map insufficient_quota to 429, derive failed-response log status from error code, and record failed-stream usage for spend accounting
insufficient_quota moves out of the 400 bucket; OpenAI returns HTTP 429 for it and the non-streaming exception mapping treats 429 as RateLimitError, so the in-stream mapping now agrees
_handle_logging_failed_response previously hardcoded APIError(status_code=500), so a rate-limited response.failed was logged to integrations as 500 while the caller saw 429; it now shares the same error-code-to-status mapping via _error_event_fields and _status_code_for_error_code
usage carried on a response.failed event is now stashed as combined_usage_object with its computed cost on the logging object before failure handlers run, reusing the mid-stream-interruption spend recovery path (_failure_handler_helper_fn, proxy post_call_failure_hook, _ProxyDBLogger), so failed streams count their billed tokens instead of logging zero cost
dedupe: TestMaybeRaiseForErrorEvent in tests/llm_responses_api_testing duplicated tests/test_litellm/responses/test_streaming_iterator_error_events.py, which is the canonical mirrored location and CI-covered via test-unit-responses-caching-types; the duplicate class is removed
* fix(responses-api): wrap retriable in-stream errors in MidStreamFallbackError and map error type field to status
Mirror chat streaming semantics from _handle_stream_fallback_error: 429 and
5xx in-stream error events now raise MidStreamFallbackError carrying the
mapped APIError so the router's FallbackResponsesStreamWrapper triggers
mid-stream fallback and cooldown; non-retriable 4xx still raise APIError
directly. Status mapping now reads both the OpenAI error type and code
fields, so type-classified client errors (e.g. invalid_request_error with
code invalid_prompt) map to 400 instead of falling through to 500.
* fix(responses-api): accumulate streamed output text so mid-stream fallback continues instead of restarting
MidStreamFallbackError was always raised with generated_content="", so the
router's stream_with_fallbacks treated every mid-stream error as pre-first-chunk
and retried with the original input, streaming duplicated content to clients
that had already received partial output. The iterators now accumulate
response.output_text.delta text (mirroring chat's response_uptil_now) and pass
it as generated_content, letting the router build a continuation input via
_build_responses_continuation_input.
* test(responses-api): pin in-stream token limit error to raised APIError
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(bedrock): gate in-place system role messages on model support for Claude Invoke
* feat(bedrock): default unmapped Claude 4.8+ to in-place system role handling via fallback rule
* fix(responses): preserve reasoning_tokens through chat->responses usage translation
Remove the unconditional else-branch that wrote reasoning_tokens=0 whenever
completion_tokens_details.reasoning_tokens was None or absent. Also change
OutputTokensDetails.reasoning_tokens from int=0 to Optional[int]=None so that
re-instantiation without explicit reasoning_tokens no longer silently zeroes out
the field, and remove the same hardcoded zero from the mock_responses_api_response
initializer.
* test(responses): update assertions to match Optional[int] reasoning_tokens default
* fix(responses): preserve explicit reasoning_tokens=0 in usage translation
Align the reasoning_tokens guard with the is-not-None guards used for
text_tokens and image_tokens: a provider-reported zero passes through
while an absent value stays omitted.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(bedrock): add jp.anthropic.claude-opus-4-8 to model cost map
* test: use apac regional profile for cost-map fallback test since jp now has an entry
* feat(router): add LLM-based classifier option to complexity router
Adds classifier_type: "heuristic" | "llm" to complexity_router_config.
When set to "llm", the router calls a configured model (e.g. a small
model like haiku) via structured output to pick the complexity tier,
falling back to the existing regex/keyword scorer on any error, empty
response, or unparseable output.
* feat(ui): add classifier_type option to complexity router UI, fix edit flow
Adds an "Advanced: Classification Method" section to ComplexityRouterConfig
with a heuristic/LLM toggle, revealing a classifier model picker and timeout
when LLM is selected.
Also fixes the auto router edit modal, which never rendered the complexity
router UI at all (it only handled the semantic router), and the "Edit Auto
Router" button visibility check, which was gated on auto_router_config and
never matched complexity router deployments.
* fix(router): attribute classifier calls to caller, raise default timeout
Forwards the original request's litellm_metadata into the classifier's
acompletion call. Without it, the proxy's cost-tracking gate sees no
user_api_key/team_id/user_id and silently drops spend logging and budget
accounting for every classifier call, letting an authenticated user rack
up unaccounted provider spend via repeated requests.
Also raises the default classifier timeout from 400ms to 3000ms (400ms
undershoots real LLM latency and would silently degrade to the heuristic
scorer on most requests) and corrects the module/class docstrings, which
still claimed zero external API calls after the llm classifier path was
added.
* fix(ci): resolve ruff strict-budget and frontend-lint failures
- Use PEP 585 generics (dict/tuple/list) in the new aclassify/_classify_with_llm
signatures instead of typing.Dict/Tuple/List, and suppress BLE001 on the
intentionally broad except in aclassify's fallback path with a reason.
- Fix prettier formatting in ComplexityRouterConfig.tsx.
- Regenerate eslint-metrics.json (was stale after the classifier UI changes).
* fix(ci): regenerate stale eslint-metrics.json
* fix(router): strip parent budget reservation from classifier metadata
The classifier's internal acompletion call previously forwarded the
parent request's full litellm_metadata, including its budget
reservation (user_api_key_budget_reservation / user_api_key_auth).
That reservation belongs to the routed completion the classifier is
deciding on, not to the classifier call itself, so it's now stripped
while key/team attribution fields are still forwarded for spend
logging.
A trailing slash on --base-url (or LITELLM_PROXY_URL) produced
double-slash URLs like https://host//sso/cli/start, which 404s. Normalize
once in the CLI's top-level group callback so every subcommand benefits.
* fix(proxy): match list/dict guardrail_mode in compliance mode checks
* test(compliance): cover ComplianceChecker guardrail_mode shapes (str/list/dict/None)
* fix(proxy): trust only Mode.default in compliance mode matching (ignore tag overrides)
* fix(proxy): match dict guardrail_mode only when every branch runs in mode (no false-compliant)
* fix(proxy): treat multi-mode guardrail_mode as unresolved (no false-compliant)
The list branch previously counted a guardrail configured with mode:
[pre_call, post_call] under every listed mode. But when the writer cannot
infer the concrete hook that fired (apply_guardrail invocations), the raw
list is logged, and an image-only request that only reaches the post-call
path still records both modes. That let a pre_call compliance check pass on
a request that only ran post_call.
Match the tightened dict semantics: a list now counts for mode only when
every listed mode equals mode. Same trade-off (under-report instead of
false-COMPLIANT). Speculative set support is dropped (spend logs are
JSON-serialized, sets do not cross the wire).
Tests updated to reflect the tightened list semantics, deduplicated (single
TestModeMatching class), and shortened. The invariant is now expressed as
a computed check: True implies every branch runs in the matched mode.
---------
Co-authored-by: Marton Schneider <marton@schneider.co.nl>
* fix(proxy): build redis usage cache from REDIS_* env when cache backend is not Redis
Selecting a semantic (or any non-Redis-KV) response cache left
redis_usage_cache unset, silently downgrading cross-pod rate limits,
parallel-request limits, spend coordination, and the pod lock manager
to per-pod in-memory state. Fall back to a standalone RedisCache built
from REDIS_* environment variables, mirroring the existing
use_redis_transaction_buffer escape hatch, which now shares the same
helper.
Resolves LIT-3861
* feat(proxy): configure the coordination redis independently of the response cache
Adds general_settings.coordination_redis, an explicit block for the Redis
the proxy uses for cross-pod rate limits, parallel-request limits, spend
tracking, the pod lock manager, and shared health checks. Resolution order
is the explicit block, then a plain-Redis response-cache backend, then the
REDIS_* environment. Cluster and sentinel targets are supported, and a
cluster target now builds a RedisClusterCache so cluster-aware consumers
take the cluster path.
Admins can configure it from the Caching page of the dashboard via
/coordination_redis/settings, which reports which source is in effect,
redacts credentials on read, and offers a connection test. Settings saved
there are read back at startup so they take effect on restart.
Also fixes redis client construction so an explicitly configured host
outranks REDIS_URL in the environment. Previously the url branch stripped
the caller's host and port, so an explicit block, or a connection test
typed into the dashboard, silently targeted whatever REDIS_URL named
* fix(ui): move coordination_redis_settings into renamed _components directory
---------
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
* feat(otel): emit the gen_ai.client.operation.exception event on failed LLM calls
The GenAI semantic conventions record failures of a GenAI client operation as
a log-based event named gen_ai.client.operation.exception, carrying the
exception.type / exception.message / exception.stacktrace trio at severity
WARN and correlated to the failed span. OTel v2 never emitted it: a failed LLM
call produced only the deprecated error.* span attributes, a generic exception
span event without a stacktrace, and the stacktrace under the vendor key
litellm.provider.error.stack_trace.
Build the logs pipeline (LoggerProvider + console/OTLP log exporters mirroring
the metrics plumbing) and record the event behind the enable_events flag, which
until now was defined but consumed nowhere. An operator-configured LoggerProvider
global is reused so the events ride their existing logs pipeline; an explicit
NoOpLoggerProvider global is honored as an opt-out and builds no recorder at all.
The existing span-side error surface (error.type, error.message, the exception
span event, and the litellm.provider.error.* detail keys) is untouched for
backwards compatibility.
* fix(otel): always ride the semconv-required exception pair on the GenAI event
Filtering the event attributes on truthiness conflated "absent" with "empty",
so an empty exception.type or exception.message would have been dropped, leaving
an event with neither semconv-required field. Build the attributes so the pair is
unconditional and only the recommended stacktrace is omitted when the payload
carries none.
* docs(otel): document the events plumbing module in the package README
* test(otel): cover the log exporter selection and logs endpoint normalization
The new logs plumbing had no coverage for exporter-kind selection, the
console fallback for an unrecognized kind, the /v1/logs signal-path rewriting
that lets one OTEL_ENDPOINT serve every signal, or the simple-vs-batch
processor split.
Selecting a semantic (or any non-Redis-KV) response cache left
redis_usage_cache unset, silently downgrading cross-pod rate limits,
parallel-request limits, spend coordination, and the pod lock manager
to per-pod in-memory state. Fall back to a standalone RedisCache built
from REDIS_* environment variables, mirroring the existing
use_redis_transaction_buffer escape hatch, which now shares the same
helper.
Resolves LIT-3861
Hoisting every role system entry into the top-level system field mutates
the cache prefix whenever a client such as Claude Code appends a new
mid-conversation system message, invalidating the prompt cache for the
entire message history on Bedrock Invoke. Bedrock only rejects a system
entry at messages.0, so hoist just the leading run and forward the rest
in place
A correctly signed JWT whose user_id or server_id claim was an empty string
passed claims validation but raised ValidationError from the EnvelopeIdentity
constructor inside open_envelope, breaking its never-raises guarantee. The
claims model now mirrors the identity's min_length constraints, so any claim
set that validates also constructs, and the empty-identity case maps to
MalformedPayload like every other bad claim shape.
Pure, unwired module: mints and opens the single client-held bearer that
carries both a litellm identity and the encrypted upstream OAuth grant with
zero server-side storage. HS256 JWT signing (same approach as the BYOK
session bearer) plus the existing encrypt_value/decrypt_value symmetric
helpers, with all key material and the clock injected as parameters. Opening
returns typed frozen error values (not_an_envelope, bad_signature, expired,
malformed_payload, decrypt_failed); minting rejects envelopes over
MAX_ENVELOPE_BYTES with a typed error instead of truncating. Error values
and reprs never carry token material.