* fix(responses-api): raise APIError on in-stream error events; widen ErrorEventError.param
- BaseResponsesAPIStreamingIterator._maybe_raise_for_error_event inspects each
chunk and raises litellm.APIError for type=error and type=response.failed events
so callers see an exception instead of a benign stream chunk
- rate_limit* codes map to 429; client error codes (invalid_request_error,
context_length_exceeded, etc.) map to 400; all other codes default to 500;
raw integer codes are never used as-is as HTTP status codes
- ErrorEventError.param widened from Optional[str] to Optional[Union[str, Dict]]
to prevent Pydantic ValidationError on dict-typed param payloads silently
dropping error events before any type inspection
* test(responses-api): add streaming iterator error event tests to CI-covered path
* test(responses-api): cover response.failed, dict-error, null-error, and sync iterator paths
* test(responses-api): set completion_start_time on mock logging objects for internal staging _process_chunk
* fix(responses-api): map insufficient_quota to 429, derive failed-response log status from error code, and record failed-stream usage for spend accounting
insufficient_quota moves out of the 400 bucket; OpenAI returns HTTP 429 for it and the non-streaming exception mapping treats 429 as RateLimitError, so the in-stream mapping now agrees
_handle_logging_failed_response previously hardcoded APIError(status_code=500), so a rate-limited response.failed was logged to integrations as 500 while the caller saw 429; it now shares the same error-code-to-status mapping via _error_event_fields and _status_code_for_error_code
usage carried on a response.failed event is now stashed as combined_usage_object with its computed cost on the logging object before failure handlers run, reusing the mid-stream-interruption spend recovery path (_failure_handler_helper_fn, proxy post_call_failure_hook, _ProxyDBLogger), so failed streams count their billed tokens instead of logging zero cost
dedupe: TestMaybeRaiseForErrorEvent in tests/llm_responses_api_testing duplicated tests/test_litellm/responses/test_streaming_iterator_error_events.py, which is the canonical mirrored location and CI-covered via test-unit-responses-caching-types; the duplicate class is removed
* fix(responses-api): wrap retriable in-stream errors in MidStreamFallbackError and map error type field to status
Mirror chat streaming semantics from _handle_stream_fallback_error: 429 and
5xx in-stream error events now raise MidStreamFallbackError carrying the
mapped APIError so the router's FallbackResponsesStreamWrapper triggers
mid-stream fallback and cooldown; non-retriable 4xx still raise APIError
directly. Status mapping now reads both the OpenAI error type and code
fields, so type-classified client errors (e.g. invalid_request_error with
code invalid_prompt) map to 400 instead of falling through to 500.
* fix(responses-api): accumulate streamed output text so mid-stream fallback continues instead of restarting
MidStreamFallbackError was always raised with generated_content="", so the
router's stream_with_fallbacks treated every mid-stream error as pre-first-chunk
and retried with the original input, streaming duplicated content to clients
that had already received partial output. The iterators now accumulate
response.output_text.delta text (mirroring chat's response_uptil_now) and pass
it as generated_content, letting the router build a continuation input via
_build_responses_continuation_input.
* test(responses-api): pin in-stream token limit error to raised APIError
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(bedrock): gate in-place system role messages on model support for Claude Invoke
* feat(bedrock): default unmapped Claude 4.8+ to in-place system role handling via fallback rule
* fix(responses): preserve reasoning_tokens through chat->responses usage translation
Remove the unconditional else-branch that wrote reasoning_tokens=0 whenever
completion_tokens_details.reasoning_tokens was None or absent. Also change
OutputTokensDetails.reasoning_tokens from int=0 to Optional[int]=None so that
re-instantiation without explicit reasoning_tokens no longer silently zeroes out
the field, and remove the same hardcoded zero from the mock_responses_api_response
initializer.
* test(responses): update assertions to match Optional[int] reasoning_tokens default
* fix(responses): preserve explicit reasoning_tokens=0 in usage translation
Align the reasoning_tokens guard with the is-not-None guards used for
text_tokens and image_tokens: a provider-reported zero passes through
while an absent value stays omitted.
---------
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
* fix(bedrock): add jp.anthropic.claude-opus-4-8 to model cost map
* test: use apac regional profile for cost-map fallback test since jp now has an entry
* feat(router): add LLM-based classifier option to complexity router
Adds classifier_type: "heuristic" | "llm" to complexity_router_config.
When set to "llm", the router calls a configured model (e.g. a small
model like haiku) via structured output to pick the complexity tier,
falling back to the existing regex/keyword scorer on any error, empty
response, or unparseable output.
* feat(ui): add classifier_type option to complexity router UI, fix edit flow
Adds an "Advanced: Classification Method" section to ComplexityRouterConfig
with a heuristic/LLM toggle, revealing a classifier model picker and timeout
when LLM is selected.
Also fixes the auto router edit modal, which never rendered the complexity
router UI at all (it only handled the semantic router), and the "Edit Auto
Router" button visibility check, which was gated on auto_router_config and
never matched complexity router deployments.
* fix(router): attribute classifier calls to caller, raise default timeout
Forwards the original request's litellm_metadata into the classifier's
acompletion call. Without it, the proxy's cost-tracking gate sees no
user_api_key/team_id/user_id and silently drops spend logging and budget
accounting for every classifier call, letting an authenticated user rack
up unaccounted provider spend via repeated requests.
Also raises the default classifier timeout from 400ms to 3000ms (400ms
undershoots real LLM latency and would silently degrade to the heuristic
scorer on most requests) and corrects the module/class docstrings, which
still claimed zero external API calls after the llm classifier path was
added.
* fix(ci): resolve ruff strict-budget and frontend-lint failures
- Use PEP 585 generics (dict/tuple/list) in the new aclassify/_classify_with_llm
signatures instead of typing.Dict/Tuple/List, and suppress BLE001 on the
intentionally broad except in aclassify's fallback path with a reason.
- Fix prettier formatting in ComplexityRouterConfig.tsx.
- Regenerate eslint-metrics.json (was stale after the classifier UI changes).
* fix(ci): regenerate stale eslint-metrics.json
* fix(router): strip parent budget reservation from classifier metadata
The classifier's internal acompletion call previously forwarded the
parent request's full litellm_metadata, including its budget
reservation (user_api_key_budget_reservation / user_api_key_auth).
That reservation belongs to the routed completion the classifier is
deciding on, not to the classifier call itself, so it's now stripped
while key/team attribution fields are still forwarded for spend
logging.
Add a RESTful PATCH /team/{team_id} that partially updates a team using RFC 7386 JSON Merge Patch. team_id comes from the path, and metadata is merged with the team's stored metadata instead of being replaced wholesale the way POST /team/update does: an omitted key is preserved, key: null deletes it, and any other value overwrites, recursing into nested objects. Every other field behaves the same as POST /team/update
The handler delegates to the existing update path, so authorization, budget checks, system-managed-key stripping, metadata encryption, cache refresh, and audit logging are shared rather than reimplemented. POST /team/update is untouched, so the change is purely additive
DataTableFilterField rendered a bare native label element. Swap it for the
shadcn Label primitive so the filter drawer's field labels share the same
typography and disabled-state handling as the rest of the form primitives
The workflow-runs Status filter leaked the internal "__all__" sentinel as its
displayed value because Base UI's Select.Value renders the raw value when no
items map or children function is given. Drop the sentinel and use Base UI's
native null handling: a null "All statuses" item plus a placeholder, with an
items map so a real selection renders its capitalized label rather than the
raw status string
The shared DataTable rendered every loading-skeleton cell as one identical
half-width bar, which read as a rigid grid instead of the table beneath it.
Vary the skeleton width per column and add a per-column skeleton shape hint
(text or twoLine) on ColumnMeta so identity columns like the workflow "Run"
cell get a two-line skeleton that matches their real content
A trailing slash on --base-url (or LITELLM_PROXY_URL) produced
double-slash URLs like https://host//sso/cli/start, which 404s. Normalize
once in the CLI's top-level group callback so every subcommand benefits.
Adds server- and client-side column filtering to the shared DataTable via a
filterMode prop that mirrors the existing sorting and pagination modes, a global
search filter, a staged filter drawer ("Apply Filters" / "Reset") built on the
shadcn Sheet primitive (added via `npx shadcn add sheet`), and a table-aware
toolbar (search, refresh, active-filter chips, and a Columns menu). Also adds
skeleton loading rows, an icon empty state, and renders toolbar, table, and
pagination as a single card.
Migrates the two pilot tables onto the design as a proof of concept:
TeamVirtualKeysTable drives server-side filtering through the list API (key alias
via the search box, user via the drawer), and WorkflowRuns filters client-side in
memory. Both keep their existing behavioral tests green, with new coverage for the
filter drawer, the toolbar, the global search, and the server-filter query mapping.
Addresses PR review: _to_dict special-cased LitellmParams and returned an
empty dict for any other pydantic model. Switch to isinstance(value,
BaseModel) so it coerces any pydantic model uniformly (BaseModel is already
imported for the response models), which also drops the now-unused
LitellmParams import. Behavior is unchanged for the current call sites.
The /guardrails/usage/{overview,detail,logs} endpoints resolved guardrails only
from the litellm_guardrailstable Prisma table, so guardrails defined in
config.yaml (which live only in IN_MEMORY_GUARDRAIL_HANDLER) were invisible:
detail 404'd, overview omitted them or rendered them as Custom/Guardrail
orphans, and logs missed their logical-name alias.
Add config-owned accessors (list_config_guardrails, get_config_guardrail_by_id)
to the in-memory handler and use them in the usage endpoints, mirroring the
union/fallback already used by list_guardrails_v2 and get_guardrail_info. Also
preserve guardrail_info when storing a config guardrail (type/description were
dropped at initialize time) and read the Prisma-row / dict / LitellmParams
shapes uniformly.
Resolves LIT-2529
* fix(proxy): match list/dict guardrail_mode in compliance mode checks
* test(compliance): cover ComplianceChecker guardrail_mode shapes (str/list/dict/None)
* fix(proxy): trust only Mode.default in compliance mode matching (ignore tag overrides)
* fix(proxy): match dict guardrail_mode only when every branch runs in mode (no false-compliant)
* fix(proxy): treat multi-mode guardrail_mode as unresolved (no false-compliant)
The list branch previously counted a guardrail configured with mode:
[pre_call, post_call] under every listed mode. But when the writer cannot
infer the concrete hook that fired (apply_guardrail invocations), the raw
list is logged, and an image-only request that only reaches the post-call
path still records both modes. That let a pre_call compliance check pass on
a request that only ran post_call.
Match the tightened dict semantics: a list now counts for mode only when
every listed mode equals mode. Same trade-off (under-report instead of
false-COMPLIANT). Speculative set support is dropped (spend logs are
JSON-serialized, sets do not cross the wire).
Tests updated to reflect the tightened list semantics, deduplicated (single
TestModeMatching class), and shortened. The invariant is now expressed as
a computed check: True implies every branch runs in the matched mode.
---------
Co-authored-by: Marton Schneider <marton@schneider.co.nl>
gemini/gemini-3.1-flash-image, vertex_ai/gemini-3-pro-image, and
vertex_ai/gemini-3.1-flash-image existed in the root pricing JSON but not in
litellm/model_prices_and_context_window_backup.json, leaving deployments with
LITELLM_LOCAL_MODEL_COST_MAP=True unprotected. Copies the root entries into
the backup verbatim and extends the regression test to cover all ten gemini
image models, asserting each exists in the local cost map so a missing backup
entry fails the test instead of passing vacuously
Render a general_settings.coordination_redis block into the litellm-helm
proxy config when the bundled Redis is enabled, gated on a new
redis.coordination.enabled value and skipped when the user already
supplies their own block. Sentinel deployments render sentinel_nodes and
service_name rather than a host/port pair.
Also fixes litellm.redis.serviceName, which gated its sentinel branch on
standalone architecture. The bundled Redis subchart only serves sentinel
in replication mode, and renders no master Service there, so REDIS_HOST
pointed at a Service that never existed for every sentinel user.
Documents the coordination redis in the componentized chart and in the
terraform modules, whose existing REDIS_* exports now feed it directly.
Adds helm-unittest coverage for both charts' redis wiring, which had none
* fix(proxy): build redis usage cache from REDIS_* env when cache backend is not Redis
Selecting a semantic (or any non-Redis-KV) response cache left
redis_usage_cache unset, silently downgrading cross-pod rate limits,
parallel-request limits, spend coordination, and the pod lock manager
to per-pod in-memory state. Fall back to a standalone RedisCache built
from REDIS_* environment variables, mirroring the existing
use_redis_transaction_buffer escape hatch, which now shares the same
helper.
Resolves LIT-3861
* feat(proxy): configure the coordination redis independently of the response cache
Adds general_settings.coordination_redis, an explicit block for the Redis
the proxy uses for cross-pod rate limits, parallel-request limits, spend
tracking, the pod lock manager, and shared health checks. Resolution order
is the explicit block, then a plain-Redis response-cache backend, then the
REDIS_* environment. Cluster and sentinel targets are supported, and a
cluster target now builds a RedisClusterCache so cluster-aware consumers
take the cluster path.
Admins can configure it from the Caching page of the dashboard via
/coordination_redis/settings, which reports which source is in effect,
redacts credentials on read, and offers a connection test. Settings saved
there are read back at startup so they take effect on restart.
Also fixes redis client construction so an explicitly configured host
outranks REDIS_URL in the environment. Previously the url branch stripped
the caller's host and port, so an explicit block, or a connection test
typed into the dashboard, silently targeted whatever REDIS_URL named
* fix(ui): move coordination_redis_settings into renamed _components directory
---------
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
A plain optional label implies the field is persisted like any other form
field; the qualifier was only there to contrast with the removed not-saved
wording
* feat(otel): emit the gen_ai.client.operation.exception event on failed LLM calls
The GenAI semantic conventions record failures of a GenAI client operation as
a log-based event named gen_ai.client.operation.exception, carrying the
exception.type / exception.message / exception.stacktrace trio at severity
WARN and correlated to the failed span. OTel v2 never emitted it: a failed LLM
call produced only the deprecated error.* span attributes, a generic exception
span event without a stacktrace, and the stacktrace under the vendor key
litellm.provider.error.stack_trace.
Build the logs pipeline (LoggerProvider + console/OTLP log exporters mirroring
the metrics plumbing) and record the event behind the enable_events flag, which
until now was defined but consumed nowhere. An operator-configured LoggerProvider
global is reused so the events ride their existing logs pipeline; an explicit
NoOpLoggerProvider global is honored as an opt-out and builds no recorder at all.
The existing span-side error surface (error.type, error.message, the exception
span event, and the litellm.provider.error.* detail keys) is untouched for
backwards compatibility.
* fix(otel): always ride the semconv-required exception pair on the GenAI event
Filtering the event attributes on truthiness conflated "absent" with "empty",
so an empty exception.type or exception.message would have been dropped, leaving
an event with neither semconv-required field. Build the attributes so the pair is
unconditional and only the recommended stacktrace is omitted when the payload
carries none.
* docs(otel): document the events plumbing module in the package README
* test(otel): cover the log exporter selection and logs endpoint normalization
The new logs plumbing had no coverage for exporter-kind selection, the
console fallback for an unrecognized kind, the /v1/logs signal-path rewriting
that lets one OTEL_ENDPOINT serve every signal, or the simple-vs-batch
processor split.
Render DcrBridgeToggle inside PassthroughAuthorizeSection, after the OAuth
client ID/secret fields and just before the Authorize & Fetch Tools button,
in both the create and edit flows. Also update the section copy to say a
configured OAuth app is saved with the server, using the same wording as the
credential lifecycle rework in #32752 so whichever PR lands second rebases
cleanly
Selecting a semantic (or any non-Redis-KV) response cache left
redis_usage_cache unset, silently downgrading cross-pod rate limits,
parallel-request limits, spend coordination, and the pod lock manager
to per-pod in-memory state. Fall back to a standalone RedisCache built
from REDIS_* environment variables, mirroring the existing
use_redis_transaction_buffer escape hatch, which now shares the same
helper.
Resolves LIT-3861
Addresses review nits on the sidebar shell. The Meter primitive now owns its bar color through a tone variant, so the usage card passes a tone instead of a bg-* class, and the near-limit tone goes back to 80% to match the previous usage indicator. The enterprise usage card is now explicitly gated on an active license, so it never renders for an unlicensed proxy rather than relying on seat/team data being null. The collapsed-rail control and the sidebar menu button render through the shadcn and Base UI button primitives instead of raw button elements; the menu button gains render/nativeButton passthrough while its group toggles stay native buttons, so the nav link leaves keep their link semantics. Adds a regression test that the card stays hidden without a license even when seat limits exist
The backup file is used by tests; the root model_prices_and_context_window.json
is what gets published to the pricing URL and loaded by the proxy at runtime.
Without this, the proxy would continue resolving supports_reasoning via the
provider-level fallback and returning true for Gemini image generation models.
Also covers vertex_ai/gemini-3-pro-image and vertex_ai/gemini-3.1-flash-image
(non-preview variants) and gemini/gemini-3.1-flash-image which exist only in
the root JSON.
vertex_ai/gemini-2.5-flash-image, vertex_ai/gemini-3-pro-image-preview,
vertex_ai/gemini-3.1-flash-image-preview, gemini/gemini-3-pro-image-preview,
and gemini/gemini-3.1-flash-image-preview were missing supports_reasoning
entries; _supports_factory then fell through to the vertex_ai provider-level
config which returns true, causing requests with reasoning_effort to be sent
to an API that rejects them.
The guard-main-branch error messages and the contributor docs still
pointed people at litellm_oss_staging. Redirect them to the current
daily OSS branch (litellm_oss_daily_YYYY_MM_DD), a fresh one of which
is cut each weekday, so contributors should target the most recent
Hoisting every role system entry into the top-level system field mutates
the cache prefix whenever a client such as Claude Code appends a new
mid-conversation system message, invalidating the prompt cache for the
entire message history on Bedrock Invoke. Bedrock only rejects a system
entry at messages.0, so hoist just the leading run and forward the rest
in place