client_ack_messages classified a websockets ConnectionClosed raised by
the client socket as the backend closing, so bidirectional_forward kept
waiting on the upstream instead of ending the session. Starlette clients
raise WebSocketDisconnect, but the realtime test client in
tests/llm_translation/realtime raises websockets.exceptions.ConnectionClosed,
which hung test_openai_realtime_simple.py until the run was killed.
Only the receive_text call now maps every exception to
CLIENT_DISCONNECTED; the loop body keeps ConnectionClosed as
BACKEND_CLOSED, since the backend socket is the only websockets socket
touched there.
The detail endpoint now returns untracked_usage_units_by_team and
untracked_usage_units_by_key next to the cost breakdowns, and the By team and
By key tables show them in an Unpriced Units column, so a row that pairs its
total units with a partial cost says how many units that cost leaves out.
The overview comparator no longer treats a missing cost as zero: guardrails
with no known cost sort last in both directions instead of mixing in with
genuinely free ones.
Refs LIT-5652
The overview table gains Usage Units and Cost columns plus a Guardrail Cost
card, and the detail page gains a Usage & Cost section that breaks units and
cost down by counter, team and key. Units the cost map could not price are
called out next to the cost they are left out of.
Both pages now read /guardrails/usage/* through $api.useQuery so the rows are
typed from schema.d.ts; the hand-written PerformanceRow and the untyped fetch
helpers are gone. fetchClient resolves fetch per request so integration tests
that stub the global see typed-client calls too.
Refs LIT-5652
When the upstream closes while the proxy is forwarding a client message,
the client loop ends before the backend relay sees the close, and the
relay skipped closing the client because it read the client loop's exit
as the client hanging up. The client loop now reports why it stopped, so
a close observed on the backend send still reaches the client with the
error event and the upstream close code
When the provider closes the realtime websocket (for example Vertex Live
refusing the session with 1008 "Publisher model ... was not found"), the
proxy swallowed the close and kept waiting on the client, so the client
sat on an open socket with nothing coming back and the session was logged
as a $0 success
The backend relay now returns the upstream close, and bidirectional_forward
sends the client an OpenAI-style error event naming the upstream code and
reason, then closes the client socket with the same code (or 1011 when the
upstream code is one a server may not send). A session the upstream refused
before sending any frame is logged through the failure handlers instead of
as a success
OpenAI and Azure realtime usage reports output_tokens == text_tokens + audio_tokens
with reasoning_tokens counted inside text_tokens, so generic_cost_per_token billed
the reasoning share twice. When the output token details sum past completion_tokens,
the nested reasoning overlap is now subtracted from text_tokens before pricing;
shapes where text_tokens already excludes reasoning are unchanged.
The classifier scores extracted text, so a turn whose complexity lives in
its image is invisible to it: a screenshot of a stack trace classifies on
its caption, and an image-only turn flattens to empty text and never
reaches the classifier at all.
classifier_llm_config.vision opts in, off by default, with max_images
bounding what one turn can add. Images are still dropped when the
classifier model is declared supports_vision false. Anthropic and
Responses image parts are rewritten into chat-completions dialect before
they reach the classifier call, since /v1/messages hands the pre-routing
hook its own dialect untranslated.
The local scorer no longer short-circuits heuristic_first or hybrid on a
turn carrying forwarded images, because it reads text alone and its
confidence describes a request it has only partly seen.
The completed-batch early return skipped both the cancel and the list
assertion while the lifecycle's covers markers still credited both cells.
List does not depend on the batch being cancellable, so it now runs either
way; cancel on a completed batch stays a documented vacuous pass
* fix(shadow_eval): tell a tool-call shadow reply apart from an empty one
Both arrive at the attempt row as the same 'shadow router returned an empty
response', because _chat_final_text returns empty for a tool-final turn by
design and for a reply that genuinely carried no text. Those are different
things: an arm that chose a tool where the real model wrote prose is a
divergence a text judge cannot score, and the sampling side already drops the
real arm's tool-final turns for exactly that reason, so the shadow side reads
as a fault where the real side reads as a filter. A job that is almost all
'empty response' gives no way to tell a tool-happy arm from a broken one.
The error now names which of the two happened, and carries the finish_reason
and the routed model so the row says what the arm was doing. Every varying
part sits behind the first semicolon: operators read these by grouping on the
error text, and interpolating the model into the leading sentence would make
each row its own group.
The outcome stays 'error'. Whether a tool-call reply should instead be its own
non-judged outcome, excluded from the loss rate the way the real arm's
tool-final turns already are, needs the four aggregation predicates that spell
judged as outcome != 'error' rewritten, and a decision on how to surface the
new bucket. That is a separate change.
* fix(shadow_eval): read the tool name of a custom tool call
A custom tool call carries its name under custom.name with no function key,
so every one of them reported as tool=unnamed.
* feat(shadow_eval): judge tool calls instead of dropping the turn
A turn where either arm called a tool was discarded before it could be
compared: the real arm's at sampling, the shadow arm's as an error row. On
agentic traffic that is most of the traffic, so a job set to sample 10% was
sampling 10% of the prose-only slice. Tool calls now serialize to text on
every surface and are judged like any other response, and the judge is told
a tool call is not a defect so it scores the choice rather than the shape.
* feat(shadow_eval): show the judge what tools were available
Both arms were offered the same tools, but the judge only ever saw the
chosen call in isolation, with no way to tell whether a better tool existed
or the arguments matched what the tool expects. Threads the request's tool
definitions (name and description only) into the judge prompt, capped and
omitted entirely on turns that offered none.
* fix(shadow_eval): read a custom tool definition's name from custom, not function
A chat-completions custom tool definition nests name and description under
custom, mirroring how a custom tool call nests them (openai.types.chat.
ChatCompletionCustomToolParam). Reading only function rendered every one as
unnamed, telling the judge nothing about what it was.
Bedrock passthrough Converse routes flattened every non-empty string under
toolConfig.tools into the guardrail INPUT texts, so tool names, tool
descriptions and JSON-schema strings (object, property names, titles, type
names, enum values) each arrived as a separate guardrail item. A request whose
only prompt was one benign user message could be blocked outright because a
denied term appeared in an app-authored tool definition.
Tool definitions are now excluded from the extracted texts, matching every
other guardrail translation handler, which carries tool definitions in the
structured tools input rather than in texts. Caller content stays scanned:
message text, toolUse.input, toolResult content and json, and
additionalModelRequestFields are unchanged.
Resolves LIT-5797
Two fixes for the zero-headroom basedpyright budget:
- arm_pre_call's data parameter is dict[str, object], not MutableMapping: the
latter is itself banned by LIT001 with no benefit, and it mismatched every
dict-typed helper (get_or_create_metadata_bucket, resolve_structured_messages,
_get_tags_from_request_kwargs), which is what the budget was actually flagging.
- Router.async_pre_routing_hook computed pre_routing_hook_response in one shot
instead of reassigning a Final-annotated local.
The remaining two reportArgumentType hits are pre-existing: LiteLLM_Params(**merged)
in _create_deployment_object already fails this check for all ~165 of its other
fields, since the merged dict's value type is partly untyped/float; adding two new
string fields to the model just grows that existing pile by two. Suppressed at the
one call site with a reason, since fixing the root typing is out of scope here.
Bedrock batch cancel (StopModelInvocationJob) and the managed list view
both work through the proxy since LIT-4774, but the batches e2e still
gated them off and the coverage registry claimed no cell for either.
Flip can_cancel/can_list for the Bedrock provider, assert cancel the
same way the OpenAI leg does, add the two registry cells the gates
select, and update COVERAGE.md
* fix(datadog_llm_obs): keep the guardrail audit record under message redaction
Redaction nulled `guardrail_information` on the span whole, so an operator
running `turn_off_message_logging` (or a caller sending
`x-litellm-enable-message-redaction`) lost the record of which guardrails ran,
what they returned, and what they masked. Four of the record's fields can quote
the prompt; the rest report what the guardrail decided without reproducing it.
Replace only those four, the way
`_sanitize_guardrail_information_for_spend_logs` already does for spend logs,
and declare the field list once in `litellm/types/utils.py` so both readers
share it.
* fix(datadog_llm_obs): keep a lone guardrail record, and test through the span
Review round 1.
A guardrail that writes the metadata key itself leaves a single record where
the type says list, which Prometheus already normalizes at
`_guardrail_overhead_seconds`. Redaction dropped that shape and the latency
extraction raised on it, so the span was lost outright. Normalize once and use
it in both places.
The new tests now drive `create_llm_obs_payload` instead of reading the module's
private helpers and the record's declared field names.
Gemini AI Studio file URIs under generativelanguage.googleapis.com/v1beta/files/
answer 403 when fetched and must pass through as file_data.file_uri, which the
sync transform already did. Give the shared walker a skip_url_prefixes parameter,
pass that prefix from the Gemini body builder, and skip it in the AI Studio
message transform for both image_url and file parts