The previous commit added two branches to apply_guardrail and pushed it from 14 to
16 against ruff-strict's max-complexity of 15, taking the codebase C901 total from
312 to 313 with the budget at 312. It was one branch under the ceiling, so the
check had to live somewhere else.
Both conditions now sit in _refuse_file_backed_images: input_type, because images
are a request-side concern, and the count. apply_guardrail is back to 14 and the
call site reads as one statement.
Worth recording why this took two tries to see. `ruff check --select C901` uses
pyproject.toml, where max-complexity is 10, and apply_guardrail was already over
that both before and after -- so a before/after comparison showed no change and I
read the failure as a local artifact problem. The gate runs ruff with
--config ruff-strict.toml, where the ceiling is 15, and that is the only threshold
the budget is counted against.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An Anthropic `{"type": "image", "source": {"type": "file", "file_id": ...}}` block
carries no bytes, so the extractor yields nothing for it while the provider still
forwards the file to the model. The docstring called that a known gap. Documented
is not the same as safe, and a silent pass is the exact failure this whole path
exists to remove, so it is now refused: blocked by default, allowed only where the
operator sets on_unscannable_image.
Resolving the file id would need a Files API client this guardrail does not have,
which is a larger change than the one this PR is making.
The check reads inputs["structured_messages"], which on /v1/messages carries the
raw Anthropic blocks, so the file source is visible to the guardrail without
altering GenericGuardrailAPIInputs. Putting a marker in inputs["images"] was the
alternative and would have reached five other guardrails that consume that field,
trading one gap for four new unknowns.
It detects the file shape specifically rather than comparing an image count against
inputs["images"]. structured_messages is already narrowed by the skip and scope
flags, so a mismatch is not by itself evidence of a dropped image, and refusing a
legitimate request would be worse than the gap being closed. A test pins that
base64 and url sources are still accepted.
Placed before the incremental and latest-message-only shortcuts, so a file-backed
image cannot be skipped by them either.
The first draft of the refusal test passed for the wrong reason: without the check
the request died on absent AWS credentials rather than on the bypass. The Bedrock
call is now stubbed, so removing the check makes it fail with DID NOT RAISE -- the
silent pass itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CI's basedpyright budget moved with the base and surfaced 12 new
reportArgumentType errors, ten of them on one line.
`request_kwargs` was typed `dict[str, object]` to keep an earlier gate happy, but
`object` does not unpack into httpx's typed parameters: `stream()` names ten of
them (content, data, files, params, headers, cookies, auth, follow_redirects,
timeout, extensions) and every one was an error. `dict[str, Any]` is what the
values actually are -- whatever `async_safe_get`'s caller passed through -- and it
lets the call be checked instead of merely tolerated.
The other two are the OUTPUT branch's `enumerate(filtered_messages)`. Making
skip_scan conditional on `image_urls` removed pyright's narrowing there, since
skip_scan used to be the only way that name could still be None. It cannot be None
at runtime -- images exist on the request side only, so a response scan returns
early exactly as before -- but the branch now says so with `or ()` rather than
resting on that indirection.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The base gained LIT010 (assignment without a Final declaration) while this PR was
open, so code that passed the gate at the old base now trips it.
Three are genuinely never rebound and take Final: `per_message`, `chunks`, and the
awaited fetch result, which is now bound separately so the unpack below has room
for its reason.
Two are real rebinding and say so. `total` is the running byte count the cap is
measured against. `mime_type` is overwritten further down when the caller passes
`format`, so it cannot be Final; the reason sits on the unpack line itself, since
a reason on the preceding line does not count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The base moved 478 commits while this PR was open. One conflict, in
test_litellm_core_utils_prompt_templates_factory.py: #38169 added
`from typing import Final` to the import block where this branch added
`import socket`. Both kept; the tests each side appended do not overlap.
Merged rather than rebased so the commit SHAs referenced in this PR's
description and review replies keep resolving.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The capped path was only ever exercised through the Bedrock guardrail, so the two
shared files it lives in were the last uncovered lines on this PR. Both now have
tests where the code does, not only where its first caller does.
url_utils: _underlying_httpx_client's TypeError had no test at all. Every existing
one goes through AsyncHTTPHandler, whose `.client` is a real httpx client, so the
isinstance guard always held. It exists because `cast` is banned here, and without
it the declared return type would be a lie. Also covered: aborting mid-transfer,
asserting on how many bytes were pulled rather than only that it raised, and the
rebuilt response dropping content-encoding and content-length -- aiter_bytes
yields decoded bytes, so carrying those over would describe the body wrongly.
factory: get_image_details_async's body never ran under test. The model-path tests
stub the whole method out, and the guardrail tests fake the transport underneath
it. One test drives the real method with a cap; the other pins the reason the
previous commit exists, calling process_image_async with no cap against a stub
written to the old one-parameter signature.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The PR's position is that anything it cannot recognise must not reach the model
unscanned, but nothing asserted that. Every unrecognised shape landed on an
untested `return None`, so a refactor letting one fall through to the text branch
would have been reported to the operator as scanned content.
One parametrized test walks the shapes a caller can actually put on the wire --
absent content, content that is not a list, a part that is not a mapping, an image
part with no url, a url that is neither string nor mapping, a part carrying neither
image nor text -- and asserts none of them becomes a content item. Stating the
contract once beats restating each branch.
The rest cover paths that are behaviour rather than defence:
- `image_url` as a bare string, which OpenAI accepts alongside `{"url": ...}` and
which reaches the model as an image either way
- a bare base64 payload whose format cannot be sniffed, left to
on_unscannable_image instead of guessed at. It goes through apply_guardrail,
the only caller that normalizes before decoding
- an oversized inline image under `allow`. Under `block` that branch raises and
never returns, so this is the only path reaching its fall-through
_get_image_url and _retained_image_bytes are exercised directly. Both are reached
only through callers that already checked the shape, so their guards are otherwise
unreachable and would be dropped in a refactor without anything noticing.
Added lines in bedrock_guardrails.py now measure at full coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The savings card carried four numbers in two stacked halves: the headline
saving with its delta on the left over the two spend rows, and avg saved per
session on the right. Give the headline the whole left half, move the two
spend rows into a rail on the right, and drop avg saved per session into the
metric row below as its first tile, with the session count as an inline hint.
Each spend row stays a description list so assistive tech keeps the label to
value association, with the shadcn Separator between the two rows. Both hero
columns are minmax(0,1fr) so a large total wraps instead of overflowing the
card, which also fixes the clipping the old 1fr columns already had. Metric
grows one optional hint slot so the new tile reuses the same presenter as its
three siblings.
Capping each image at 4 MB does not bound a request.
_create_bedrock_input_content_request gathers over every message and
_build_input_content_items then gathers over every part, so the fetches all start
together, and _MAX_IMAGES_PER_APPLY_GUARDRAIL_CALL is not consulted until
_bin_pack_bedrock_content runs on items that are already resident. 200 urls at
4 MB is 800 MB, chosen by the caller.
Two request-scoped bounds. A byte budget of _MAX_IMAGE_BYTES times
_MAX_IMAGES_PER_APPLY_GUARDRAIL_CALL, held for one content-request build and
passed down rather than kept on the guardrail, which is a callback instance shared
by every request. And a semaphore of 4, so the reserved bytes are not all in
flight at once.
Each fetch reserves a cap and refunds what a usable image did not take. A response
that decoded to nothing is charged in full: the transfer happened, and refunding it
would let a url serving megabytes of junk be repeated down the whole list for free,
which is the exhaustion being guarded against. An earlier draft refunded the whole
reservation and bounded only in-flight bytes; the regression test caught that the
decoded images still accumulated without limit.
Inline base64 draws on neither budget nor gate. Those bytes arrived in the request
body the proxy already accepted, so charging them to a download quota would refuse
inline images for no reason.
The total is what a single ApplyGuardrail call would accept anyway (20 images at
4 MB), so no request the API would take in one call is refused. A conversation
chunked across several calls can exceed it, and images past the budget are then
unscannable and left to on_unscannable_image, which blocks by default.
Without the budget the regression test fetches 209,715,200 bytes for one request.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`process_image_async` passed `max_bytes=` on every call. The parameter has a
default, so the signature stayed compatible, but the *call* did not: an override
or test stub written against the previous signature now gets an unexpected
keyword and raises.
test_url_with_format_param caught it. It stubs `get_image_details_async` with a
one-parameter fake, so the model call died on TypeError and the provider mock it
asserts on was never reached.
Omitting the keyword when there is nothing to cap leaves the model-call image
paths byte-for-byte as they were, which is what the parameter being additive was
supposed to mean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(ui): carry a preset's per-tier litellm_params through the prefill
buildPresetPrefill rebuilt the complexity router config field by field and
never emitted tier_model_params, so a bundled preset that declares per-model
litellm_params (reasoning_effort, for instance) lost them before the create
form ever saw them. Both halves of the round trip already existed:
hydrateTierModelParams reads either storage shape, and serializeTierModelConfigs
writes them back on submit.
Hydrating alone is not enough. Tier entries get rewritten to the caller's
registered model spelling, which can differ from the preset's literal string by
version-separator punctuation, while the params stay keyed on what the preset
spelled. serializeTierModelConfigs then drops any param whose key is not in the
tier, silently. The param keys go through the same resolver as the tier entries.
* test(ui): catch a preset spelling the same model two ways in one tier
buildPresetPrefill resolves every model reference through normalizeModelName,
so two spellings of the same model in one tier (e.g. "claude-sonnet-4-5" and
"claude-sonnet-4.5") collapse to one key. For tier_model_configs that means one
model's litellm_params silently overwrites the other's - flagged by Greptile
on #38453 (P2, confirmed real via a throwaway repro, not a regression: on the
merge base both param sets were already dropped).
Nothing else validates preset authoring, and these are trusted, checked-in
JSON, so the fix is a static test over the bundled data rather than runtime
code. Exports normalizeModelName so the test exercises the actual resolution
rule instead of a hand-rolled copy of it. Verified the test fails when a
preset is mutated to spell one model two ways, and passes clean on the real
bundled presets.
* feat(newrelic): per-team cost and usage metrics via team callbacks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(newrelic): retry transient 429/408 metric posts instead of dropping
* fix(newrelic): drop only records queued when the drain began, not mid-drain arrivals
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
`test_gemini_chat_returns_content_and_logs_cost` asks gemini-2.5-flash to
"reply with the single word pong" under `max_tokens=32`, and has been seen
returning no content at all:
completion_tokens=29, reasoning_tokens=29, content=None
gemini-2.5-flash defaults to dynamic thinking, and `max_tokens` maps to
`maxOutputTokens`, which on the 2.5 family counts thinking tokens as well as
visible output. So the model is free to spend the entire budget on thoughts and
emit nothing, which is exactly what the usage above shows.
Raising the limit alone does not fix this. Dynamic thinking on 2.5 Flash is
documented up to 24576 tokens, so no budget small enough to be reasonable for a
one-word smoke test is safe. The fix is to take thinking out of the picture:
`reasoning_effort="none"` maps to `thinkingConfig.thinkingBudget=0` for the 2.5
family, so the whole limit is available to visible output. Verified against this
checkout:
get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini",
max_tokens=32)
-> {'max_output_tokens': 32} # no thinkingConfig at all
get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini",
max_tokens=64, reasoning_effort="none")
-> {'max_output_tokens': 64,
'thinkingConfig': {'thinkingBudget': 0, 'includeThoughts': False}}
This mirrors what the OpenAI tool tests in this same file already do with
gpt-5.6 for the same failure mode. `max_tokens` goes to 64 for headroom; with
thinking disabled that is ample for a one-word answer.
Neither `covers` claim changes: the call still exercises the gemini chat
translation path and still produces a costed SpendLogs row.
`_cacheable_system_block` embedded the per-run marker in all 300 paragraphs, so
the block's token count moved with the marker's own tokenization. Measured over
40 random markers the size ranged 3611-5408 tokens (median 4509): 15% of runs
landed under the 4096-token minimum cacheable prefix of Haiku 4.5, despite the
docstring claiming the prompt was comfortably above it.
When the system block is under the minimum, no cache entry is written at the
system breakpoint. The entry at the second breakpoint still gets written,
because system + first user turn clears the minimum -- which is why the failures
report a large cache_creation with cache_read stuck at 0
(`cache_creation_input_tokens=5610 cache_read_input_tokens=0`, and 5610 is the
whole prefix, not the user turn's share). `_prime_prompt_cache` rotates the user
turn on every attempt, so that second entry never prefix-matches the next
attempt either. Every attempt re-creates the full prefix, cache_read never rises
above 0, and the loop burns its 60s deadline:
prompt cache never became readable in full within 60.0s
That is the single most frequent flake in the e2e suite, 9 of 38 runs, and it
hits all three provider classes identically because they share this helper.
Move the marker out of the repeated paragraph so it appears once, and size the
block at 1500 paragraphs. The prefix is now 8056-8060 tokens across markers --
spread 4 tokens instead of 1797, and 1.97x the minimum in the worst case. The
same marker-per-repetition pattern in `_first_turn_user_text` is fixed the same
way. Both copies of the helpers stay byte-identical.
The keywords feed the scorer's technical dimension, so they change tier decisions
on any router that scores. The control rendered only for classifier_type
'heuristic', while the scoring knobs right below it already gated on
heuristicScoringRole(value) !== 'never'. The two disagreed, so an operator could
edit boundaries and weights on a router whose keywords they could neither see nor
set.
That hid the control on an LLM classifier using the default heuristic fallback,
and on heuristic_first, which runs the scorer on every request to decide whether
to short-circuit. Both now read the same predicate as the panel below them.
Google withdrew gemini-live-2.5-flash-preview-native-audio-09-2025 from the
Vertex Live API. Every session dies at setup:
received 1007 (invalid frame payload data)
gemini-live-2.5-flash-preview-native-audio-09-2025 is not supported in the live api.
The client sees session.created (the proxy synthesizes it on connect) and then
nothing, so both vertex_ai realtime tests time out waiting for session.updated.
Confirmed by probing the Vertex Live endpoint directly with the e2e stack's own
credentials:
gemini-live-2.5-flash-preview-native-audio-09-2025 -> 1007, not supported
gemini-live-2.5-flash-native-audio -> setupComplete
so this swaps to the non-preview sibling, which is the same native-audio class
and is what the cost map already carries for vertex_ai.
Not a litellm regression. The suspicion fell on #38395 because it removed the
native-audio speechConfig strip, but the setup payload this suite sends is
byte-identical either side of that change: the strip only fires when a client
sends a voice, and the e2e SessionConfig has no voice field. Google's rejection
names the model, not a field.
The gemini (Google AI Studio) provider keeps the -09-2025 id, which still works
there; only the Vertex endpoint dropped it.
* feat(ui): add Teams list CSV export with budgets, model grants, and rate limits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): neutralize formula-leading values in teams CSV export
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.
Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].
The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.
Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
_build_image_content_item hands a caller-supplied url to
BedrockImageProcessor.process_image_async, which buffers the whole body through
async_safe_get before returning. The 4 MB check runs on bytes that are already
resident, so it rejects an oversized image without preventing the allocation --
and _build_input_content_items gathers these concurrently, so one request with
several urls multiplies it. An arbitrarily large or indefinitely chunked response
is enough to exhaust proxy memory.
async_safe_get takes an optional max_bytes and, when given one, streams the body
and aborts past the cap with PayloadTooLargeError. Omitted, it keeps the previous
buffering, so every existing caller -- including the model-call image paths in
factory.py -- is byte-for-byte unchanged. Only the guardrail passes it.
The rebuilt response drops content-encoding and content-length: aiter_bytes
yields decoded bytes, so carrying those over would describe the body wrongly.
PayloadTooLargeError subclasses ValueError, like SSRFError, so callers already
treating a bad remote response as a rejected fetch need no new except arm. The
guardrail names it explicitly anyway, so an operator reading the log sees "too
large" rather than "could not be read".
The two existing remote-url tests now stub `stream` rather than `get`, which is
the transport the capped path uses. The new test asserts on how many bytes were
pulled -- without that, an unstubbed transport would raise for the wrong reason
and the test would pass against the unfixed code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
apply_guardrail reads inputs["images"], but two optimizations return before that
point and both decide what to scan from `texts` alone:
_apply_incremental_request_scan (only_scan_new_messages) sends just the new text
segments and returns. Benign text plus a policy-violating image, on a session
whose text is already cached, is never scanned -- and the proxy still reports the
guardrail as having run. Nothing caches images either, so "already seen" cannot
be established for them in the first place.
_select_messages_for_apply_guardrail (experimental_use_latest_role_message_only)
marks a latest user message with no text as skip_scan. An image-only message is
exactly that shape, so the whole request was dropped from the scan.
An image now forces the full image-aware path in both cases. The optimization is
lost for image-carrying requests -- an incremental turn with an image rescans its
text rather than skipping it -- which is the safe direction: correctness over the
optimization, failing closed rather than silently open.
Falling through skip_scan leaves filtered_messages None; the existing
`*(filtered_messages or ())` unpack and _merge_masked_texts's empty-input guard
already handle that, so the image-only scan needs no other change.
Both tests fail on the code without this change, for the reason named in each.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(complexity_router): heuristic-first classifier chaining
Adds classifier_type 'heuristic_first', which scores locally on every request and
only calls the LLM classifier for traffic the scorer could not place at or below
heuristic_first_max_tier. A request short-circuits when the scorer landed at or
below the threshold and produced at least one signal; everything else escalates.
The signal requirement is load-bearing. A prompt where no dimension fires scores
exactly 0.0, which is under simple_medium, so the score-to-tier mapping calls it
SIMPLE by default rather than by evidence, and that is about half of general
traffic. Gating on the tier alone would route it to the cheapest model without
ever consulting the classifier.
Introduces uses_llm_classifier as the single owner of 'does this router call the
classifier model', replacing the classifier_type == 'llm' comparisons in the
config validator, the prompt prebuild, the health dependency graph, the
routing-test authorizer, and six dashboard sites.
* fix(complexity_router): reuse the heuristic verdict on classifier failure, load the threshold on edit
Three review findings, one push.
The heuristic-first fallback re-scored the prompt after a classifier failure,
which the README already documented as a reuse. The outcome computed before
escalation is now handed to the failure path, so the scorer runs once per request.
The edit modal never hydrated heuristic_first_max_tier, while save rebuilds every
managed key from form state, so opening a heuristic-first router and saving it
dropped a field the proxy requires. The dropdown's display fallback hid it. Both
are fixed, and the hydration is extracted into a pure function so a test can pin
the invariant: every managed key present in a stored config survives an untouched
open-and-save. That test also covers every field added later.
Classifier radio labels lost their em dashes, per the repo writing convention.
AWS caps images at 4 MB each and 20 per request
(https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails-mmfilter.html).
Nothing checked either; BEDROCK_APPLY_GUARDRAIL_CHUNK_BUDGET_CHARS is the text-unit
quota and says nothing about images. An oversized image is now reported through
on_unscannable_image with the measured size instead of surfacing as an opaque AWS
error.
_bin_pack_bedrock_content measured only text, so an image counted as 0 and
`used + 0 <= budget` always held: 45 images packed into a single batch of 45. That
measurement was complete when a content item could only be text; putting images in
the payload is what invalidated it. Packing now carries a second dimension for the
image count. Image bytes are deliberately not charged against `budget`, which is a
different quota.
_apply_guardrail_content_with_chunking splits up front when the image count is over
the limit. Chunking is otherwise reactive, and the substrings
_is_input_too_large_error matches ("text unit", "too long", ...) are all text-shaped,
so an image-count rejection may never reach that fallback. Bisection could not
rescue it either: _split_bedrock_content halves a lone item by its "text", which is
empty for an image, so it gives up and re-raises the original error. Recursion
terminates because every batch _bin_pack_bedrock_content returns is already within
the image limit.
Nested images needed no work here: _extract_tool_result already collects them out of
tool_result blocks, so with apply_guardrail reading inputs["images"] an image inside
a tool result is scanned.
Four tests. The oversized case builds a real 5 MB data URI rather than patching the
decoder, so the size check runs against what the decoder actually produces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
_image_sources returned source["data"] only. An Anthropic image block has three
shapes (types/llms/anthropic.py:259) and only the base64 one carries "data", so
{"type": "url", "url": ...} yielded nothing and the image never reached any
guardrail at all.
This is not Bedrock-specific. Five guardrails consume
GenericGuardrailAPIInputs["images"] (vigil_guard, custom_code, deepkeep, straiker,
generic_guardrail_api) and every one of them was blind to url sources on
/v1/messages.
base64 now returns a data URI rather than the bare payload. A consumer otherwise
has no way to recover media_type, and an API like Bedrock's ApplyGuardrail needs
the format to build its request.
The file shape stays unresolvable here: the bytes live behind the Files API and
this extractor has no client to fetch them. Documented rather than silently
dropped, so a consumer treating a missing entry as "no image to scan" is a known
gap and not a surprise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two gaps the create form and the edit modal share today.
The submit gate never asked for a classifier model. Choosing the LLM classifier
and no model leaves Test Routing and Add Auto Router enabled, so Test Routing
posts a config the backend rejects and only the later save says why.
The keyword-rule gate only looked for empty keyword rows. A rule's tier has been
a free string since #37413, and the backend matches it exactly, so a rule naming
a tier the router does not have cleared the gate and failed the save as a raw
400.
Both gates now live in build_complexity_router_config.ts, and each form's submit
handler reads the same blocked reason the button reads instead of re-deriving
its own list, so a disabled button and a refused submit cannot disagree.
* fix(logging): stop billing and logging response reads as LLM calls
Retrieving, deleting or cancelling a stored response, and vector store management calls, run through the same logging lifecycle as inference. A retrieved response replays the usage of the call that created it, so every read priced it again and wrote a second spend log row for the same tokens. Non-inference calls now cost 0, report no usage, log no placeholder chat message, and get a litellm.responses_management operation name instead of reading as chat.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep billing background response jobs after the poll
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): use an empty list for read-call messages
A tuple matches no branch in the loggers that walk this value, so lunary's
parse_messages falls through to clean_message and raises AttributeError on the
success hook. An empty list reads as no messages everywhere: it satisfies the
isinstance(list) checks in newrelic, mlflow and datadog, iterates zero times in
traceloop and helicone, and is what StandardLoggingPayload.messages is typed to
hold. None would be type-legal too but is not iterable, so it trades one crash
for another in mlflow and traceloop.
* fix(otel): stop the legacy emitter reporting replayed tokens on response reads
The zeroing so far lands in the standard logging payload, which the legacy
OpenTelemetry emitter does not read for usage: it takes prompt, completion and
total tokens straight off the response object, so a retrieval span still carried
the token counts of the call that produced the response, and the token usage
histogram still recorded them. That emitter is the default, so the spend row said
zero while the trace said otherwise. The background cost poller keeps its counts,
the same exemption the pricing path already makes.
* fix(logging): keep billing a background response when its retrieval is read
A response created with background=true comes back queued and carries no usage, so
its create bills nothing. The retrieval that first sees the finished job is the only
place that job's tokens are ever visible, and pricing every read at zero therefore
loses the spend outright rather than deduplicating it. On a proxy without the
enterprise cost poller a background job ended up costing $0 end to end.
is_unbilled_non_inference_call now takes the response it is deciding about and treats
a background response the same way it already treats the poller's own read, which is
the same exemption seen from the other side. The legacy OpenTelemetry emitter's time
per output token metric picks up the read gate it was missing, so it stops dividing a
read's latency by the replayed completion token count.
* test(proxy): pass the read response to the non-inference predicate
The poller test called is_unbilled_non_inference_call with the pre-background signature, so it broke when the predicate gained the response it classifies. It now hands the predicate a foreground read, and asserts that the same read is free without the origin stamp, so the stamp is what the test proves.
* fix(otel): stop the v2 metrics recorder reporting replayed tokens on response reads
The v2 span builder sources usage from the standard logging payload, so the
earlier fix already zeroes it there. The metrics recorder reads response_obj
directly, so a responses-management read still recorded the original
generation's tokens into gen_ai.client.token.usage and divided generation time
by them for gen_ai.server.time_per_output_token.
The read still records operation and response duration, under the
litellm.responses_management operation, so it stays observable.
* fix(proxy): keep the response-cost headers on calls priced at zero
Pricing responses reads and vector-store management routes at zero dropped the whole
x-litellm-response-cost family off those replies. The header build reads a falsy zero as
a cost this response never recorded and filters it out, and a call that returns before
pricing stores no cost breakdown for the component headers to read, so a client parsing
the cost off a read got a KeyError where it had previously been handed a number.
Those calls now advertise the family at zero. Retrieving a background response, and the
cost poller's read of one, still report their real cost.
The params-taking form of the predicate moves from opentelemetry into
internal_call_metadata so the proxy header build and the OTEL recorders share one copy.
* fix(proxy): report a zero cost split only under a zero cost total
The component headers were filled from call-type membership alone, while the
total they sit beside keeps its real value when the read priced normally, so a
breakdown that had not landed by the time headers were built could advertise a
real total next to an all-zero split. The split is now reported as zero only
when the total agrees with it, and is otherwise left absent.
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
For non-Anthropic models served over /v1/messages, the outer wrapper recomputes
cost over the adapter-translated Anthropic response dict. That dict dropped every
web search usage signal, so the recompute overwrote the correct cost breakdown
with a token-only one: x-litellm-response-cost-tool-usage read 0.0 and
x-litellm-response-cost-original excluded the search cost, while the total kept it.
The adapter now maps web search request counts (from Usage.server_tool_use or
Gemini's prompt_tokens_details) into usage.server_tool_use.web_search_requests,
matching the Anthropic API shape, and the Gemini web search cost calculator falls
back to server_tool_use when prompt_tokens_details carries no count. The shared
get_web_search_requests helper is now public since five modules consume it.
Resolves LIT-6288
The case was skipped because /budget/update 500d on any model_max_budget.
#38430 fixes that by serializing the update payload before the write, so
the case now passes against a proxy carrying that change and there is
nothing left for the skip to hide.
Merge this after #38430; on staging alone the case still fails with the
same 500 it was skipped for.
Reverts #37725. The field existed so SDK callers that cannot read
`x-litellm-model-id` could tell which tier an auto-router picked, and the
framework that motivated it was LangChain. `@langchain/openai` builds
`additional_kwargs` and `response_metadata` from fixed key allowlists and drops
unknown fields at both the chunk top level and inside `delta`, so no
proxy-side placement of a namespaced key can reach a LangChain caller.
The complexity router's existing `return_raw_model_name` already covers that
case: it puts the resolved model in the standard `model` field, which
LangChain does propagate (`model_name` is on its metadata allowlist), and the
proxy honors it on both the streaming and non-streaming paths.
Keeps the unrelated cleanup from #37725 that dropped the redundant
function-local `ProxyBaseLLMRequestProcessing` import shadowing the
module-level one in `async_data_generator`.
`TestModelGroupAliasReachesPreRoutingStrategies` asserted on the marker as a
proof of strategy dispatch; the surviving `response.model == "gemini-flash"`
assertion already proves it.