A realtime session's usage is rebuilt from its own response.done event, so a counter
that does not survive the round trip is invisible to the cost path. Both directions
copied a fixed allow-list, which meant a grounded Gemini Live session reported its
query on the Usage object and then lost it before anything could bill it.
Gemini reads the grounding counters off the input token details while Anthropic reads
its own server_tool_use field, so carrying these two cannot move an Anthropic bill.
Absent counters stay absent, so no provider starts paying a fee it did not incur.
(cherry picked from commit ffd6c723e2a65dfff1860a913c34e41bc72ff11f)
(cherry picked from commit 1590822f6893c57086b72965963d23f4f367b424)
A client that named a bare gateway alias logged the session as "unknown" and billed nothing,
because the model was read off the raw setup frame and the extractor only yields a name when the
string already contains "/models/". The rewriter qualifies that same model a few lines later for
the upstream, so the supported client form, an alias, was the one that went unbilled.
Resolving through the rewriter first means the real model reaches the logging object, and from
there the cost map. A route with no rewriter, which is every non-Live passthrough, hands the frame
over untouched.
(cherry picked from commit 573982803df612fd94144e2e06dd647f8530d4e8)
Live reports grounding in the server frames and never in usageMetadata, so nothing
set the counter the cost path reads and the per-query charge was missing from every
grounded session. Google bills a grounded Live prompt on top of its tokens, and that
fee dwarfs the token cost, so a non-zero spend check could never catch it.
Both Live surfaces now read serverContent.groundingMetadata where they build usage,
and reuse the chat path's own classifier so web search and Maps keep their separate
SKUs rather than being counted together.
Separately, a client sending turn_detection: null reached a membership test against
None and took the session down with no traceback, while the branch immediately above
already guards for it. Live emits grounding and usage on the same frame, verified
against Vertex directly, so the realtime counter is set where usage is built.
(cherry picked from commit c997436be34beb2e84f8468286b8016c195eca92)
(cherry picked from commit 26c8d4822fc1c5c44fe8f72f4f57473a2ce1acbf)
The four session helpers this PR added were unannotated. Typing them needs a
name for the (text, audio) pair each turn carries, so _LiveTurn is a TypedDict
rather than a Mapping union that would leave sum() over a prompt pair
ill-typed, and AUDIO_SESSION is declared with it. The message list reuses
list[dict[str, object]], the annotation the passthrough already uses where it
collects those messages
toolUsePromptTokenCount was the one prompt-side total not named in
_AGGREGATED_FIELDS, so it rode the unknown-key pass-through and took the first
frame's value while promptTokenCount, candidatesTokenCount and totalTokenCount
beside it were summed. Live's frames grow over a session, so the first frame is
the smallest number in the series and a grounded session under-reported its
tool-use tokens by everything after turn one. It is now summed like its three
neighbours.
This is reporting only, and pricing these tokens is deliberately left out. Google
charges tool-use prompt tokens at the input token rate, but generic_cost_per_token
reads the input bill out of prompt_tokens_details and only falls back to
prompt_tokens when the details carry no text or a cache hit overlaps them.
Measured on the native-audio entry with 500 tool-use tokens: adding them to
prompt_tokens moves an ordinary Live turn's bill by $0.0000000000, and on a turn
with a cache hit it moves it by $0.0002650000 where the tokens are worth
$0.0002500000, because it perturbs the cache-overlap correction. Pricing them
belongs beside the modality terms in the shared input-cost path, in its own
change that fixes the same latent no-op on the ordinary Gemini path.
Not verified against a live capture: no Vertex Live session we have captured
reported toolUsePromptTokenCount at all, so the summing convention is inferred
from the three prompt-side totals that accumulate the same way.
The two helpers this branch adds took Sequence[Mapping[str, Any]], which the
repo forbids, and only typechecked because Any is compatible with everything.
Both now take Mapping[str, object] and the raw *TokensDetails value is narrowed
to its mapping entries at each of the three call sites.
TypedDicts are the wrong tool here: _merged_modality_totals reads count_key and
details_key as runtime strings, and the aggregation deliberately passes unknown
keys straight through, so both need a mapping whose keys are not literals.
The narrowing is not cosmetic. The handler's only failure path returns no result
at all, so a *TokensDetails value that was not a list of objects used to raise
while being read and cost the whole session its bill.
The Live passthrough builds Usage from the TEXT-modality counts alone, so audio,
image and video tokens never reach the cost calculator and bill as nothing. A
one-turn audio session reported 13 text and 127 audio input tokens and billed the
13; a camera session reported 1043 prompt tokens and billed 11.
Reporting the full per-modality breakdown fixes it, because the shared Gemini
input and output cost path already prices audio, image and video from
prompt_tokens_details and completion_tokens_details. On the native-audio entry
that is a 6x difference per token in both directions, which is the whole gap.
Aggregation across turns is unchanged. Google charges per turn for every token in
the Live session context window, current turn plus all accumulated tokens from
previous turns, so the existing summing is what Vertex bills and it stays as it
is. That is worth stating because the cumulative promptTokensDetails looks like a
restatement of one running total, and treating it that way would under-bill a
multi-turn session. See the LiveAPI context-window note on
https://cloud.google.com/vertex-ai/generative-ai/pricing.
Live can also name the modality carrying the rest of a turn and omit its
tokenCount. Reading that absent key as zero left the tokens inside
candidatesTokenCount but outside the breakdown, so real speech was charged at the
text output rate. A lone unpriced entry now takes whatever the turn's declared
count leaves over. Two or more cannot be told apart, so they are still left to the
calculator's text remainder.
Server-side toolUsePromptTokenCount is now reported in prompt_tokens_details. It
is deliberately kept out of prompt_tokens: no Gemini route prices tool-use tokens,
and adding them there instead suppresses the cache-overlap correction and raises
the bill for no extra work.
Removes _calculate_live_api_cost, whose result never reached the bill. It set
kwargs["response_cost"], which the standard logging path recomputes from the
ModelResponse, and on a measured audio session it returned $0.000487 against a
$0.0000425 row. Now that the modality counts reach the standard calculator,
keeping a second hand-rolled pricing path would only ever double-charge.
The rewrite of the aggregator is arithmetically identical to what it replaced. It
sums the same three counts and the same per-modality details, still takes the
remaining fields from the first turn, and drops nine LIT010, one C901 and 42
basedpyright findings in the process.
Reasoning items the Responses API returned for a /v1/messages turn were rebuilt
from their summary text on every replay, so the prompt the model saw changed
between turns and the prompt cache never matched. The bridge now asks for
reasoning.encrypted_content, carries it in the thinking signature (or as a
redacted_thinking block when there is no summary), and replays it verbatim as
the reasoning item's encrypted_content. Anthropic replay paths drop those
tagged blocks so a cross-model resume never forwards OpenAI bytes to Anthropic
Replace the thread semaphore around HuggingFace encodes with an anyio CapacityLimiter
applied at every offloaded count site through offload_token_count, so waiting counts no
longer hold slots in the shared 40-thread pool and inline counts on the event loop never
block on the bound. Rename TOKEN_COUNTER_MAX_CONCURRENT_HF_ENCODES to
TOKEN_COUNTER_MAX_CONCURRENT_COUNTS
get_optional_params_image_gen only forwarded the per-call drop_params flag
to provider configs, so litellm_settings drop_params: true never dropped
the n the MAI generations endpoint ignores. The MAI edits config also
advertised and forwarded size, which that endpoint ignores. Non-numeric
and non-positive n now surface as a 400 instead of a 500 or a pass-through.
Every test this PR adds now annotates its fixture and parametrize
parameters. The submit-time warning for a tag-matched deferred policy
names only the policies, since a wildcard attachment pattern would let
caller-provided tag text reach the log.
A guardrail that removes or adds a tool call on an ended stream used to be
silently ignored: every handler substitutes the original list on a count
mismatch and the executor skipped its observer once the translation could
deliver rewrites. The executor now tracks the count change on the observer
and releases the original chunks with the discard warning on every
translation, matching what the merge base did for any tool call rewrite
* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods
* fix(mcp): preserve proxy logging and authorization coverage
* test(mcp): respect the proxy FastAPI import boundary
Extract one immutable helper for reading the deployment's model_info off the
logging object, stop rebinding the model_info parameter inside ocr_cost, drop
the explanatory comment blocks, and move the OCR custom pricing regression
tests into tests/test_litellm/test_cost_calculator.py
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.
Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.
Fixes#39145
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge
Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): drop docstrings from the two regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun proxy-infra after the flaky test_check_migration process-tree test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>