The Live passthrough builds Usage from the TEXT-modality counts alone, so audio,
image and video tokens never reach the cost calculator and bill as nothing. A
one-turn audio session reported 13 text and 127 audio input tokens and billed the
13; a camera session reported 1043 prompt tokens and billed 11.
Reporting the full per-modality breakdown fixes it, because the shared Gemini
input and output cost path already prices audio, image and video from
prompt_tokens_details and completion_tokens_details. On the native-audio entry
that is a 6x difference per token in both directions, which is the whole gap.
Aggregation across turns is unchanged. Google charges per turn for every token in
the Live session context window, current turn plus all accumulated tokens from
previous turns, so the existing summing is what Vertex bills and it stays as it
is. That is worth stating because the cumulative promptTokensDetails looks like a
restatement of one running total, and treating it that way would under-bill a
multi-turn session. See the LiveAPI context-window note on
https://cloud.google.com/vertex-ai/generative-ai/pricing.
Live can also name the modality carrying the rest of a turn and omit its
tokenCount. Reading that absent key as zero left the tokens inside
candidatesTokenCount but outside the breakdown, so real speech was charged at the
text output rate. A lone unpriced entry now takes whatever the turn's declared
count leaves over. Two or more cannot be told apart, so they are still left to the
calculator's text remainder.
Server-side toolUsePromptTokenCount is now reported in prompt_tokens_details. It
is deliberately kept out of prompt_tokens: no Gemini route prices tool-use tokens,
and adding them there instead suppresses the cache-overlap correction and raises
the bill for no extra work.
Removes _calculate_live_api_cost, whose result never reached the bill. It set
kwargs["response_cost"], which the standard logging path recomputes from the
ModelResponse, and on a measured audio session it returned $0.000487 against a
$0.0000425 row. Now that the modality counts reach the standard calculator,
keeping a second hand-rolled pricing path would only ever double-charge.
The rewrite of the aggregator is arithmetically identical to what it replaced. It
sums the same three counts and the same per-modality details, still takes the
remaining fields from the first turn, and drops nine LIT010, one C901 and 42
basedpyright findings in the process.
Reasoning items the Responses API returned for a /v1/messages turn were rebuilt
from their summary text on every replay, so the prompt the model saw changed
between turns and the prompt cache never matched. The bridge now asks for
reasoning.encrypted_content, carries it in the thinking signature (or as a
redacted_thinking block when there is no summary), and replays it verbatim as
the reasoning item's encrypted_content. Anthropic replay paths drop those
tagged blocks so a cross-model resume never forwards OpenAI bytes to Anthropic
Replace the thread semaphore around HuggingFace encodes with an anyio CapacityLimiter
applied at every offloaded count site through offload_token_count, so waiting counts no
longer hold slots in the shared 40-thread pool and inline counts on the event loop never
block on the bound. Rename TOKEN_COUNTER_MAX_CONCURRENT_HF_ENCODES to
TOKEN_COUNTER_MAX_CONCURRENT_COUNTS
get_optional_params_image_gen only forwarded the per-call drop_params flag
to provider configs, so litellm_settings drop_params: true never dropped
the n the MAI generations endpoint ignores. The MAI edits config also
advertised and forwarded size, which that endpoint ignores. Non-numeric
and non-positive n now surface as a 400 instead of a 500 or a pass-through.
Every test this PR adds now annotates its fixture and parametrize
parameters. The submit-time warning for a tag-matched deferred policy
names only the policies, since a wildcard attachment pattern would let
caller-provided tag text reach the log.
With enable_azure_ad_token_refresh, every keyless image request built a new DefaultAzureCredential and fetched a token. Cache the provider per scope like the Entra ID one.
A guardrail that removes or adds a tool call on an ended stream used to be
silently ignored: every handler substitutes the original list on a count
mismatch and the executor skipped its observer once the translation could
deliver rewrites. The executor now tracks the count change on the observer
and releases the original chunks with the discard warning on every
translation, matching what the merge base did for any tool call rewrite
* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods
* fix(mcp): preserve proxy logging and authorization coverage
* test(mcp): respect the proxy FastAPI import boundary
Extract one immutable helper for reading the deployment's model_info off the
logging object, stop rebinding the model_info parameter inside ocr_cost, drop
the explanatory comment blocks, and move the OCR custom pricing regression
tests into tests/test_litellm/test_cost_calculator.py
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.
Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.
Fixes#39145
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge
Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): drop docstrings from the two regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun proxy-infra after the flaky test_check_migration process-tree test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>