The call site repeated _resolved_vertex_live_setup's own docstring almost verbatim, which is the
duplication the repo's comment rule exists to prevent.
A realtime session's usage is rebuilt from its own response.done event, so a counter
that does not survive the round trip is invisible to the cost path. Both directions
copied a fixed allow-list, which meant a grounded Gemini Live session reported its
query on the Usage object and then lost it before anything could bill it.
Gemini reads the grounding counters off the input token details while Anthropic reads
its own server_tool_use field, so carrying these two cannot move an Anthropic bill.
Absent counters stay absent, so no provider starts paying a fee it did not incur.
(cherry picked from commit ffd6c723e2a65dfff1860a913c34e41bc72ff11f)
(cherry picked from commit 1590822f6893c57086b72965963d23f4f367b424)
A client that named a bare gateway alias logged the session as "unknown" and billed nothing,
because the model was read off the raw setup frame and the extractor only yields a name when the
string already contains "/models/". The rewriter qualifies that same model a few lines later for
the upstream, so the supported client form, an alias, was the one that went unbilled.
Resolving through the rewriter first means the real model reaches the logging object, and from
there the cost map. A route with no rewriter, which is every non-Live passthrough, hands the frame
over untouched.
(cherry picked from commit 573982803df612fd94144e2e06dd647f8530d4e8)
Live reports grounding in the server frames and never in usageMetadata, so nothing
set the counter the cost path reads and the per-query charge was missing from every
grounded session. Google bills a grounded Live prompt on top of its tokens, and that
fee dwarfs the token cost, so a non-zero spend check could never catch it.
Both Live surfaces now read serverContent.groundingMetadata where they build usage,
and reuse the chat path's own classifier so web search and Maps keep their separate
SKUs rather than being counted together.
Separately, a client sending turn_detection: null reached a membership test against
None and took the session down with no traceback, while the branch immediately above
already guards for it. Live emits grounding and usage on the same frame, verified
against Vertex directly, so the realtime counter is set where usage is built.
(cherry picked from commit c997436be34beb2e84f8468286b8016c195eca92)
(cherry picked from commit 26c8d4822fc1c5c44fe8f72f4f57473a2ce1acbf)
The four session helpers this PR added were unannotated. Typing them needs a
name for the (text, audio) pair each turn carries, so _LiveTurn is a TypedDict
rather than a Mapping union that would leave sum() over a prompt pair
ill-typed, and AUDIO_SESSION is declared with it. The message list reuses
list[dict[str, object]], the annotation the passthrough already uses where it
collects those messages
toolUsePromptTokenCount was the one prompt-side total not named in
_AGGREGATED_FIELDS, so it rode the unknown-key pass-through and took the first
frame's value while promptTokenCount, candidatesTokenCount and totalTokenCount
beside it were summed. Live's frames grow over a session, so the first frame is
the smallest number in the series and a grounded session under-reported its
tool-use tokens by everything after turn one. It is now summed like its three
neighbours.
This is reporting only, and pricing these tokens is deliberately left out. Google
charges tool-use prompt tokens at the input token rate, but generic_cost_per_token
reads the input bill out of prompt_tokens_details and only falls back to
prompt_tokens when the details carry no text or a cache hit overlaps them.
Measured on the native-audio entry with 500 tool-use tokens: adding them to
prompt_tokens moves an ordinary Live turn's bill by $0.0000000000, and on a turn
with a cache hit it moves it by $0.0002650000 where the tokens are worth
$0.0002500000, because it perturbs the cache-overlap correction. Pricing them
belongs beside the modality terms in the shared input-cost path, in its own
change that fixes the same latent no-op on the ordinary Gemini path.
Not verified against a live capture: no Vertex Live session we have captured
reported toolUsePromptTokenCount at all, so the summing convention is inferred
from the three prompt-side totals that accumulate the same way.
The two helpers this branch adds took Sequence[Mapping[str, Any]], which the
repo forbids, and only typechecked because Any is compatible with everything.
Both now take Mapping[str, object] and the raw *TokensDetails value is narrowed
to its mapping entries at each of the three call sites.
TypedDicts are the wrong tool here: _merged_modality_totals reads count_key and
details_key as runtime strings, and the aggregation deliberately passes unknown
keys straight through, so both need a mapping whose keys are not literals.
The narrowing is not cosmetic. The handler's only failure path returns no result
at all, so a *TokensDetails value that was not a list of objects used to raise
while being read and cost the whole session its bill.
The Live passthrough builds Usage from the TEXT-modality counts alone, so audio,
image and video tokens never reach the cost calculator and bill as nothing. A
one-turn audio session reported 13 text and 127 audio input tokens and billed the
13; a camera session reported 1043 prompt tokens and billed 11.
Reporting the full per-modality breakdown fixes it, because the shared Gemini
input and output cost path already prices audio, image and video from
prompt_tokens_details and completion_tokens_details. On the native-audio entry
that is a 6x difference per token in both directions, which is the whole gap.
Aggregation across turns is unchanged. Google charges per turn for every token in
the Live session context window, current turn plus all accumulated tokens from
previous turns, so the existing summing is what Vertex bills and it stays as it
is. That is worth stating because the cumulative promptTokensDetails looks like a
restatement of one running total, and treating it that way would under-bill a
multi-turn session. See the LiveAPI context-window note on
https://cloud.google.com/vertex-ai/generative-ai/pricing.
Live can also name the modality carrying the rest of a turn and omit its
tokenCount. Reading that absent key as zero left the tokens inside
candidatesTokenCount but outside the breakdown, so real speech was charged at the
text output rate. A lone unpriced entry now takes whatever the turn's declared
count leaves over. Two or more cannot be told apart, so they are still left to the
calculator's text remainder.
Server-side toolUsePromptTokenCount is now reported in prompt_tokens_details. It
is deliberately kept out of prompt_tokens: no Gemini route prices tool-use tokens,
and adding them there instead suppresses the cache-overlap correction and raises
the bill for no extra work.
Removes _calculate_live_api_cost, whose result never reached the bill. It set
kwargs["response_cost"], which the standard logging path recomputes from the
ModelResponse, and on a measured audio session it returned $0.000487 against a
$0.0000425 row. Now that the modality counts reach the standard calculator,
keeping a second hand-rolled pricing path would only ever double-charge.
The rewrite of the aggregator is arithmetically identical to what it replaced. It
sums the same three counts and the same per-modality details, still takes the
remaining fields from the first turn, and drops nine LIT010, one C901 and 42
basedpyright findings in the process.
* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods
* fix(mcp): preserve proxy logging and authorization coverage
* test(mcp): respect the proxy FastAPI import boundary
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.
Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.
Fixes#39145
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge
Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): drop docstrings from the two regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun proxy-infra after the flaky test_check_migration process-tree test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_sentinel): pin batch_size as a per-request bound under concurrent events
Adds a regression test to the mapped Azure Sentinel test file for the concurrency scenario from LIT-6920: 40 records logged concurrently at batch_size=5 while each ingestion request is still in flight. Asserts no request carries more than batch_size records, every record arrives exactly once in order, and the queue is empty afterwards. Runs for both the standard log queue and the audit log queue.
The test fails on the tree before #39880 (whole shared queue serialized per threshold send, then cleared) and passes on current staging. It is independent of the size-split coverage that #39880 added for LIT-5899.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_sentinel): gate the first send on events so later records provably arrive while it is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Turning the tier off, or switching to a classifier that cannot emit it, dropped
the flag and the pool but left plan_mode_min_tier naming a tier that is no longer
active. The backend rejects that on save, and the switch is disabled after a
classifier change, so the operator had no way to clear it.
Both paths now release the floor when it points at the cleared tier. An orphaned
keyword rule is left alone on purpose: getKeywordTierRulesError already names it
at the save gate, which is how a removed custom tier behaves.
* feat(otel v2): send a key's or team's whole trace to its own destination
A key or team that configures its own Langfuse, Arize, Weave or New Relic
credentials used to get a single detached span in its account while the rest of
the request trace stayed on the operator's backend, so neither side held a
complete trace. Resolve the destination during auth, forward every span of the
request to it, and hold the same request back from the operator's exporter for
that backend, so the tenant gets the tree the operator would have seen and the
operator gets nothing for that request.
Also let a credential-mandatory preset build without the operator's own env
credentials. Without that, a proxy whose teams each bring their own account fell
back to the legacy integration and never ran a line of the v2 path.
* fix(otel v2): validate tenant destinations and match each backend's own endpoint
Review round on the tenant destination routing.
- A key/team Langfuse host is user-supplied input, so it goes through the
proxy's SSRF guard. A private address is refused, the operator keeps the
trace, and the warning names user_url_allowed_hosts. The operator's own
LANGFUSE_HOST is not checked.
- Arize and Weave destinations now resolve their endpoint and transport
through the backend's own config, so an ARIZE_HTTP_ENDPOINT collector and
a self-hosted WANDB_HOST are honoured instead of the cloud default.
- A half-configured backend no longer resolves: several dynamic header
builders gate each credential separately, so an api key with no space id
produced a non-empty but unusable header set that suppressed the
operator's exporter.
- A callback_type of "failure" no longer takes over the trace. The
destination is resolved during auth, before the outcome is known.
- The fan-out cache evicts without shutting the processor down, matching
ArizePhoenixLogger: a concurrent on_end may still hold it.
- The stdout placeholder is identified by what it does rather than by
equality with an import-time default, so an operator's OTEL_EXPORTER_OTLP_*
collector survives the credential-less path.
* refactor(otel v2): reuse the proxy's own destination allowlist for tenant hosts
A tenant-supplied Langfuse host is the same threat as a URL-valued `model`, so
it now goes through `is_url_destination_allowed_by_host` against
`provider_url_destination_allowed_hosts` instead of a second, DNS-based check
of its own. The DNS lookup would have blocked the asyncio auth path on a
hostname the caller picked, and its cached verdicts could blackhole a real host
after one resolver blip.
Evicting a destination processor now retires it to drain rather than shutting it
down, since `on_end` hands a processor back and exports outside the lock. The
retirees are capped so they cannot accumulate a thread each.
`credential_gated_exporters` tells the synthesized stdout placeholder from a
real exporter by transport rather than by the literal kind `console`, so an
unrecognized kind is not mistaken for a configured collector, and an exporter
the operator did configure survives. That also stops a weave test's env writes
from making this look like a real OTLP exporter later in the same CI worker.
* fix(otel v2): read the tenant's stored callback config the way the sibling parser does
Three divergences between the destination resolver and
`convert_key_logging_metadata_to_callback`, which read the same stored config:
- A key whose callbacks are disabled stores an empty list, and `or` treated that
as "the key configured nothing", so the request inherited the team's
destination. The sibling parser treats an empty list as configured.
- Two entries naming one backend now merge their `callback_vars` last-wins,
matching the sibling, instead of the resolver taking the first entry and the
per-request tracer routing taking the last.
- `credential_gated_exporters` dropped any exporter whose kind had no transport,
which also dropped an `in_memory` exporter the operator asked for. The
placeholder is the spec with every field still at its default, so that is what
the predicate now says.
Arize's `allow_missing_credentials` branch was unreachable: `get_arize_config`
resolves every credential with `os.environ.get` and always supplies an endpoint,
so it never raises. Dropped it and corrected the protocol docstring.
* fix(otel v2): keep the destination merge immutable
The per-backend var merge seeded a plain dict and the gated exporter list a plain
list, both of which the LIT budget counts. Wrap the merge in MappingProxyType and
hand the exporters back as a tuple.
* fix(otel v2): scope the fan-out to its own backend and close shed processors off the export path
Three problems in the fan-out, two of them in the eviction added last round:
- Every v2 logger carries its own provider and emits its own copy of a gen-AI
span, so a proxy running two of them handed the tenant the same model call
twice. A provider now forwards only destinations for the backend it speaks
for; the tenant's own backend always has a logger, since naming it in the key
or team config is what builds one. Reproduced live against a self-hosted
Langfuse on an arize-only proxy and on the bare `otel` callback.
- Eviction could close a processor another thread was still exporting through,
which drops that span. Exports are now counted, and a retired processor is
closed only once its count reaches zero.
- That close ran inside `on_end`, where `shutdown` flushes over the network, so
one unreachable tenant collector stalled every other tenant's spans. It now
runs on a short-lived thread, which also retires the retiree cap: a retiree
drains as soon as its export finishes.
* fix(otel v2): deliver tenant destinations from the published global provider
Scoping the fan-out by callback name in the previous commit left every backend
that is not the canonical logger with a one-span trace: only the published
global provider sees the FastAPI server span, the auth span and the post-call
database spans, so an arize-only proxy handed a team's Langfuse just the model
call. Attach the fan-out once, to that provider, and let it forward every
destination.
An overridden backend now skips per-request tracer routing outright rather than
only clearing its credential headers, since a key or team otel_service_name was
still enough to detach the model call onto a second provider. The destination
carries that service name as a resource attribute instead.
Shed processors drain on a two-thread pool rather than a thread each, so a
tenant cycling its destination config cannot spawn threads as fast as it sends
requests.
* fix(otel v2): drain shed destination processors on daemon workers
A ThreadPoolExecutor joins its workers at interpreter exit, so one unreachable
tenant collector would hold the whole proxy open for its export timeout on the
way down. Two long-lived daemon workers off a queue keep the thread count
bounded without blocking shutdown.
* fix(otel v2): give the fan-out its own drain pool instead of a lazy singleton
functools.lru_cache does not hold a lock across the call it caches, so
concurrent first evictions each finish building a queue and start its
workers, and every queue but the winner is abandoned with two daemon
threads blocked on it forever.
* fix(otel v2): close no destination processor under a span still in flight
The fan-out now refuses new work once shutdown starts and waits out the
spans already being forwarded, so teardown neither drops a trace mid-forward
nor hands the next caller an exporter nothing will ever close. The wait is
bounded so a dead collector cannot hold the proxy open.
* fix(otel v2): retire the drain workers with the fan-out that started them
A proxy that rebuilds its telemetry builds another fan-out, so workers that
outlive the one that started them are two more threads per reload. Shutdown
now retires them once everything queued is closed, and a processor shed
afterwards is closed inline rather than queued to nobody.
* fix(otel v2): guard the fan-out's closed state with the lock that gates it
An Event read on its own leaves room for shutdown to run in the gap. A cache
miss then inserted a live exporter into a map that had been cleared, and a
shed processor landed behind sentinels every drain worker had exited on.
The drain pool takes its queue by injection so both interleavings are
reachable from a test without patching.
* feat(otel v2): let a tenant destination export alongside the operator's own
Override stays the default: a key or team destination replaces the operator's
exporter for that backend. Operators running one org-wide backend across every
team set litellm_settings.otel_tenant_destination_mode to additive, and the
same trace lands in both places. A team that names the operator's own project
is still written once, since the fan-out skips a destination the operator's
exporter is already sending that span to.
* fix(otel v2): let a straggling export close its own destination processor
Shutdown waits out the exports in flight, but the wait has to be bounded or a
tenant collector that stops answering holds the proxy open on the way down.
Past the bound it closed everything anyway, which is the case it was written to
avoid: a processor closed under the span it is carrying loses that span.
Keep the bound and retire the stragglers instead. The thread still exporting one
closes it through the drain as soon as its export returns, so teardown stays
bounded and no span is dropped mid-forward.
* fix(otel v2): identify a destination account by its credentials, not its header names
Under additive the fan-out skips a destination the operator's own exporter
already writes to, so the same account is not written twice. It compared header
names as well as values, and one account answers to more than one spelling:
the operator's Arize exporter sends space_id where a team destination sends
arize-space-id, so every span landed in the operator's own space twice.
The credentials are the identity. Compare those and leave the spelling to each
backend.
* fix(otel v2): keep the credential's role in a destination's account identity
Comparing values alone folds two accounts together whenever they hold the same
strings in different roles, and the second team would then get no trace at all.
Compare the credential under a normalized name instead, and fold the one alias
that actually exists: Arize's space_id and arize-space-id.
* fix(otel v2): build one destination processor per destination, not per racing span
Building outside the cache lock meant a cold cache met by a burst of concurrent
requests constructed an exporter per thread, kept one, and handed the rest to the
drain, so a batch worker and a connection pool per losing thread sat in a queue
two workers service.
Build under the lock that reads the cache. Opening an exporter connects to
nothing, so the lock is held for a constructor, once per destination, and the
race it was avoiding stops existing.
* fix(otel v2): bound the teardown that closes a destination, not the one that never blocks
The five-second bound guarded the wait for spans still inside on_end, but a
batching processor's on_end only queues the span and returns, so that counter is
empty and the bound engaged against nothing. The blocking half was the serial
close, which flushes over the network and joins the SDK's own worker thread with
no timeout of its own, so a single tenant collector that answers and never
finishes held process teardown open for as long as it liked.
Hand every close to the drain, whose workers are daemons, and give the whole
teardown one deadline.
* fix(otel v2): preserve operator spans on destination failure
* fix(otel v2): anchor destinations off the published provider, refuse headerless tenant transports
set_tracer_provider keeps the first provider it is handed, so a process whose
OTel global was claimed before the proxy published (auto-instrumentation, a
legacy logger) had no fan-out on the global and auth anchored no destination.
Auth now reads the fan-out off the registered logger's own provider.
A destination whose protocol maps to a headerless exporter kind is no longer
buildable: the console fallback would drop the tenant's credentials and print
the spans to stdout while the operator's exporter stood down for them.
* fix(otel): anchor tenant fan-out to the published provider
A legacy v1 logger can occupy proxy_server.open_telemetry_logger, in which case
the proxy publishes with registered=None and the fan-out lands on a v2 logger
taken from _in_memory_loggers. Reading the registered slot found no v2 logger
and the OTel global belonged to v1, so auth refused every tenant destination.
* fix(otel): preserve registered provider fallback
* test(otel): cover pre-publish provider fallback
* fix(otel): attach fan-out on fallback provider
* fix(otel): serialize first fan-out attach
* fix(otel): keep the operator's database endpoint out of tenant traces
A database span forwarded to a key or team destination carried the proxy's own
Postgres host, port and schema, and on failure the Prisma error text naming them.
The fan-out now hands tenants a view of each database span without those keys,
its events or its status text, while the operator's own copy is untouched and
model endpoints such as server.address on the LLM span still travel
* fix(otel): keep relabelled spans in the fan-out and honour disabled callbacks for destinations
A key or team otel_service_name used to move a backend's span onto a second
provider even when another backend had a destination, so the fan-out never saw
the model call and the tenant's trace lost it. A service name alone now stays on
the published provider whenever the request has a destination; credential and
project routing to a tenant's own account is unchanged
Destinations now skip a backend the request disabled dynamically, reading the
x-litellm-disable-callbacks header and the key's litellm_disabled_callbacks with
the same precedence and premium gate dispatch applies, so a disabled backend is
neither delivered to nor withheld from the operator
* test(otel): project routing survives a sibling backend destination
* docs(otel): state why a disabled backend still routes its own span
* fix(otel): keep a degraded backend's spans off a collector another v2 logger already serves
* test(otel): a credentialed preset beside another v2 logger keeps every exporter
* fix(otel): keep credentialless fallback on base path
* test(otel): cover legacy callback carrier rejection
* fix(otel): preserve valid exporter beside gated preset
* fix(otel): avoid console export without operator destination
* refactor(otel): share the console placeholder check with the presets
* fix(otel): bound shed destination processors waiting on a dead collector
* fix(otel): preserve explicit console exporters
* fix(otel): avoid mutable field-set construction
* fix(otel): close drain saturation race
* fix(otel): drop captured request headers from tenant spans
* test(otel v2): give the newrelic dispatch tests operator credentials, since a credential-less preset now falls back
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): rebuild an anchored destination's processor past drain saturation
A destination deliverable() accepted at auth can be evicted by other tenants' auths
before its request's spans end, and that eviction is what tips the drain over. The
saturation gate then refused the rebuild at on_end, and with the operator's exporter
already stood down for that backend the span went nowhere. The gate now applies only
while a request decides whether to anchor
* fix(otel): hold destination eviction while the drain is saturated
An anchored destination evicted by other tenants' auths is rebuilt on its
next span, and that rebuild evicted another anchored one, so with more
destinations in flight than the cache holds every span cost one more
processor, one more batch thread and one more close queued behind a collector
that never answers. Eviction now holds while the drain is saturated, so the
cache keeps one entry per destination in flight and trims back to its cap on
the next hit or build once the drain has room
* fix(otel): keep the proxy's own error text out of tenant traces
A tenant destination received every span the request produced, error text
included, so a Prisma failure during auth handed a team admin's collector the
operator's Postgres endpoint, and the exception event on any failed span carried
a stack trace naming the proxy's install paths.
Spans the tenant's own call produced (the model call, MCP, guardrails) keep their
error text. Every other span keeps the failure without the prose: its type, its
provider error code and its status code, with the message, the events and the
status description dropped. Stack traces come off every span, attribute and event
alike.
A destination's resource attributes now merge onto the span's resource instead of
rebuilding one per span, which was re-running resource detection on every export.
* fix(otel): redact tenant URL query parameters
* fix(otel): close final tenant routing gaps
* fix(otel): refresh destinations for stateful MCP messages
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>