Commit graph

183 commits

Author SHA1 Message Date
mateo-berri
4f04e59ca0 fix: harden vertex live passthrough against client model forms and dict credentials
- accept the Live SDK's models/<id> and LiteLLM's vertex_ai/<id> when rewriting the setup model
- keep a dict service account intact instead of stringifying it
- treat same-target deployments holding different credentials as ambiguous
- guard both websocket states before every close so a second close cannot raise
- build the sendable close codes from the public CloseCode enum
2026-08-20 03:05:48 -07:00
mateo-berri
d434787a20 fix: refuse to guess a vertex project when live passthrough has no model hint 2026-08-20 02:19:38 -07:00
mateo-berri
b9d977aeee fix: guard vertex live passthrough provider lookup and close-code relay 2026-08-20 02:15:17 -07:00
mateo-berri
021a09b156 fix(passthrough): resolve vertex live credentials from db model deployments
The /vertex_ai/live WebSocket passthrough only ever looked at
default_vertex_config and the DEFAULT_VERTEXAI_* env vars, so a proxy whose
Vertex credentials live in the DB as a model entry with use_in_pass_through
had nothing to authenticate with. The upgrade still succeeded and the socket
then closed with a bare 1000 on the first client frame, which gave the client
no way to tell a misconfiguration from a normal end of session.

Credentials now also resolve from the router deployments flagged
use_in_pass_through, preferring the one matching the requested model, and a
failure to mint an access token closes 1011 with a reason naming both ways to
configure it. Upstream closes other than a plain 1000 are relayed to the client
with their code and reason, so Google's own errors reach the caller. The setup
frame's model is rewritten to the full projects/.../publishers/google/models
resource path, which is what Vertex expects and what lets a bare model id or a
gateway alias work over this route.
2026-08-20 02:00:29 -07:00
mateo-berri
1140366bee fix(vertex_ai): resolve passthrough serving location in the logging cost recompute 2026-08-19 17:24:44 -07:00
mateo-berri
c549cddada fix(vertex_ai): price passthrough calls on the URL's serving location 2026-08-19 16:44:52 -07:00
mateo-berri
81914ebc31 fix(proxy): log spend for OpenAI passthrough embeddings with unmapped models 2026-08-18 19:48:02 -07:00
yuneng-jiang
c33b3a32a6
feat(bedrock): add a config toggle to disable agent-runtime pass-through (#37386)
* feat(bedrock): add a config toggle to disable agent-runtime pass-through

The /bedrock pass-through dispatches agents, knowledge bases, flows, rerank,
retrieveAndGenerate, generateQuery and optimize-prompt to bedrock-agent-runtime,
so an operator who only wants to expose model invoke and converse has no way to
narrow that surface

Adds general_settings.disable_bedrock_agent_runtime_passthrough. When set, those
routes are rejected with a 403 before credentials are fetched or the request is
signed. Plain bedrock-runtime model pass-through is unaffected, and the setting
defaults to off, so existing deployments behave exactly as before

The branch is inverted to an early return for the non-agent-runtime case so the
toggle can reject outright instead of falling through to model extraction, which
would surface a confusing 400 about an unparseable model

* style(bedrock): drop redundant docstrings from the agent-runtime toggle
2026-08-18 17:05:40 -07:00
Yassin Kortam
1857f5d04b
fix(proxy): send SSE keepalives while a slow upstream is still silent (#37322)
A model with a long time-to-first-token leaves the proxy's response completely
idle, so any hop with an idle read timeout (AWS ALB and nginx both default to
60s) drops a connection that is perfectly healthy and would have delivered its
tokens shortly after.

The keepalive engines LiteLLM already ships wrap the response object, so they
fill a gap once the upstream has answered and then gone quiet. They cannot fill
the gap before it answers at all, and that is where the whole wait is spent:
measured against api.openai.com/v1/chat/completions with gpt-5.6 at
reasoning_effort high, the response headers and the first body byte both arrive
at 37.90s. Nothing has entered the ASGI response phase by then.

The upstream call is now raced against the keepalive interval, and when it
loses, the SSE response is opened immediately and ": ping" comments, which every
conformant SSE client ignores, fill the wire until the real response is ready to
be replayed onto it. One seam per funnel: base_process_llm_request covers every
native route, create_pass_through_route covers every passthrough route.

Committing the status line that early is the cost. A failure discovered after
the first ping reaches the client as an SSE error frame under a 200 rather than
as an HTTP error status, and LiteLLM's own x-litellm-* response headers are not
yet known. keepalive_ping_has_fired already documents the same trade-off for the
existing engines. Both are why this stays off until an operator sets
litellm_settings.sse_keepalive_ping_interval_seconds.

Separately, the passthrough relay reached neither engine even for mid-stream
gaps, which is the shape of #32491 and #24929, so the relayed bytes get the same
treatment, gated on the upstream declaring text/event-stream and only emitted
between complete frames so a binary transport (AWS event streams on /bedrock)
and a stall halfway through a frame are both left alone.

Fixes #34819
2026-08-18 14:43:01 -07:00
Mateo Wang
9cd7696156
Merge pull request #37229 from BerriAI/litellm_comprehend_medical_passthrough
feat(proxy): add Amazon Comprehend Medical passthrough provider
2026-08-17 17:25:01 -07:00
mateo-berri
7f42c84f57 fix(passthrough): dispatch Comprehend Medical logging on the provider tag only
Config-driven pass_through_endpoints pointed at a comprehendmedical.*.amazonaws.com
target were being claimed by the Comprehend Medical logging handler through the
hostname arm, which overrode their operator-set cost_per_request and relabeled
their spend rows. Only the built-in /comprehendmedical routes tag the provider,
so match on that alone.

Also mirror /comprehendmedical into the helm ingress and terraform gateway
prefix lists that hand-copy gateway/routes/allowlist.py
2026-08-17 17:07:31 -07:00
mateo-berri
915a1cabcd feat(proxy): add Amazon Comprehend Medical passthrough provider 2026-08-17 15:44:06 -07:00
mateo-berri
91c12ec810 test(proxy): run managed passthrough limit tests in CI 2026-08-17 15:39:46 -07:00
mateo-berri
6e55a21ebb fix(proxy): enforce batch list limit bounds on managed passthrough listings 2026-08-17 12:15:22 -07:00
mateo-berri
4ba9d6b136 fix(proxy): expose url join helper at module level for websocket route 2026-08-16 14:29:21 -07:00
mateo-berri
81aefe4b3c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr36151_ws_passthrough
# Conflicts:
#	litellm/proxy/pass_through_endpoints/pass_through_endpoints.py
#	tests/test_litellm/proxy/pass_through_endpoints/test_pass_through_endpoints.py
2026-08-16 14:21:57 -07:00
mateo-berri
a258b2b130 fix(proxy): harden OpenAI websocket passthrough
- decode upstream first frame as utf-8 instead of ascii
- reject model-restricted keys at connect to match HTTP model enforcement
- log the actual request path for /openai_passthrough traffic
2026-08-16 13:51:50 -07:00
mateo-berri
90493a217f fix(passthrough): protect accept-encoding from x-pass- forwarding 2026-08-15 15:34:28 -07:00
mateo-berri
d0be6eee8a fix(passthrough): stop forwarding client Accept-Encoding upstream 2026-08-15 15:22:06 -07:00
mateo-berri
0a81e1b222 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_azure_ai_docs_index_write_grant_rc 2026-08-14 13:58:28 -07:00
lostmartian
7a519e26ec
fix(proxy): track spend for OpenAI passthrough /v1/embeddings (#36660)
* fix(proxy): track spend for OpenAI passthrough /v1/embeddings

OpenAI passthrough embeddings returned 200 but wrote no spend because the
route was unsupported and Cohere's /v1/embed prefix stole the match.

* fix(proxy): clear embeddings lint and Greptile comment nits

Inline embeddings cost tracking to avoid new LIT001/002 hits, trim
redundant doc comments, and cover the Cohere /v1/embeddings collision.

* fix(proxy): drop unreachable embeddings TypeError guard

convert_to_model_response_object with response_type=embedding already
returns EmbeddingResponse; the isinstance check was dead patch coverage.
2026-08-13 20:48:16 -07:00
Noah Nistler
f8fccec108 fix(azure_ai): enforce admin-only index create on the passthrough route
POST /azure_ai/indexes carries no index name, so
get_azure_ai_search_index_from_endpoint returns None,
is_vector_store_index never matches any segment, and the request falls
through to the generic Azure passthrough on the proxy's own
AZURE_API_BASE and AZURE_API_KEY without ever reaching
is_allowed_to_call_vector_store_endpoint. A non-admin could therefore
create a Search index whenever AZURE_API_BASE points at the Search
service.

The earlier lifecycle commit made this look covered. Its test asserts
that POST /indexes?api-version=... is refused with "Only proxy admins can
create", but it calls the permission gate directly, and that gate is
exactly what the route skips for a path with no index name, so the guard
was verified in isolation while the route stayed open.

Gate the service-level create on the route itself, before the segment
loop, with assert_proxy_admin_for_vector_store_index_management. Scope it
to POST on a path whose last segment is indexes, mirroring the
endswith("/indexes") branch the lifecycle helper already uses, so the
managed-index paths and ordinary Azure OpenAI passthrough traffic are
untouched.

Add route-level tests: a non-admin is refused with the admin-only message
and never reaches the passthrough handler, an admin still creates, and the
new predicate is parametrized over the service-level, per-index, and
non-Search paths.
2026-08-13 17:56:57 +00:00
Noah Nistler
bdc80b11ac fix(azure_ai): authorize the targeted Search index, not any matching path segment
The Azure passthrough scanned every URL segment for one matching a registered
index, authorized against that, then forwarded the original path. A caller with
a grant on a managed index named e.g. "index" or "docs" could send
POST /azure_ai/indexes/{victim}/docs/index: the scan matched the trailing
segment and authorized on the caller's own index while Azure applied the batch
write to {victim} on the same Search service, enabling cross-index document
uploads or deletions.

Resolve the index positionally from the /indexes/{name} segment and require
that exact name to be the one authorized and credentialed, so the authorized
index and the physical target can never diverge. Add a pure helper plus
regression tests covering positional extraction and the route-level cross-index
attack.
2026-08-13 17:56:57 +00:00
yucheng-berri
0e9da56f89
fix(batches): strip NUL bytes from passthrough batch tags before the managed object write (#36688)
PostgreSQL rejects NUL in jsonb with 22P05, and the tags go into the managed
object's CREATE payload, so one poisoned tag aborts the whole row insert rather
than just that column. With no LiteLLM_ManagedObjectTable row, CheckBatchCost
never discovers the batch, so a batch that really ran and billed at the provider
produces no spend at all. The create-time write is fire and forget, so nothing
retries it.

This regressed in #36468, which started passing request_tags and
persist_attribution from the Anthropic passthrough; before that no
caller-supplied string reached the column.

Sanitize in the shared helper that builds the value, matching how
spend_tracking_utils already handles LiteLLM_SpendLogs.request_tags. Both the
Anthropic and the Vertex passthrough build tags through that one helper, so this
covers both. Rename it to _sanitized_str_tuple since it no longer merely
coerces.
2026-08-12 13:31:51 -07:00
Yassin Kortam
258fe3e4ba
fix(passthrough): carry the budget reservation into request metadata (#36592)
A successful pass-through request left its pre-call budget reservation in
the shared Redis spend counter. `_init_kwargs_for_pass_through_endpoint`
built the request metadata from the sanitized key fields only, so
`_PROXY_track_cost_callback` resolved `budget_reservation = None` and
`increment_spend_counters` added the actual cost on top of a reservation
nobody released. The counter drifted above real spend on every request
until the key falsely tripped BudgetExceededError, while the Postgres
spend stayed far below the limit. The failure path was unaffected because
it releases `user_api_key_dict.budget_reservation` directly.

The reservation is now set alongside the other internal keys, after the
client-supplied metadata merge, so a request body cannot forge one that
names arbitrary counter keys.
2026-08-12 12:34:13 -07:00
Mateo Wang
23b805d5a4
Merge pull request #36447 from BerriAI/litellm_anthropic_fast_mode_speed_usage
fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through
2026-08-12 00:05:59 -07:00
Mateo Wang
7e80e094c4
Merge pull request #36529 from william-xue/fix-responses-passthrough-stream-cost
fix(proxy): track streamed passthrough Responses cost
2026-08-11 21:58:42 -07:00
mateo-berri
08a73740ec fix(passthrough): keep prompt/completion token split for streamed OpenAI rows 2026-08-11 21:28:55 -07:00
mateo-berri
5e14649c54 fix(passthrough): bill streamed Responses calls that end failed
A stream can terminate with a response.failed event that still reports
consumed tokens; those were rebuilt as None and logged at zero spend.
Parse response.failed alongside completed and incomplete, matching the
buffered path, which prices any terminal response that reports usage.
2026-08-11 21:01:18 -07:00
yucheng-berri
8bfb7772e4
fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags (#36468)
The Anthropic batch create never persisted the creating key's hashed token or its
request tags on the managed object, so when CheckBatchCost billed the batch hours
later there was nothing to attribute it to. Key spend, key budgets and tag spend
never moved for batch usage.

Persist both from the create, the way the Vertex passthrough already does, and
register the batch only from the collection route. An id-scoped route cannot
rebuild the unified object id, because it embeds the model and the model comes
from the create's request body, so it could only claim a row it did not create or
fail the model_object_id unique constraint.

The shared metadata helpers, the route predicate and the registration-result
logging now live in batch_attribution instead of being copied per provider. The
Anthropic write previously logged success unconditionally, before the
fire-and-forget task had run.

Resolves LIT-5288
2026-08-11 20:52:43 -07:00
mateo-berri
2df121c821 fix(passthrough): bill streamed Responses calls that end incomplete
Streams that terminate with response.incomplete (e.g. max_output_tokens
reached) carry real usage in the terminal event but were rebuilt as None
and logged at zero spend, letting callers bypass budget enforcement.
Parse response.incomplete alongside response.completed when
reconstructing the streamed response.
2026-08-11 20:21:17 -07:00
mateo-berri
dc30e1816d refactor(passthrough): move Responses stream terminal-event parsing into OpenAI provider config
Addresses review feedback: the ResponseCompletedEvent SSE parsing now lives
in OpenAIResponsesAPIConfig next to the other Responses stream event handling,
and the proxy logging handler calls it. Adds coverage for streams that end
without a response.completed event.
2026-08-11 20:11:53 -07:00
william-xue
c8655c3825 fix(proxy): track streamed passthrough Responses cost 2026-08-11 18:30:42 +08:00
mateo-berri
e7c8cff3b7 fix(proxy): preserve crlf line endings when injecting streamed usage cost 2026-08-11 00:45:32 -07:00
mateo-berri
938396ef90 fix(proxy): recognize crlf sse frame boundaries in passthrough reassembly 2026-08-10 21:37:07 -07:00
mateo-berri
46fb1cd514 fix(proxy): reassemble fragmented SSE frames and inject logging dependency 2026-08-10 20:07:57 -07:00
mateo-berri
426b909447 fix(proxy): inject streaming usage cost on openai passthrough streams 2026-08-10 19:56:34 -07:00
Devin AI
f0c3d8dcda chore: merge litellm_internal_staging into litellm_anthropic_fast_mode_speed_usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-10 19:28:09 +00:00
mateo-berri
bd719c21dc test(passthrough): annotate route match scope as Final 2026-08-09 11:49:20 -07:00
mateo-berri
85c1b5d04a Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_fix_openai_passthrough_files_route_36086
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-08-08 18:01:59 -07:00
yucheng-berri
efc4e6f28c
fix(batches): keep batch state in sync on a poll without claiming attribution (#34456)
A poll of a Vertex passthrough batch wrote nothing to the managed-object row,
so status and file_object stayed frozen at the create-time snapshot and
GET /v1/batches served a stale status and an empty output file id for the life
of the batch. Only the create may claim a batch, but every observation of one
may refresh its state.

store_unified_object_id takes create_if_missing, which the poll clears: it
refreshes status and file_object through update_many, and leaves a row that is
absent absent rather than creating one owned by the observer, since created_by
and team_id are written by whoever reaches the create branch. The update payload
is now shared with the upsert so it cannot drift into writing api_key,
request_tags, created_by or team_id.

The passthrough identity re-assertion that was previously part of this PR ships
separately in #36121, so this PR keeps only the batch attribution work.

The creating key owns user_api_key_alias only when it actually has one. Guarding
the overwrite on the presence of a key rather than on a resolved alias nulled the
field out for every key generated without key_alias, and for any key rotated or
deleted before its batch finished, losing the creating user's alias that the spend
row previously carried. The guard now matches the team-alias line below it.
2026-08-08 16:01:47 -07:00
Devin AI
5a5bb8c9d8 fix(proxy): stop /{provider}/v1/files from capturing /openai_passthrough
The native files and batches routes declare /{provider}/v1/... and their routers are mounted before the passthrough router, so /openai_passthrough/v1/files and /openai_passthrough/v1/batches matched them with provider="openai_passthrough" and 500'd on the LlmProviders lookup instead of reaching openai_proxy_route.

Move the dedicated /openai_passthrough prefix onto its own router mounted ahead of the batches and files routers. /openai/... and every other provider prefix keep their current behavior.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-06 15:29:39 +00:00
mateo-berri
1fefd80925 fix(proxy): resolve pass-through credentials live from router deployments 2026-08-04 23:01:37 -07:00
Devin AI
81a80b8c63 fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through
Fast mode is priced with a provider-specific multiplier applied off usage.speed, but only chat completions kept that field. The Messages route rebuilt usage with empty optional params, stream reassembly dropped speed and inference_geo, and the pass-through handler never read speed off the request body, so fast-mode spend was logged at the standard rate.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-02 01:02:27 +00:00
mateo-berri
47ebc964eb test: patch unified guardrail mapping global instead of loader to fix order-dependent flake 2026-07-30 20:55:32 -07:00
mateo-berri
0b09588685 refactor(batches): aggregate batch output cost, usage, and models in a single pass
Completed-batch cost tracking parsed the whole output file into a list of
dicts, pretty-printed it into debug strings even with debug logging off, and
walked the list three times (cost, usage, models), so a large batch output
could pin a worker's memory. The output is now folded line by line into small
per-line stats records via _aggregate_batch_cost_usage_models, the eager
json.dumps debug calls are gone, and the raw-vertex path computes cost and
usage in one call instead of two. _get_batch_output_file_content_as_dictionary
becomes _fetch_batch_output_file_content (returns bytes); the superseded
three-pass helpers are deleted and their tests migrated
2026-07-29 22:00:01 -07:00
Tin Chi Lo
970ea2949e fix(vertex): decide rawPredict passthrough streaming from the request body
Vertex passthrough classified any target URL containing "stream" as a streaming
request. `:streamRawPredict` carries that substring, so a unary Claude-on-Vertex
call whose body omits `stream` was routed through the streaming logging path.
That path never consults the response content-type, so a complete
`"type": "message"` JSON body was handed to the Anthropic SSE chunk parser,
which recognises none of it; the spend log recorded 0 prompt tokens,
0 completion tokens and zero cost

Streaming for the rawPredict family now comes from the request body, which is
what the Anthropic Messages contract uses for those endpoints. The
generateContent family keeps its URL signal because the Gemini REST body has no
`stream` field, and `?alt=sse` is still appended for every request that is
classified as streaming, so Gemini framing and its usage parsing are unchanged

Both passthrough streaming predicates read `.get("stream")` off a body that is
only annotated as a dict; `_read_request_body` returns whatever the JSON parser
produced, so an array body raised AttributeError. The two predicates are now one
owner that answers False for any non-object body, which covers the vertex,
mistral, anthropic, vllm and azure passthrough routes
2026-07-25 18:09:29 -07:00
Yuneng Jiang
9c48ad41ac
fix(passthrough): honor the zero fallback and aggregate-only TPM usage
Two follow-ups from review on the upstream-reported usage contract.

An unusable cost header fell through to the endpoint's flat cost_per_request
instead of the zero the contract promises, so a target that contradicted itself
got billed an estimate it had just disowned. A target that speaks this contract
now owns the cost for the request whether or not the value it sent parsed.

The reported total also cannot be split into prompt and completion, so reading
one out of it under token_rate_limit_type input or output yielded zero and left
the TPM window uncharged; pass-through traffic then ran past a limit it is
meant to share with the general API. Usage that carries no split now charges
its total under every limit type, while usage that does carry one is untouched.
2026-07-24 18:49:41 -07:00
Yuneng Jiang
ab44c8a8ce
fix(passthrough): let an upstream-reported cost outrank cost_per_request
PassThroughGenericEndpoint.cost_per_request defaults to 0.0, so every
config-defined endpoint forwards a flat 0.0 even when the operator never
configured one, and the success handler applied it over whatever cost was
already established. That silently zeroed the cost an upstream reported for
the request. The flat value is an estimate for targets LiteLLM cannot price,
so it now yields to a target that priced the request itself; it still applies
unchanged when no cost was reported.
2026-07-24 18:26:12 -07:00
Yuneng Jiang
838c7a7ea7
feat(passthrough): record upstream-reported cost and token usage
A pass-through target that fans a single HTTP request out to several models
internally cannot be priced from its response body, so LiteLLM had nothing to
record and every such request landed in the spend logs with zero cost and zero
tokens. The target now reports the totals for the whole request in
x-litellm-response-cost and x-litellm-total-tokens response headers, and
LiteLLM records those values as-is rather than recomputing them.

The headers are read on every upstream response, so a request that burned
tokens before failing still books its spend on the failure row instead of
being dropped for having a 4xx/5xx status. Only what the upstream actually
reported is written, so a target that sends a cost but no token count keeps
the token count LiteLLM derived on its own; a target that sends neither header
is untouched, which is the normal case for Anthropic, Vertex and friends.

Two supporting fixes fall out of this. The rate limiter only pulled token
counts off response shapes it models, so pass-through usage never charged the
TPM window and a team could exceed its shared token limit through pass-through
traffic alone; it now falls back to combined_usage_object. And the streaming
success path reset response_cost unconditionally before the assembled response
recomputed it, which discarded any cost a pass-through handler had already
established (the pass-through branch right below it has always intended to
preserve exactly that).
2026-07-24 18:10:15 -07:00