litellm/tests/e2e
devin-ai-integration[bot] c9e8a04139
feat(vertex): native batch JSONL passthrough with cost tracking (#42810)
* feat(vertex): native batch JSONL passthrough with cost tracking

Add a per-request `passthrough=true` multipart field on `POST /v1/files`
(and the same kwarg on `litellm.create_file`) that uploads a native
Vertex AI batch JSONL to the deployment's GCS bucket unchanged, so rows
using `googleSearch` and other Gemini-only features run as written and
the output, `groundingMetadata` included, comes back untouched.

Passthrough is sticky through the GCS object path
(`litellm-vertex-files/passthrough/...`), so batch create and output
retrieval inherit it without new state. Native output rows are costed
from their `usageMetadata` with the deployment's model and model_info,
in the polling and retrieve paths and for the existing global
`disable_vertex_batch_output_transformation` flag, which billed $0
before.

The proxy requires the target to resolve to vertex_ai deployments only,
refuses `passthrough` with a non-batch purpose, a non-default
`target_storage`, or pre-call guardrails, and validates native rows on
`request` instead of the OpenAI batch keys.

* refactor(vertex): keep native batch row pricing inside the Vertex adapter

Moves native Vertex batch row detection, response parsing, and per-row
pricing from litellm/batches/batch_utils.py into
litellm/llms/vertex_ai/batches/transformation.py, so batch_utils only
aggregates the rows it gets back. Adds tests/test_litellm/files to the
misc unit shard so the new test directory is claimed by a shard.

* fix(files): say what a passthrough batch upload takes when a row is not native

The missing-key 400 listed bare key names, so an OpenAI-shaped row under
passthrough=true read "Each line must be a JSON object with keys request".
The batch line shape now carries its own hint, and the passthrough one says
a passthrough upload takes native Vertex batch rows with a request key

* fix(batches): bill native Vertex embedding batch rows on the native cost path

A native Vertex output row whose response holds an embedding was validated as a
generateContent response, so the documented tokenCount-only shape counted as a failed
row. Price embedding rows from their own usage (promptTokenCount, else tokenCount) with
the helper the transformed embeddings path already used, and drop the prompt-details
helper nothing calls anymore.

* fix(batches): keep modality batch rates on native Vertex embedding rows

An embedding row that carries usageMetadata was billed from promptTokenCount alone, so
its promptTokensDetails no longer reached the audio, image, and video batch rates the
way it did before the native cost path. Run every row with usageMetadata through the
Gemini usage parser and keep the flat tokenCount fallback for embedding rows without it.

* fix(batches): price native Vertex batch rows by modelVersion under a wildcard deployment

A `vertex_ai/*` deployment hands the batch cost path `*` as the deployment model, which
no cost map resolves, so every native (passthrough or flag-on) row was billed at $0. A
wildcard deployment model now defers to the row's own `modelVersion`, the way the
transformed path already prices by the row's `model`.

Also moves the native passthrough tests under tests/test_litellm, the tree codecov
reads, and covers the raw upload chunking, the embedding output translation, the
unpriceable-row path, and the flag-on dispatch.

* fix(batches): keep explicit deployment prices for native Vertex rows without a modelVersion

Under a wildcard deployment a native batch row that carries no modelVersion (an embedding
row, or a generateContent row Vertex returned without one) was billed at $0 even when the
deployment's model_info sets explicit batch prices, because the cost calculator was never
called. The row now falls back to the wildcard name, which the cost calculator prices from
the explicit model_info, and only a row with neither a modelVersion nor a deployment model
is billed at $0 with the warning

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 12:35:34 -07:00
..
a2a test(e2e): settle control-plane writes across every replica, not just one 2026-08-07 19:36:30 -07:00
access_control Merge remote-tracking branch 'github/main' into litellm_rust_bridge_declarative_route_catalog 2026-09-17 11:08:08 -07:00
batches feat(vertex): native batch JSONL passthrough with cost tracking (#42810) 2026-09-24 12:35:34 -07:00
claude_code chore(e2e): move the compat-matrix populator from a GCE VM to a Render cron job (#42608) 2026-09-22 19:43:31 -07:00
coverage_registry feat(vertex): native batch JSONL passthrough with cost tracking (#42810) 2026-09-24 12:35:34 -07:00
gateway test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
guardrails fix(policy_engine): keep inherited parent guardrails when a child policy condition misses (#42548) 2026-09-22 23:27:38 -07:00
llm_translation test(e2e): tolerate provider-side flakes on five full-suite cells (#42628) 2026-09-22 20:12:14 -07:00
load feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
logging fix(otel): honor SSL_CERT_FILE and ssl_verify in OTLP HTTP exporters (#42106) 2026-09-22 20:50:53 +00:00
management fix(proxy): write key deleted audit logs for cascade and alias key deletions (#42446) 2026-09-22 11:52:08 -07:00
mcp test(e2e): restore LIT-3467 implementation for rework 2026-09-19 16:21:53 -07:00
migrations test(migrations): close the gaps the upgrade assertions left open 2026-09-21 13:15:46 -07:00
other fix(proxy): list key and team model aliases in GET /v1/models (#42908) 2026-09-24 06:41:13 -07:00
quota_management feat(logging): add normalized_error cluster key to error_information (#41715) 2026-09-22 15:56:50 -07:00
router test(e2e): tolerate provider-side flakes on five full-suite cells (#42628) 2026-09-22 20:12:14 -07:00
secret_manager test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
ui test(e2e/ui): give the logout specs their own admin session (#42930) 2026-09-24 11:55:07 -07:00
AGENTS.md test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
conftest.py test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
CONTRIBUTING.md test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
e2e_config.py test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
e2e_db.py test(e2e): guard destructive spend-log truncate behind an explicit opt-in (#33751) 2026-07-20 08:47:39 -07:00
e2e_http.py feat(vertex): native batch JSONL passthrough with cost tracking (#42810) 2026-09-24 12:35:34 -07:00
fixture_bundle.py test: add strict stateless provider replay identity 2026-09-14 16:55:43 -07:00
fixture_canonical.py feat(e2e): key the provider cache per test and mount Bedrock behind it 2026-09-16 02:15:46 -07:00
fixture_mode.py fix(e2e): own a shared fixture's deployment by the fixture's node, not the first test 2026-09-16 17:35:05 -07:00
fixture_profile.py test: preserve strict replay numeric spelling 2026-09-14 17:11:56 -07:00
idp.py test(e2e): restore LIT-3467 implementation for rework 2026-09-19 16:21:53 -07:00
idp_realm.json test(e2e): harden JWT fixtures and cover management lifecycles 2026-09-11 16:36:12 -07:00
junit_properties.py test: bind management E2E callers and isolate JWT actors 2026-09-12 13:29:04 -07:00
lifecycle.py feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
memory_readings.py test(e2e): hold every worker under an idle RSS budget before any traffic (#42552) 2026-09-22 14:47:14 -07:00
models.py fix(proxy): list key and team model aliases in GET /v1/models (#42908) 2026-09-24 06:41:13 -07:00
otel_client.py test(e2e): harden the suite against response-cache cross-talk, slow providers and single upstream blips (#37957) 2026-08-22 14:47:03 -07:00
PROVIDER_CACHE.md fix(e2e): own a shared fixture's deployment by the fixture's node, not the first test 2026-09-16 17:35:05 -07:00
provider_cache.py fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process 2026-09-16 17:08:17 -07:00
provider_cache_redis.py chore(e2e): report the key components behind a mount that never converges 2026-09-16 15:01:08 -07:00
provider_cache_routing.py revert(e2e): unmount Gemini, its api_base means two things 2026-09-16 08:37:25 -07:00
provider_edge.py fix(responses): forward safety_identifier through the chat completion bridge 2026-09-21 04:09:59 +00:00
provider_edge_bedrock.py feat(e2e): cache the responses and embeddings endpoints behind the edge 2026-09-16 02:34:09 -07:00
proxy_client.py test(e2e): hold every worker under an idle RSS budget before any traffic (#42552) 2026-09-22 14:47:14 -07:00
pytest.ini test(e2e): add secret manager lanes for HashiCorp Vault and CyberArk Conjur (#42503) 2026-09-22 18:02:32 -07:00
stack_lock.py test(e2e): run the memory cell alone on the shared stack (#42518) 2026-09-22 15:13:37 -07:00
test_e2e_http.py test: bind management E2E callers and isolate JWT actors 2026-09-12 13:29:04 -07:00
test_fixture_bundle.py feat(e2e): record and replay streamed provider responses chunk-for-chunk 2026-08-24 12:51:44 -07:00
test_fixture_canonical.py test(e2e): pin query params and multipart form fields as replay match-key identity 2026-08-20 15:59:34 -04:00
test_fixture_mode.py feat(e2e): move record/replay to the provider edge (LIT-5745) 2026-08-19 18:39:15 -07:00
test_idp.py test: enforce isolated actors and stop OIDC process groups 2026-09-12 13:49:49 -07:00
test_junit_properties.py test(e2e): read JUnit properties off the real collected pytest Item 2026-09-01 19:12:07 -07:00
test_provider_edge.py fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process 2026-09-16 17:08:17 -07:00
test_proxy_client.py test(e2e): assert a cooldown reaches a sibling replica within the 1s Redis read interval (#42422) 2026-09-22 12:42:41 -07:00
test_stack_lock.py test(e2e): run the memory cell alone on the shared stack (#42518) 2026-09-22 15:13:37 -07:00
transport.py fix(e2e): route credential, cost map, and UI login calls to the control plane (#42506) 2026-09-22 12:59:56 -07:00