litellm/tests/e2e/batches/COVERAGE.md
Sameer Kankute a16d9c6f9e
test(e2e): add live batches suite across providers and routing scenarios (#30958)
* tests: add e2e tests for spend, budgets and llms

* style: make chained comparison of status_code clearer

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* remove e2e_tests folder

* test: add spend tracking tests

* fix: p0 issues, added types and shared functions for each test suite

* style: carry clearer status_code comparison into renamed e2e dir

* refactor: migrate to gateway client

* fix: add new tests, split gateway

* test(e2e): add live batches suite across providers and routing scenarios

* test(batches): cover real cost tracking on completed batch retrieve

* test(e2e): assert managed vs raw file and batch id shapes per routing scenario

* test(e2e): assert full response shape of each batches and files endpoint

* test(e2e): only accept transitional statuses for a freshly created batch

* test(prompt-factory): make test_convert_url deterministic with a data URL

picsum.photos is down (HTTP 522), so test_convert_url failed on every
run. Swap the live external image for an inline data: URL and assert the
round-trip through convert_url_to_base64 genuinely.

A data URL is already inline base64 image data, so convert_url_to_base64
now short-circuits it instead of attempting an impossible HTTP fetch;
add a regression for that branch in the mapped image_handling test

* fix: pass through async image data urls

* fix(image-handling): short-circuit data URLs in async path too

Bugbot flagged that convert_url_to_base64 returns data: base64 URLs
unchanged but async_convert_url_to_base64 still tried to fetch them,
so async OCR flows (Bedrock, Azure) would reject inline images the sync
path accepts. Add the same guard to the async function and a regression
test that asserts the async path returns the data URL without touching
the HTTP client

* Fix: openai batches lifecycle

* Fix: add e2e azure openai tests

* Fix e2e for vertex ai

* Add all models for testing

* test(managed-files): assert idempotent upsert in store_unified_file_id

store_unified_file_id switched from create to upsert to avoid
UniqueViolationError when re-storing the same unified_file_id (e.g.
batch output files stored before metadata is available). Update the
unit test to assert the upsert call and its create payload instead of
the removed create call.

* test(batches): reconcile vertex_ai native batch-id comment with fallback guard

* fix(test-config): keep rust-ocr models in model_list by moving files_settings after it

* fix(test-config): move batch models after OCR block to keep merge with internal_staging clean

* fix(batches): use '24hrs' completion window and allow managed-files listing with provider filter

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* style: ruff format transformation.py and endpoints.py

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(e2e/batches): set Azure raw_model to gpt-4.1-mini-batch to match deployed model

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(vertex-ai/batches): correct completion_window to 24h per Literal type definition

* test(vertex-ai/batches): align completion_window assertion to 24h

* fix: update managed file metadata on upsert

---------

Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-02 08:05:23 -07:00

4.4 KiB

Batches Test Coverage Matrix

Live e2e coverage of the Batches API over a real proxy, real provider keys, and real cost. Synchronous tier only: a batch's completion window is 24h, so these tests never wait for completed. They assert the proxy accepts, routes, retrieves, cancels, and lists a batch; everything created is deleted on teardown.

Provider x operation

Only supported cells are tested. The capability table in capabilities.py holds one row per supported (provider, scenario) pair, so there are no skipped cells in the parametrized run.

Provider create retrieve cancel list file backing
OpenAI yes yes yes yes OpenAI Files
Azure yes yes yes yes Azure Files
Vertex AI yes yes yes yes GCS bucket
Bedrock yes yes no (limited upstream) no S3 bucket
Anthropic no yes (env-gated) no no Anthropic Files

Bedrock cancel is unreliable upstream and list is unsupported, so both are gated off (can_cancel=False, can_list=False). Anthropic cannot create/cancel/list through litellm, so it has a standalone retrieve test that skips unless ANTHROPIC_BATCH_ID points at a real Anthropic batch.

Routing scenarios (per litellm/proxy/batches_endpoints/endpoints.py)

Each create-capable provider runs all four. The test asserts the returned file id and batch id carry the shape that scenario must produce (matches_id_shape):

Scenario How the batch is routed File id Batch id
encoded upload with ?model= -> model-encoded file id -> create with just that id model-encoded model-encoded
unified upload with target_model_names= -> unified managed file id -> create with that id managed managed
model_param raw file (provider-fallback upload) -> create with model in the body raw model-encoded
provider_fallback raw file -> POST /{provider}/v1/batches, env creds, no model raw raw (native provider shape)

"managed" ids base64-decode to a litellm_proxy marker; "model-encoded" ids keep the provider prefix and base64-encode litellm:<id>;model,<model>; "raw" ids are the provider's native ids. Asserting these catches a proxy that returns a raw id where it should manage it, or vice versa. On top of the id shape, a misroute to the wrong provider also fails create (the file id / model do not belong there), and the provider_fallback raw batch id is additionally checked against the provider's native shape (raw_id_matches_provider).

Key model restriction

test_batch_key_model_access_denied mints a key restricted to one model (resources.key(models=[...])) and proves the proxy returns 403 key_model_access_denied both when that key uploads a file for a disallowed model (files endpoint) and when it creates a batch for a disallowed model (batches endpoint).

Per-endpoint output assertions

Each endpoint's full response is validated, not just the id. File upload asserts object=="file", purpose=="batch", a positive bytes, a status, and a created-at. Batch create / retrieve assert object=="batch", endpoint=="/v1/chat/completions", completion_window=="24h", a non-empty input_file_id, and a created-at; retrieve additionally cross-checks that id and input_file_id match the created batch. Cancel asserts the same id, object=="batch", and a cancelling/cancelled status. List asserts the object=="list" envelope and that the created batch is present as a batch. File delete asserts object=="file" and deleted==True.

This suite's files

File Covers
batch_client.py typed file upload/download + batch create/retrieve/cancel/list/delete over the shared Gateway; denial helpers
capabilities.py the provider x scenario matrix + id-shape classifiers + per-provider raw-id assertion
test_batches_e2e.py parametrized lifecycle with per-endpoint output assertions, file upload/delete outputs, key-model-access denial, anthropic retrieve

Out of scope (intentionally)

Driving a batch to completed, cost tracking on completion, and the DB write-back are not covered here; the 24h window makes them unfit for a synchronous gate. That logic belongs in a DI-stubbed proxy integration test under tests/test_litellm/proxy/ where the provider client is injected to return completed deterministically.