* tests: add e2e tests for spend, budgets and llms * style: make chained comparison of status_code clearer Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * remove e2e_tests folder * test: add spend tracking tests * fix: p0 issues, added types and shared functions for each test suite * style: carry clearer status_code comparison into renamed e2e dir * refactor: migrate to gateway client * fix: add new tests, split gateway * test(e2e): add live batches suite across providers and routing scenarios * test(batches): cover real cost tracking on completed batch retrieve * test(e2e): assert managed vs raw file and batch id shapes per routing scenario * test(e2e): assert full response shape of each batches and files endpoint * test(e2e): only accept transitional statuses for a freshly created batch * test(prompt-factory): make test_convert_url deterministic with a data URL picsum.photos is down (HTTP 522), so test_convert_url failed on every run. Swap the live external image for an inline data: URL and assert the round-trip through convert_url_to_base64 genuinely. A data URL is already inline base64 image data, so convert_url_to_base64 now short-circuits it instead of attempting an impossible HTTP fetch; add a regression for that branch in the mapped image_handling test * fix: pass through async image data urls * fix(image-handling): short-circuit data URLs in async path too Bugbot flagged that convert_url_to_base64 returns data: base64 URLs unchanged but async_convert_url_to_base64 still tried to fetch them, so async OCR flows (Bedrock, Azure) would reject inline images the sync path accepts. Add the same guard to the async function and a regression test that asserts the async path returns the data URL without touching the HTTP client * Fix: openai batches lifecycle * Fix: add e2e azure openai tests * Fix e2e for vertex ai * Add all models for testing * test(managed-files): assert idempotent upsert in store_unified_file_id store_unified_file_id switched from create to upsert to avoid UniqueViolationError when re-storing the same unified_file_id (e.g. batch output files stored before metadata is available). Update the unit test to assert the upsert call and its create payload instead of the removed create call. * test(batches): reconcile vertex_ai native batch-id comment with fallback guard * fix(test-config): keep rust-ocr models in model_list by moving files_settings after it * fix(test-config): move batch models after OCR block to keep merge with internal_staging clean * fix(batches): use '24hrs' completion window and allow managed-files listing with provider filter Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * style: ruff format transformation.py and endpoints.py Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(e2e/batches): set Azure raw_model to gpt-4.1-mini-batch to match deployed model Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(vertex-ai/batches): correct completion_window to 24h per Literal type definition * test(vertex-ai/batches): align completion_window assertion to 24h * fix: update managed file metadata on upsert --------- Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com> Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
4.4 KiB
Batches Test Coverage Matrix
Live e2e coverage of the Batches API over a real proxy, real provider keys, and
real cost. Synchronous tier only: a batch's completion window is 24h, so these
tests never wait for completed. They assert the proxy accepts, routes, retrieves,
cancels, and lists a batch; everything created is deleted on teardown.
Provider x operation
Only supported cells are tested. The capability table in capabilities.py holds one
row per supported (provider, scenario) pair, so there are no skipped cells in the
parametrized run.
| Provider | create | retrieve | cancel | list | file backing |
|---|---|---|---|---|---|
| OpenAI | yes | yes | yes | yes | OpenAI Files |
| Azure | yes | yes | yes | yes | Azure Files |
| Vertex AI | yes | yes | yes | yes | GCS bucket |
| Bedrock | yes | yes | no (limited upstream) | no | S3 bucket |
| Anthropic | no | yes (env-gated) | no | no | Anthropic Files |
Bedrock cancel is unreliable upstream and list is unsupported, so both are gated off
(can_cancel=False, can_list=False). Anthropic cannot create/cancel/list through
litellm, so it has a standalone retrieve test that skips unless ANTHROPIC_BATCH_ID
points at a real Anthropic batch.
Routing scenarios (per litellm/proxy/batches_endpoints/endpoints.py)
Each create-capable provider runs all four. The test asserts the returned file id
and batch id carry the shape that scenario must produce (matches_id_shape):
| Scenario | How the batch is routed | File id | Batch id |
|---|---|---|---|
encoded |
upload with ?model= -> model-encoded file id -> create with just that id |
model-encoded | model-encoded |
unified |
upload with target_model_names= -> unified managed file id -> create with that id |
managed | managed |
model_param |
raw file (provider-fallback upload) -> create with model in the body |
raw | model-encoded |
provider_fallback |
raw file -> POST /{provider}/v1/batches, env creds, no model |
raw | raw (native provider shape) |
"managed" ids base64-decode to a litellm_proxy marker; "model-encoded" ids keep the
provider prefix and base64-encode litellm:<id>;model,<model>; "raw" ids are the
provider's native ids. Asserting these catches a proxy that returns a raw id where it
should manage it, or vice versa. On top of the id shape, a misroute to the wrong
provider also fails create (the file id / model do not belong there), and the
provider_fallback raw batch id is additionally checked against the provider's native
shape (raw_id_matches_provider).
Key model restriction
test_batch_key_model_access_denied mints a key restricted to one model
(resources.key(models=[...])) and proves the proxy returns 403
key_model_access_denied both when that key uploads a file for a disallowed model
(files endpoint) and when it creates a batch for a disallowed model (batches
endpoint).
Per-endpoint output assertions
Each endpoint's full response is validated, not just the id. File upload asserts
object=="file", purpose=="batch", a positive bytes, a status, and a created-at.
Batch create / retrieve assert object=="batch", endpoint=="/v1/chat/completions",
completion_window=="24h", a non-empty input_file_id, and a created-at; retrieve
additionally cross-checks that id and input_file_id match the created batch.
Cancel asserts the same id, object=="batch", and a cancelling/cancelled status. List
asserts the object=="list" envelope and that the created batch is present as a batch.
File delete asserts object=="file" and deleted==True.
This suite's files
| File | Covers |
|---|---|
batch_client.py |
typed file upload/download + batch create/retrieve/cancel/list/delete over the shared Gateway; denial helpers |
capabilities.py |
the provider x scenario matrix + id-shape classifiers + per-provider raw-id assertion |
test_batches_e2e.py |
parametrized lifecycle with per-endpoint output assertions, file upload/delete outputs, key-model-access denial, anthropic retrieve |
Out of scope (intentionally)
Driving a batch to completed, cost tracking on completion, and the DB write-back
are not covered here; the 24h window makes them unfit for a synchronous gate. That
logic belongs in a DI-stubbed proxy integration test under tests/test_litellm/proxy/
where the provider client is injected to return completed deterministically.