litellm/tests
ryan-crabbe-berri ee5efaed7b Give e2e tests typed metadata and a recorded step log
A test's JUnit report says what happened to it, never what it was about.
`@pytest.mark.covers("cell.id")` is a registry key, not a description: it
cannot answer "which tests drive /v1/responses on Anthropic", and nothing
in the report says where a failing test actually died.

Two halves, deliberately separated, both riding out as JUnit <property>
entries the downstream emitter already knows how to read.

DECLARED - `@meta(Subject(...))` from the new tests/e2e/e2e_metadata.py.
One frozen dataclass, every field a closed enum (domain, route, provider,
model, capabilities, mode), so a typo is a basedpyright error at the call
site rather than a property that silently never appears. Serialization is
one pass over `dataclasses.asdict`, so a new scalar field needs no
serializer edit; `capabilities` is deduped and sorted at declaration so
committed run artifacts diff cleanly whatever order a test spelled it in.
Empty fields emit nothing - the suite does not pad every testcase with
five empty entries.

RECORDED - `@step("POST /chat/completions")` on harness helpers, never on
tests. Each call appends its label to the running test's user_properties
in call order, so the list IS the test's user story and cannot drift from
what the test did. The label is recorded BEFORE the wrapped call, so a
helper that raises still leaves its own label last: a failing test's last
step is where it died. Consecutive duplicates collapse and the log caps at
50, so a poll loop is one beat of the story rather than fifty.

Steps cannot be attached where the other properties are -
`pytest_collection_modifyitems` runs before any test body, so the recorder
is empty there. They attach from the existing `pytest_runtest_makereport`
wrapper on the call phase, which is what puts them on failures too, and an
autouse fixture empties the log at setup. The attach drops any prior step
entries first, because the suite runs `--reruns 1` and a retry would
otherwise stack a second copy of the story behind the first.

`covers` is untouched: the marker is separate because a dataclass passed
to `covers` would be dropped silently by `dedupe_covers` and would hard-
fail collection in tests/integration/conftest.py. The fixed
package/covers/source prefix stays byte-identical, and `@meta` goes BELOW
`@covers` so `Item.location` still anchors at the first decorator and
every `source` deep link keeps pointing where it pointed.

`Provider` mirrors litellm's `LlmProviders` values instead of importing
them, so nothing here - the module or its call sites - imports litellm.
tests/e2e is a black-box HTTP suite that is copied to the runner image on
its own, so a `from litellm...` at the top of a test module would make the
package a COLLECTION-time dependency: where it is absent, every test in
the suite errors out before running rather than importing slowly. The
mirror cannot drift silently - `TestProviderMirrorsLitellm` asserts every
value is a real `LlmProviders` value wherever litellm is importable, and
skips where it is not, which is the property it is guarding.

A declared `model` names the constant the test drives, never a copy of its
value: `CHEAP_ANTHROPIC_MODEL` and `CHEAP_OPENAI_MODEL` are env-overridable
(`E2E_CHEAP_ANTHROPIC_MODEL`, `E2E_CHEAP_OPENAI_MODEL`), so a hardcoded
default would have reported a model the run never touched. The same holds
for a file's own `BACKEND`/`MODEL` constant, where the copy was merely
waiting to drift.

Pilot: tests/e2e/quota_management, all 85 tests annotated and its three
clients plus cost_rows @step-decorated, to prove the API against real
tests rather than a toy. The rest of the suite is a later backfill.

Verified: 42 harness unit tests in test_junit_properties.py (19 new), 550
harness unit tests green, basedpyright over tests/e2e at the same 29
pre-existing errors as origin/main (all in the untouched mcp/
oauth_chat_client.py), ruff clean, the coverage-registry collector
byte-identical before and after (473/586), and the test-quality gate OK
against origin/main. Collection was also run with the litellm package
blocked at the import hook: 1339/1350 collected either way, the one error
being the pre-existing missing `httpx2` in mcp/. The live e2e tests need a
deployed proxy and real provider keys and were not run.
2026-09-21 10:56:52 -07:00
..
agent_tests fix(a2a): keep the upstream status on card discovery failures and inject the card client in tests 2026-09-16 23:55:57 +00:00
audio_tests
base_sdk_tests fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
basic_proxy_startup_tests
batches_tests
benchmarks
code_coverage_tests Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-19 23:59:26 -07:00
documentation_tests fix(docs-test): restore line-anchored table regex in router settings check 2026-09-18 20:10:20 +00:00
e2e Give e2e tests typed metadata and a recorded step log 2026-09-21 10:56:52 -07:00
enterprise fix(projects): persist explicit budget cap clears 2026-09-12 13:43:55 -07:00
guardrails_tests Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context" 2026-09-19 17:45:15 +00:00
image_gen_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
integration Merge pull request #42097 from BerriAI/litellm_budget_exceeded_422 2026-09-21 08:35:30 -05:00
litellm-proxy-extras Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-18 20:55:05 -07:00
litellm_utils_tests Merge remote-tracking branch 'origin/main' into litellm_invalid_tool_choice_400 2026-09-19 02:35:29 -07:00
llm_responses_api_testing test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
llm_translation test(bedrock): expect canonical session tags in the dynamic auth params propagation test 2026-09-18 11:17:34 -07:00
load_tests
local_testing Merge pull request #42105 from BerriAI/litellm_flip_v2_migration_resolver_default 2026-09-21 10:05:59 -07:00
logging_callback_tests Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
mcp_tests Merge pull request #42019 from BerriAI/litellm_master_key_boot_enforcement 2026-09-19 19:04:02 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix 2026-09-18 21:27:03 +00:00
openai_endpoints_tests test(responses): fix stale Anthropic smoke request 2026-09-11 13:56:27 -07:00
otel_tests feat(cli): deprecate the litellm-proxy entrypoint in favour of lite 2026-09-17 14:05:31 -07:00
pass_through_tests fix(mcp): preserve legacy behavior on SDK2 and streamline verification 2026-09-18 22:28:31 -07:00
pass_through_unit_tests test(pass_through): shorten protocol-constrained route docstring 2026-09-17 00:24:29 +00:00
proxy_admin_ui_tests
proxy_behavior Merge pull request #41906 from BerriAI/litellm_team_member_budget_source_reset 2026-09-21 10:34:08 -07:00
proxy_e2e_anthropic_messages_tests fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1 2026-09-19 23:33:29 +00:00
proxy_migration_tests Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-19 23:59:26 -07:00
proxy_security_tests refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
proxy_unit_tests Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
router_unit_tests Merge remote-tracking branch 'origin/main' into litellm_pr38499_batch_retrieve_model_group 2026-09-19 02:12:25 -07:00
rust-python-harness test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
search_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
spend_tracking_tests
store_model_in_db_tests test(proxy): expect the sanitized unknown-model message in the spend-log error test 2026-09-11 19:18:10 -07:00
test_gateway
test_litellm Merge pull request #42121 from BerriAI/litellm_utils_model_info_lookup 2026-09-21 10:42:28 -07:00
test_litellm_rust test(rust): pin child interpreters to the parent's litellm and lint for it 2026-09-19 18:17:54 +00:00
unified_google_tests test(unified_google_tests): import ReadOnly from typing_extensions and cover the Vertex global endpoint 2026-09-19 12:15:34 -07:00
unit test(a2a): sort imports in merged bedrock agentcore test 2026-09-21 16:10:40 +00:00
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py refactor(test): validate cost-map entries into a typed model 2026-09-16 14:32:08 -07:00
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
AGENTS.md ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py refactor(rust): remove gateway, config, router, realtime, and trace-parity infrastructure 2026-09-16 16:00:07 +00:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.