litellm/tests
ryan-crabbe-berri d0e37d39c4 Record each e2e test's steps from the harness it calls
A test's JUnit report says whether it passed, never what it did or where a failing test died. This records that from the harness, so nothing about it is hand-written and it cannot drift from what the test actually ran

`@step("create team with a budget")` from the new tests/e2e/e2e_metadata.py goes on harness helpers, never on tests, and appends its label to the running test's step log in call order. The label is recorded before the wrapped call, so a helper that raises still leaves its own label last: a failing test's last step is where it died. Every public harness method that performs an action now carries one, 355 across the client modules, lifecycle, idp, the logging readers, migrations and the claude_code driver

Only the outermost step records, tracked per thread. Harness layers call each other (ResourceManager.key goes through ProxyClient.generate_key, a domain client wraps the shared ProxyClient), so every layer carries a label and the story still reads at the level the test called in at, one beat per action. A step above @contextmanager holds the guard through __enter__ and __exit__, so a context's cleanup never lands behind the step a test died on, and a bare generator function is refused at import because its body interleaves with its caller's. Consecutive duplicates collapse and the log caps at 50, so a poll loop is one beat rather than fifty. The wrapper is a frame, so the eight cleanup and retry warnings raised directly inside decorated helpers use stacklevel=2 + STEP_FRAMES to keep reporting at their caller

The log is emptied first thing in pytest_runtest_setup and attached from the existing pytest_runtest_makereport wrapper after setup and again after call, so a test that errors in a fixture keeps the steps recorded before the crash. Teardown does not attach: finalizer steps are cleanup. Each attach drops the item's earlier step entries, so the second attach and a --reruns 1 retry replace the story rather than doubling it

Steps ride out as repeated <property name="step"> entries behind the fixed package/covers/source prefix, which stays byte-identical. The project-releaser emitter already regroups them into the results JSON's steps array. test_junit_report.py runs real pytest with --junitxml against this conftest, in-process and under -n 2, and pins the passing, failing, setup-error, rerun and wide-scope-fixture cases on the parsed XML
2026-09-21 18:48:59 -07:00
..
agent_tests fix(a2a): keep the upstream status on card discovery failures and inject the card client in tests 2026-09-16 23:55:57 +00:00
audio_tests
base_sdk_tests fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
basic_proxy_startup_tests
batches_tests
benchmarks
code_coverage_tests Merge pull request #42143 from BerriAI/litellm_e2e_changed_keep_pytest_log 2026-09-21 11:39:25 -07:00
documentation_tests fix(docs-test): restore line-anchored table regex in router settings check 2026-09-18 20:10:20 +00:00
e2e Record each e2e test's steps from the harness it calls 2026-09-21 18:48:59 -07:00
enterprise fix(projects): persist explicit budget cap clears 2026-09-12 13:43:55 -07:00
guardrails_tests Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context" 2026-09-19 17:45:15 +00:00
image_gen_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
integration Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 21:04:53 +00:00
litellm-proxy-extras Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-18 20:55:05 -07:00
litellm_utils_tests Merge remote-tracking branch 'origin/main' into litellm_invalid_tool_choice_400 2026-09-19 02:35:29 -07:00
llm_responses_api_testing test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
llm_translation test(bedrock): expect canonical session tags in the dynamic auth params propagation test 2026-09-18 11:17:34 -07:00
load_tests
local_testing fix(ci): let the install smoke test boot its key-less proxy config 2026-09-21 12:56:21 -07:00
logging_callback_tests Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
mcp_tests fix(mcp): preserve Python 3.10 imports and integration test seams 2026-09-21 12:58:06 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix 2026-09-18 21:27:03 +00:00
openai_endpoints_tests test(responses): fix stale Anthropic smoke request 2026-09-11 13:56:27 -07:00
otel_tests test: fix stale budget-status and bad-database-url assertions 2026-09-21 15:00:54 -07:00
pass_through_tests fix(mcp): preserve legacy behavior on SDK2 and streamline verification 2026-09-18 22:28:31 -07:00
pass_through_unit_tests test(pass_through): shorten protocol-constrained route docstring 2026-09-17 00:24:29 +00:00
proxy_admin_ui_tests
proxy_behavior Merge pull request #42026 from BerriAI/litellm_user_jwt_savings 2026-09-21 16:41:16 -07:00
proxy_e2e_anthropic_messages_tests fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1 2026-09-19 23:33:29 +00:00
proxy_migration_tests Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-19 23:59:26 -07:00
proxy_security_tests refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
proxy_unit_tests Merge pull request #39578 from Louis-Vauterin/jwt-key-mapping-token-id 2026-09-21 18:22:54 -07:00
router_unit_tests Merge pull request #42283 from BerriAI/litellm_mid_stream_fallback_walks_full_list 2026-09-21 15:14:04 -07:00
rust-python-harness test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
search_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
spend_tracking_tests
store_model_in_db_tests test(proxy): expect the sanitized unknown-model message in the spend-log error test 2026-09-11 19:18:10 -07:00
test_gateway feat(proxy): offload spend tracking to a pod-local collector sidecar (#40545) 2026-09-10 17:14:13 -07:00
test_litellm Merge pull request #42057 from BerriAI/litellm_classifier_forecast_cards 2026-09-21 17:55:25 -07:00
test_litellm_rust fix(cache): keep native Redis semantic binding and Qdrant batch writes after merge 2026-09-22 00:41:24 +00:00
unified_google_tests test(google): boot the unified Google proxy fixture with a real master key 2026-09-20 08:48:39 +00:00
unit chore: merge main into litellm_cherry_pick_password_breach_reset 2026-09-21 19:54:40 +00:00
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py refactor(test): validate cost-map entries into a typed model 2026-09-16 14:32:08 -07:00
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
AGENTS.md ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py refactor(rust): remove gateway, config, router, realtime, and trace-parity infrastructure 2026-09-16 16:00:07 +00:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.