litellm/tests
Yuneng Jiang 274a489c46
test(e2e/ui): automate the RC checklist's Presidio guardrail walk
The guardrail section of the release checklist is done by hand every cut:
create a Presidio guardrail through the wizard, send a sentence with PII from
the playground, then open Logs and check the guardrail caught it. Nothing
covered that path, so a break anywhere along it surfaced only when someone
happened to repeat the steps.

Adds a spec that walks it once and turns the eyeball checks into assertions.
The strongest of them is the leak check: it reads the request back and fails if
the stored prompt still carries the raw address or number, which is what the
manual step is really looking for.

The stack gains a Presidio stand-in that answers the two routes the guardrail
calls, detecting a fixed regex set with a Luhn check on card numbers. Real
Presidio's detection quality is Presidio's business, and pinning the UI lane to
it would mean two heavy containers with spaCy models on every CI run for a test
that is about LiteLLM's integration. The real analyzer stays covered in the
Python lane. The stand-in returns the same entities at the same spans as the
real one for the checklist's sentence, and driving the real guardrail against it
produces the same record shape, so a test written against it is written against
the product's real behavior.

Making the stand-in return the text unmasked turns the spec red on the raw
address reaching the spend log, so the leak assertion reads live data.

run_e2e.sh and both CircleCI UI jobs start the stand-in alongside the mock LLM.
Its port is overridable like the others so two checkouts can run at once.
2026-09-06 02:36:46 -07:00
..
agent_tests
audio_tests
base_sdk_tests
basic_proxy_startup_tests
batches_tests merge(litellm_internal_staging): reconcile batch observability with per-line resilience 2026-08-24 19:08:40 -04:00
benchmarks
code_coverage_tests fix(e2e-changed): keep the gate off suites the stack cannot run 2026-09-05 21:03:50 -07:00
documentation_tests Merge remote-tracking branch 'origin/main' into litellm_bedrock_messages_disconnect_billing 2026-08-31 08:58:37 -07:00
e2e test(e2e/ui): automate the RC checklist's Presidio guardrail walk 2026-09-06 02:36:46 -07:00
enterprise fix(batches): register ownership for every batch create path (#39810) 2026-09-04 23:59:51 -07:00
guardrails_tests fix(logging): blocked requests no longer report guardrail_status=success in multi-guardrail configs (#39596) 2026-09-03 17:33:31 -07:00
image_gen_tests
integration
litellm-proxy-extras test(proxy-extras): fake run_prisma instead of subprocess.run in the migrate deploy harness 2026-09-03 16:10:44 -07:00
litellm_utils_tests test(aiohttp): pin NO_PROXY so proxy env cannot hijack the refused-port probe 2026-08-28 12:45:46 -07:00
llm_responses_api_testing fix(guardrails): deliver modify_response block as valid SSE on streaming chat and Responses 2026-08-31 16:01:39 -07:00
llm_translation test: exempt MockTransport request-shape embedding tests from VCR replay 2026-09-01 13:29:12 -07:00
load_tests feat(proxy): per-worker admission control that rejects excess requests with 503 (#39352) 2026-09-03 18:19:04 -07:00
local_testing fix(cache): use sync Redis batch reads (#39358) 2026-09-03 14:37:48 -07:00
logging_callback_tests Merge branch 'litellm_internal_staging' into litellm_fix_failing_request_slowdown 2026-09-03 00:20:51 -07:00
mcp_tests fix(mcp): scope allow-all servers to virtual keys (#39531) 2026-09-03 10:32:03 -07:00
multi_instance_e2e_tests
ocr_tests fix(vertex): avoid duplicate DeepSeek OCR model namespace 2026-09-01 14:39:15 -07:00
openai_endpoints_tests refactor(tests): assign the streamed id and lock poll once instead of rebinding 2026-09-03 13:49:43 -07:00
otel_tests
pass_through_tests fix(logging): key bridged /v1/messages rows on the id the caller received 2026-09-03 01:15:06 -07:00
pass_through_unit_tests test(websearch): carry a reasoned test-quality suppression on the router patch 2026-08-31 22:32:05 -07:00
proxy_admin_ui_tests refactor(tests): assign the streamed id and lock poll once instead of rebinding 2026-09-03 13:49:43 -07:00
proxy_behavior feat(team): report per-user spend within a team for JWT traffic (#39771) 2026-09-04 12:00:47 -07:00
proxy_e2e_anthropic_messages_tests
proxy_migration_tests fix(proxy-extras): kill the whole Prisma process group when a command times out 2026-09-02 18:29:55 -07:00
proxy_security_tests
proxy_unit_tests fix(jwt): evict mapping cache after DB write in /jwt/key/mapping update and delete 2026-09-04 14:35:42 -07:00
router_unit_tests test: repair two CI tests broken by intentional changes 2026-09-05 12:02:48 -07:00
rust-python-harness test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
search_tests fix(search): harden bing_grounding auth, result cap, status, and cost 2026-08-24 12:26:07 -07:00
spend_tracking_tests
store_model_in_db_tests test(store_model_in_db): accept both 400 shapes in the unknown-model spend log test 2026-09-04 18:52:17 -07:00
test_litellm Merge pull request #39723 from Atharva-Kanherkar/fix/anthropic-responses-refusal-translation 2026-09-06 00:03:06 -07:00
unified_google_tests
vector_store_tests fix(vector-store): carry request metadata into the Router executor built from the router kwarg 2026-09-02 16:53:05 -07:00
windows_tests
__init__.py
_fake_openai_endpoint_server.py test(timeout): time out against the local fake endpoint instead of api.openai.com 2026-09-03 09:53:30 -07:00
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py test: fix staging CI regressions from #38182, #38144, #38265, #37962, and #37969 2026-08-25 23:01:20 -07:00
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
test_new_vector_store_endpoints.py
test_openai_endpoints.py
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.