litellm/tests
tin-berri 4e48d74455
feat(shadow_eval): measure both arms' cost so a job reports what the router would have saved (#38631)
The attempt row now prices the real arm (the payload's response_cost plus its own
routing classifier when it routed) beside the shadow arm (completion plus the
classifier cost the routing decision writes back), and flags turns litellm's
response cache served. A per-leg funnel table counts the eligible requests that
produced no row (lost the sampling dice, unjudgeable shape, concurrency shed),
so results can weigh judged rows against the traffic they stand for. Job results
gain per-slice and overall arm spends plus the coverage counts, the budget gates
charge the shadow arm's classifier spend against max_budget, and the dashboard
shows the measured cost comparison beside the win rate

Resolves LIT-6358

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 15:13:19 -07:00
..
agent_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
audio_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
base_sdk_tests test(cli): cover the keyless token record and keep keyring to the cli extra 2026-08-20 03:45:00 -07:00
basic_proxy_startup_tests
batches_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
benchmarks
code_coverage_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260821 2026-08-27 09:17:37 +00:00
documentation_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
e2e test(e2e): cover Together reasoning_effort=none, json_schema, and cache-read pricing 2026-08-28 12:49:09 -07:00
enterprise test(proxy): narrow pytest.raises to HTTPException/ProxyException 2026-08-27 21:05:48 -07:00
guardrails_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
image_gen_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
integration test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
litellm-proxy-extras fix: keep schema reconciliation from fighting a partitioned LiteLLM_SpendLogs (#38452) 2026-08-27 12:52:52 -07:00
litellm_utils_tests test(aiohttp): pin NO_PROXY so proxy env cannot hijack the refused-port probe 2026-08-28 12:45:46 -07:00
llm_responses_api_testing test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
llm_translation test(together_ai): assert fail-open supported params for models missing from the registry 2026-08-27 01:27:32 -07:00
load_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ruff_dead_test_code 2026-08-24 09:46:56 -07:00
local_testing test(batch): make the upstream-failure tolerance actually reachable 2026-08-28 00:25:53 -07:00
logging_callback_tests test: add litellm_gateway_injected_cache to gcs_pub_sub spend fixture 2026-08-28 12:36:49 -07:00
mcp_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260821 2026-08-27 09:17:37 +00:00
multi_instance_e2e_tests test: say whether a match= pattern is a regex or a literal (ruff RUF043) 2026-08-21 16:25:33 -07:00
ocr_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
openai_endpoints_tests test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
otel_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
pass_through_tests test(passthrough): drop the Azure assistants test, retired upstream on 2026-08-26 2026-08-27 23:58:23 -07:00
pass_through_unit_tests test(passthrough): name the test for what it now covers 2026-08-28 09:11:48 -07:00
proxy_admin_ui_tests test(access-groups): give each xdist worker its own fixture ids 2026-08-28 09:35:12 -07:00
proxy_behavior fix(proxy): reset a stuck team member's budget (#37971) 2026-08-25 09:50:09 -07:00
proxy_e2e_anthropic_messages_tests ci: lint the test tree for undefined names and fix all 30 (#37671) 2026-08-20 13:30:34 -07:00
proxy_migration_tests fix(ui): boot the UI image as an arbitrary uid by anchoring nginx writes under /tmp (#37982) 2026-08-24 11:57:36 -07:00
proxy_security_tests
proxy_unit_tests refactor: dedupe server_tool_use web search reads and type fresh test locals 2026-08-27 08:01:52 +00:00
router_unit_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
search_tests fix(search): harden bing_grounding auth, result cap, status, and cost 2026-08-24 12:26:07 -07:00
spend_tracking_tests
store_model_in_db_tests test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
test_litellm feat(shadow_eval): measure both arms' cost so a job reports what the router would have saved (#38631) 2026-08-28 15:13:19 -07:00
unified_google_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
vector_store_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
windows_tests test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
__init__.py
_fake_openai_endpoint_server.py test(ci): serve /moderations from the canned OpenAI mock (#37739) 2026-08-20 16:58:38 -07:00
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_wait_helpers.py test: replace blind sleeps with deadline waits in callback and caching tests (#37660) 2026-08-20 18:48:43 +00:00
_ws_vcr.py
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py ci: lint the test tree for undefined names and fix all 30 (#37671) 2026-08-20 13:30:34 -07:00
test_fallbacks.py test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py test: fix staging CI regressions from #38182, #38144, #38265, #37962, and #37969 2026-08-25 23:01:20 -07:00
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_openai_endpoints.py test: point the live web search, groq and vertex image suites at models that still exist (#37733) 2026-08-20 17:03:35 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_resource_cleanup.py
test_service_logger_otel.py fix(langfuse): send v4 ingestion header for otel callback (#33907) 2026-07-18 20:36:51 -07:00
test_spend_logs.py
test_team.py test: add six ruff rules that catch tests which cannot fail (#37709) 2026-08-20 14:21:26 -07:00
test_team_logging.py test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
test_team_members.py test: reject assertions on a caught error inside except (ruff PT017) 2026-08-21 13:35:08 -07:00
test_users.py test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.