litellm/tests
tin-berri a613773fca
feat(auto-router)!: scope shadow eval jobs to multiple keys (#37251)
* feat(auto-router): scope shadow eval jobs to multiple keys

A shadow eval job now covers a set of keys instead of exactly one, and each
key carries its own max_turns budget, so one key exhausting its budget leaves
its siblings sampling. The existing job row already is the per-key unit
(api_key_id, max_turns, stopped_at, and the one-active-per-key-and-direction
partial unique index all live on it), so multi-key is grouping rather than
schema surgery: a new group_id column ties N sibling rows written atomically
by one create_many, the API's job id becomes the group id, and pre-existing
jobs backfill group_id = id so their ids keep resolving. The sampler hot path
is untouched; its test file has a zero-line diff

Results come back pooled plus a per-key breakdown and responses list every key
with its own budget, stop state and read-time labels. The dashboard is adapted
minimally to the new shapes (the picker stays single-key and submits a one-key
list); the multi-select picker and per-key table land in the stacked UI PR

* fix(shadow_eval): derive completed from spent budgets and record operator stops

* fix(shadow_eval): stamp stops atomically and freeze counts at the stamp

The stop endpoint wrote stopped_by and stopped_at as two separate updates, so
a failure between them left a job reading stopped while its unstamped legs
kept sampling, and the retry got 400 already stopped. One UPDATE now stamps
stopped_by and every missing stopped_at together, preserving the stopped_at a
leg earned from its own budget via COALESCE

Attempt counts now exclude attempts that land after a leg's stopped_at, so an
in-flight attempt finishing just after an operator stop can never push a
legacy pre-stopped_by job over its budget and flip it from stopped to
completed at read time

* fix(shadow_eval): backfill stopped_by so legacy stops never read as completions

* chore(ui): regenerate api types for the shadow eval stop fields

* fix(shadow_eval): let the stop statement pick one winner under racing stops

Two operators can both pass the derived-status guard in the race window. The
stop UPDATE now claims only legs with stopped_by still null and the endpoint
judges by its row count, so exactly one caller ever gets the 200 and the loser
gets the same already-stopped 400 a late caller gets

* refactor(shadow_eval): make the stop statement the whole state machine

The status guard ran before the UPDATE, so a stop racing the last budgeted
attempt still claimed the job and it read stopped forever instead of
completed. The statement now claims the job only while a leg still samples
inside the window with no stop recorded, and the endpoint reads once after
writing: a racing operator, a same-instant budget spend, and a repeat stop all
get the 400 naming the status the job actually holds. The pre-write guard and
the hand-built response go away

* chore(ui): regenerate api types for the stop route description
2026-08-19 14:02:15 -07:00
..
agent_tests test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
audio_tests test: remove tests that never execute 2026-08-12 10:45:38 -07:00
base_sdk_tests fix(deps): ship boto3 with the base SDK so bedrock works out of the box (#36568) 2026-08-11 14:39:10 -07:00
basic_proxy_startup_tests
batches_tests test: build redaction and batch limiter fixtures the way production does (#37416) 2026-08-18 19:37:50 -07:00
benchmarks
code_coverage_tests feat(search): add Nimble as a search provider (#36347) 2026-08-14 17:09:58 -07:00
documentation_tests feat(vector_stores): add Valkey as a managed vector store provider (#37002) 2026-08-18 21:45:22 +00:00
e2e test(e2e): pin the tag-routing denial to its actual cause 2026-08-18 20:49:12 -07:00
enterprise Merge pull request #37198 from BerriAI/litellm_lit5660_batches_limit_400 2026-08-17 15:53:46 -07:00
guardrails_tests fix(guardrails): honor configured timeout in Zscaler AI Guard (#36110) 2026-08-07 00:25:52 +00:00
image_gen_tests test: point the live gemini and groq conformance suites at models that still exist (#37422) 2026-08-19 02:50:46 +00:00
integration
litellm fix(proxy): deny agent access when key and team grants resolve to nothing (#36221) 2026-08-07 20:44:11 +00:00
litellm-proxy-extras test: remove tests that never execute 2026-08-12 10:45:38 -07:00
litellm_core_utils
litellm_utils_tests refactor(responses): drop commentary from the tool_choice fix 2026-08-17 13:04:45 -07:00
llm_responses_api_testing test: refresh three suites that drifted from the code they cover 2026-08-15 15:37:09 -07:00
llm_translation test: point the live gemini and groq conformance suites at models that still exist (#37422) 2026-08-19 02:50:46 +00:00
load_tests
local_testing test: move the remaining live groq call sites off the retired llama models (#37426) 2026-08-18 20:26:17 -07:00
logging_callback_tests test: build redaction and batch limiter fixtures the way production does (#37416) 2026-08-18 19:37:50 -07:00
mcp_tests fix(mcp): keep REST tool listing in step with key/team grant enforcement 2026-07-30 22:13:10 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): update Azure DI supported-params assertion for req_format 2026-08-18 19:21:46 -07:00
old_proxy_tests/tests
openai_endpoints_tests fix(batches): register managed output files on batch cancel 2026-08-05 18:28:52 -07:00
otel_tests
pass_through_tests
pass_through_unit_tests test: allow protocol-constrained pass-through routes to declare fewer methods (#37415) 2026-08-18 19:38:05 -07:00
proxy_admin_ui_tests fix(access groups): sync assigned_team_ids from the team write paths (#36825) 2026-08-14 04:45:36 +00:00
proxy_behavior feat(proxy): add /team/daily/activity/aggregated and switch the Usage team tab to it (#36562) 2026-08-18 11:29:57 -07:00
proxy_e2e_anthropic_messages_tests
proxy_migration_tests test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot (#36136) 2026-08-07 09:57:37 -07:00
proxy_security_tests
proxy_unit_tests fix(batches): price poller-tracked batches from the deployment's registered rates 2026-08-17 15:36:39 -07:00
router_unit_tests feat(complexity_router): custom classifier plugins via classifier_type 'custom' (#37249) 2026-08-18 14:09:19 -07:00
scim_tests
search_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr32448_tinyfish_headers 2026-08-18 13:03:33 -07:00
spend_tracking_tests
store_model_in_db_tests test: refresh three suites that drifted from the code they cover 2026-08-15 15:37:09 -07:00
test_litellm feat(auto-router)!: scope shadow eval jobs to multiple keys (#37251) 2026-08-19 14:02:15 -07:00
unified_google_tests
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_ws_vcr.py
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: refresh three suites that drifted from the code they cover 2026-08-15 15:37:09 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.