litellm/tests
Yuneng Jiang 6553bc8957
fix(proxy): serialize model reconciles so concurrent writes stop evicting each other
A model write is a read-modify-write of the shared `llm_router` global: read the
db into a snapshot, then make the router match that snapshot. Nothing serialized
it, so two of them interleaving was not a lost update but an eviction --
_delete_deployment removes every live deployment absent from the snapshot it was
handed, so the request holding the older snapshot reconciles the newer request's
model straight back out of the router. The row survives in the db, which is what
makes it easy to miss: the pod simply stops serving a model it was told to serve
until some later reload happens to put it back.

clear_cache compounds it. It deletes every db model from the router before
reloading them, so for the width of that reload the pod serves none of them --
and any concurrent write sampling the router in that window sees the hole.

Fix is one lock (MODEL_RECONCILE_LOCK) held across both, so each reconcile reads
the db and applies it atomically and no stale snapshot can evict a newer model.
clear_cache holds it across wipe+reload and calls the already-locked
_add_deployment_locked, since asyncio.Lock is not reentrant and routing back
through the public add_deployment would deadlock the pod's whole model-write
path.

The verdict needed the same treatment. raise_if_reload_degraded_serving compared
a desired-set read during the reload against a router snapshot taken after it,
so a neighbouring reconcile's in-flight wipe was reported to the caller as
collateral damage from its own reload -- a 500 on a create that had in fact
succeeded. Reconciles now return a ReconcileOutcome carrying both the desired set
and the post-reconcile serving state, captured before the lock is released, and
the verdict judges against that. Omitting live_after keeps the old live re-read,
which stays correct for the no-reconcile-ran case.

Found by running the e2e suite with pytest-xdist at 8 workers: three unrelated
tests failed together on "Previously served model id(s) [...] are also no longer
being served by this pod", which is this. Serial runs concurrent enough to hit it
are rare, which is why 78 minutes of sequential e2e never surfaced it -- but any
customer provisioning models in parallel (terraform, CI) is in exactly this race.

test_reconciles_serialize_so_no_stale_snapshot_can_evict fails with 5 == 1
without the lock.
2026-08-12 10:43:29 -07:00
..
agent_tests test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
audio_tests
base_sdk_tests fix(deps): move pydantic-settings into the base dependencies 2026-08-01 15:25:26 -07:00
basic_proxy_startup_tests
batches_tests test(batches): run Responses coverage in CI 2026-07-31 10:24:10 -04:00
benchmarks test(benchmarks): run shared logging executor inline to make CodSpeed measurements deterministic (#32435) 2026-07-09 11:14:22 +03:00
code_coverage_tests fix(guardrails): scan /v1/messages tool traffic 2026-08-05 14:11:27 -07:00
documentation_tests fix(ci): make the env-key doc gate see bare get_secret and get_secret_str reads (#35996) 2026-08-05 14:49:55 -07:00
e2e Merge pull request #36277 from BerriAI/litellm_make_check_fallback 2026-08-08 12:08:34 -07:00
enterprise Merge pull request #36049 from BerriAI/litellm_list_batches_resolves_unified_ids 2026-08-07 17:59:09 -07:00
guardrails_tests fix(guardrails): honor configured timeout in Zscaler AI Guard (#36110) 2026-08-07 00:25:52 +00:00
image_gen_tests test(bedrock): switch image gen live test off EOL Titan to Nova Canvas (#31937) 2026-07-01 22:57:29 -07:00
integration
litellm fix(proxy): deny agent access when key and team grants resolve to nothing (#36221) 2026-08-07 20:44:11 +00:00
litellm-proxy-extras
litellm_core_utils
litellm_utils_tests test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
llm_responses_api_testing test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
llm_translation test(bedrock): port the openai-route invoke tests onto the live config 2026-07-29 20:59:08 -07:00
load_tests
local_testing test: repair three failing suites on litellm_internal_staging 2026-08-04 16:07:02 -07:00
logging_callback_tests test(logging): pin routing_decision and internal_call_origin in the gcs pubsub spend log fixture 2026-08-01 14:34:04 -07:00
mcp_tests fix(mcp): keep REST tool listing in step with key/team grant enforcement 2026-07-30 22:13:10 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): use mistral-document-ai-2512 in azure_ai OCR tests 2026-07-15 18:13:22 -07:00
old_proxy_tests/tests
openai_endpoints_tests fix(batches): register managed output files on batch cancel 2026-08-05 18:28:52 -07:00
otel_tests
pass_through_tests test(pass-through): de-flake vertex spend-log test by routing through the proxy (#31689) 2026-06-30 15:27:48 -07:00
pass_through_unit_tests fix(proxy): re-assert the authenticated identity on passthrough requests (#36121) 2026-08-07 00:41:29 +00:00
proxy_admin_ui_tests Revert "chore: remove _experimental/out (#31546)" 2026-07-01 13:25:47 -07:00
proxy_behavior test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
proxy_e2e_anthropic_messages_tests
proxy_migration_tests test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot (#36136) 2026-08-07 09:57:37 -07:00
proxy_security_tests
proxy_unit_tests Merge pull request #36048 from BerriAI/litellm_cancelled_batch_unified_output_ids 2026-08-06 10:40:07 -07:00
router_unit_tests test(router): assert the auto-router max_input_chars kwarg 2026-08-06 11:26:32 -07:00
scim_tests
search_tests feat(tinyfish): make search provider permissive, attribute errors (#31997) 2026-07-03 10:17:11 -07:00
spend_tracking_tests
store_model_in_db_tests feat(team): custom metadata validation hook for team create and update (#33353) 2026-08-03 18:37:45 -07:00
test_litellm fix(proxy): serialize model reconciles so concurrent writes stop evicting each other 2026-08-12 10:43:29 -07:00
unified_google_tests
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_ws_vcr.py test(realtime): record and replay websocket traffic in redis vcr cassettes (#32390) 2026-07-08 00:19:06 -07:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py
test_organizations.py
test_otel_thread_leak.py
test_passthrough_endpoints.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_service_logger_otel.py fix(langfuse): send v4 ingestion header for otel callback (#33907) 2026-07-18 20:36:51 -07:00
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.