mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-12 23:01:41 +00:00
A model write is a read-modify-write of the shared `llm_router` global: read the db into a snapshot, then make the router match that snapshot. Nothing serialized it, so two of them interleaving was not a lost update but an eviction -- _delete_deployment removes every live deployment absent from the snapshot it was handed, so the request holding the older snapshot reconciles the newer request's model straight back out of the router. The row survives in the db, which is what makes it easy to miss: the pod simply stops serving a model it was told to serve until some later reload happens to put it back. clear_cache compounds it. It deletes every db model from the router before reloading them, so for the width of that reload the pod serves none of them -- and any concurrent write sampling the router in that window sees the hole. Fix is one lock (MODEL_RECONCILE_LOCK) held across both, so each reconcile reads the db and applies it atomically and no stale snapshot can evict a newer model. clear_cache holds it across wipe+reload and calls the already-locked _add_deployment_locked, since asyncio.Lock is not reentrant and routing back through the public add_deployment would deadlock the pod's whole model-write path. The verdict needed the same treatment. raise_if_reload_degraded_serving compared a desired-set read during the reload against a router snapshot taken after it, so a neighbouring reconcile's in-flight wipe was reported to the caller as collateral damage from its own reload -- a 500 on a create that had in fact succeeded. Reconciles now return a ReconcileOutcome carrying both the desired set and the post-reconcile serving state, captured before the lock is released, and the verdict judges against that. Omitting live_after keeps the old live re-read, which stays correct for the no-reconcile-ran case. Found by running the e2e suite with pytest-xdist at 8 workers: three unrelated tests failed together on "Previously served model id(s) [...] are also no longer being served by this pod", which is this. Serial runs concurrent enough to hit it are rare, which is why 78 minutes of sequential e2e never surfaced it -- but any customer provisioning models in parallel (terraform, CI) is in exactly this race. test_reconciles_serialize_so_no_stale_snapshot_can_evict fails with 5 == 1 without the lock. |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm | ||
| litellm-proxy-extras | ||
| litellm_core_utils | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| old_proxy_tests/tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| scim_tests | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _ws_vcr.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_config.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_entrypoint.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_passthrough_endpoints.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.