litellm/tests/test_litellm/proxy/management_endpoints
yuneng-jiang 32535987e8
fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other (#36687)
* fix(proxy): serialize model reconciles so concurrent writes stop evicting each other

A model write is a read-modify-write of the shared `llm_router` global: read the
db into a snapshot, then make the router match that snapshot. Nothing serialized
it, so two of them interleaving was not a lost update but an eviction --
_delete_deployment removes every live deployment absent from the snapshot it was
handed, so the request holding the older snapshot reconciles the newer request's
model straight back out of the router. The row survives in the db, which is what
makes it easy to miss: the pod simply stops serving a model it was told to serve
until some later reload happens to put it back.

clear_cache compounds it. It deletes every db model from the router before
reloading them, so for the width of that reload the pod serves none of them --
and any concurrent write sampling the router in that window sees the hole.

Fix is one lock (MODEL_RECONCILE_LOCK) held across both, so each reconcile reads
the db and applies it atomically and no stale snapshot can evict a newer model.
clear_cache holds it across wipe+reload and calls the already-locked
_add_deployment_locked, since asyncio.Lock is not reentrant and routing back
through the public add_deployment would deadlock the pod's whole model-write
path.

The verdict needed the same treatment. raise_if_reload_degraded_serving compared
a desired-set read during the reload against a router snapshot taken after it,
so a neighbouring reconcile's in-flight wipe was reported to the caller as
collateral damage from its own reload -- a 500 on a create that had in fact
succeeded. Reconciles now return a ReconcileOutcome carrying both the desired set
and the post-reconcile serving state, captured before the lock is released, and
the verdict judges against that. Omitting live_after keeps the old live re-read,
which stays correct for the no-reconcile-ran case.

Found by running the e2e suite with pytest-xdist at 8 workers: three unrelated
tests failed together on "Previously served model id(s) [...] are also no longer
being served by this pod", which is this. Serial runs concurrent enough to hit it
are rare, which is why 78 minutes of sequential e2e never surfaced it -- but any
customer provisioning models in parallel (terraform, CI) is in exactly this race.

test_reconciles_serialize_so_no_stale_snapshot_can_evict fails with 5 == 1
without the lock.

* fix(tests): return a ReconcileOutcome from the PTU test's add_deployment mock

test_ptu_model_settings.py stubs proxy_config.add_deployment with
AsyncMock(return_value=None). Now that add_deployment returns a
ReconcileOutcome, add_new_model reads .still_desired off that None and
the two PTU gate tests fail with "'NoneType' object has no attribute
'still_desired'".

Return ReconcileOutcome(still_desired=None, live_after=None), matching
the other reconcile mocks. Both fields None means no reconcile state was
captured, so the serving verdict falls back to reading the router live,
which is what the test's mock_router already drives — the PTU assertions
are unchanged.

Two sibling test files were updated for this in the parent commit; this
one was missed because the local env cannot collect four modules under
tests/test_litellm/proxy (prisma generate artifacts), so the full shard
only ran in CI.

Also applies ruff format to proxy_server.py: the new add_deployment
wrapper's single call fits on one line under the project's line length.

* fix(proxy): lock the delete evictions and stop clear_cache wiping deployments

Two follow-ups to MODEL_RECONCILE_LOCK, both found by review.

1. delete_model and delete_team_models evict from llm_router directly,
   outside the lock. The db row is gone by then, but a reconcile that
   snapshotted the db BEFORE the delete still lists that id as desired and
   upserts the deployment straight back, so the pod keeps serving a model
   the database no longer has until some later reconcile notices. Taking
   the lock orders the eviction after any in-flight reconcile's re-add.
   Both new tests fail without the lock ("did not wait for
   MODEL_RECONCILE_LOCK") and pass with it.

2. clear_cache no longer wipes deployments. It used to delete_deployment()
   every db model before the reload restored them, which left the router
   serving ZERO db models for the entire width of the reload -- every
   inference request landing in that window fell into a real hole, and
   serializing reconciles made the aggregate outage additive rather than
   overlapping. The wipe was also redundant: _delete_deployment evicts
   exactly the ids the db no longer lists, and upsert_deployment
   pops-and-re-adds a deployment whose params changed while no-opping one
   that did not, so the reconcile converges to the same state on its own.
   Every mutation is visible to that comparison (blocked, and updated_at
   for premium, are written into model_info).

   The auto-router pops are NOT redundant and stay: they are keyed by
   model_name, which no deployment-id reconcile touches.

The new tests patch their own lock rather than contending the module-level
one: asyncio.Lock binds to the event loop of its first contended acquire
and raises on every other loop after that, which would poison the next
asyncio test in the process. The proxy has a single event loop for its
lifetime so this is test-only, but it is a trap worth naming for whoever
writes the next concurrency test here.

* fix(proxy): scope the clear_cache wipe to auto-router deployments

Review caught a regression in the previous commit. Dropping the wipe
entirely stranded every db-backed auto-router on the pod.

The strategy registries (auto_routers, complexity_routers,
adaptive_routers, quality_routers) are keyed by model_name, which no
deployment-id reconcile touches, so clear_cache pops them and relies on
the reload to rebuild them. But the rebuild only happens on the ADD path:
Router.upsert_deployment returns early when a deployment is unchanged and
never reaches add_deployment -> _add_deployment ->
init_auto_router_deployment, which is what repopulates them. With the wipe
gone the deployment was always unchanged, so the pop was permanent: ANY
unrelated model write -- a team admin patching one team-owned model --
left every db-backed auto, complexity, adaptive and quality router
unroutable across tenants until a restart.

Restore the wipe for exactly the auto_router/* db deployments, whose
strategy entries are the ones being popped. Deleting them forces upsert
down the add path so both the deployment and its strategy entry come back.
Ordinary db models stay un-wiped, which is the point of the previous
commit: wiping them un-served every db model for the width of the reload,
and the reconcile converges without it.

test_clear_cache_wipes_auto_routers_but_leaves_ordinary_db_models pins
both halves against each other, since fixing either one naively breaks the
other. Both clear_cache tests fail with the pop-without-delete version.

* refactor(clear_cache): fold auto-router wipe into the classification pass

The auto-router scoping added in 5deddfd introduced two new mutable-collection
constructions, pushing LIT002 five over its budget ceiling.

Rather than suppress, do the work in the single pass that already walks
current_models: detect and delete the auto_router/* db deployments while
classifying, accumulating names into a set that replaces the old
db_router_deployments comprehension. Net-zero LIT002, same behaviour.

Comment updated to describe where the wipe actually happens now.
2026-08-12 13:42:26 -07:00
..
management_v1 test(proxy): guard management_v1 against fastapi names removed in supported releases 2026-08-08 21:15:56 -07:00
policy_endpoints style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
scim fix(scim): stop provisioning nested group ids as internal users (#34997) 2026-07-29 09:38:17 -07:00
search_endpoints test(proxy): type the search tool test helpers and record lookups with AsyncMock 2026-08-05 23:40:16 -07:00
usage_endpoints fix(proxy): reject user_id=None on non-admin analytics endpoints (cross-tenant disclosure) 2026-04-30 23:54:01 +02:00
test_access_group_endpoints.py update test cases to match new behaviour. The earlier test cases assumed the cache stores a pydantic object 2026-04-28 21:08:46 +00:00
test_access_group_management.py fix(proxy): report when a model write does not survive the post-write reload 2026-07-28 18:52:06 -07:00
test_activity_tenant_scoping.py fix(proxy): deny agent access when key and team grants resolve to nothing (#36221) 2026-08-07 20:44:11 +00:00
test_auto_router_endpoints.py feat(auto-router): track turns per complexity tier (LIT-5302) (#36209) 2026-08-07 17:03:34 -07:00
test_budget_endpoints.py fix(reset_budget_job): atomic budget cascade with chunked reset scans (#36287) 2026-08-10 14:42:36 -07:00
test_cache_settings_endpoints.py fix(ui): reflect REDIS_* env cache config and stop the UI overwriting the stored password (#34160) 2026-07-21 18:36:26 -07:00
test_callback_management_endpoints.py fix(galileo): use ingest traces API and standard logging payload (#29651) 2026-06-05 09:03:17 -07:00
test_common_daily_activity.py feat(ptu): gate PTU flat-cost attribution behind an opt-in env var (#36138) 2026-08-10 12:23:20 -07:00
test_common_utils.py fix(reset_budget_job): atomic budget cascade with chunked reset scans (#36287) 2026-08-10 14:42:36 -07:00
test_compliance_endpoints.py fix(proxy): match multi-mode guardrail_mode without false-COMPLIANT (#32832) 2026-07-10 16:41:28 -07:00
test_config_override_endpoints.py fix(audit): label vault POST as updated when DB row exists 2026-05-01 02:44:47 +00:00
test_coordination_redis_endpoints.py fix(proxy): give proxy_admin_viewer read parity with proxy_admin (#35851) 2026-08-05 18:33:55 +00:00
test_cost_tracking_settings.py fix(proxy): honor model_info custom pricing in /cost/estimate 2026-08-11 23:25:01 -07:00
test_credential_migration.py feat(proxy): add AES-256-GCM at-rest credential encryption with versioned format and re-encryption migration (#31215) 2026-06-29 20:14:22 +02:00
test_customer_budget.py feat(proxy): type Customer Management response_model for OpenAPI coverage (#31043) 2026-06-30 09:58:01 -07:00
test_customer_endpoints.py fix(reset_budget_job): atomic budget cascade with chunked reset scans (#36287) 2026-08-10 14:42:36 -07:00
test_delete_callbacks_endpoint.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_delete_verification_tokens_failed.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_encryption_endpoints.py feat(proxy): add AES-256-GCM at-rest credential encryption with versioned format and re-encryption migration (#31215) 2026-06-29 20:14:22 +02:00
test_entraid_app_roles.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_gateway_request_endpoints.py feat(sgr): make the gateway middleware the source of truth for successful requests (#35717) 2026-08-05 12:40:47 -07:00
test_internal_user_endpoints.py test: remove tests that never execute 2026-08-12 10:45:38 -07:00
test_key_management_endpoints.py test: remove tests that never execute 2026-08-12 10:45:38 -07:00
test_mcp_management_endpoints.py fix(mcp): annotate connected-app reachability on the gateway connect page (#34867) 2026-07-31 05:38:54 +00:00
test_model_management_endpoints.py fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other (#36687) 2026-08-12 13:42:26 -07:00
test_org_admin_team_access.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_organization_endpoints.py feat(organization): add RESTful PATCH /v2/organization/{organization_id} (#32350) 2026-07-23 04:53:20 +00:00
test_policy_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_project_org_authz.py fix(tests): use canonical litellm_enterprise import path (#27699) 2026-05-12 12:32:57 -07:00
test_ptu_model_settings.py fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other (#36687) 2026-08-12 13:42:26 -07:00
test_router_settings_endpoints.py feat: routing groups ui 2026-05-04 18:09:14 -07:00
test_saml_sso.py feat(proxy): add SAML 2.0 SSO for the admin UI (#31429) 2026-07-24 12:51:28 -07:00
test_tag_management_endpoints.py refactor(test): tighten typing on the tag list verification token double 2026-07-31 09:52:46 -07:00
test_team_callback_endpoints.py fix(team-callbacks): actually stop logging when disable_logging is called (#35520) 2026-08-03 11:02:51 -07:00
test_team_default_params.py feat(teams): apply default organization to new teams from default team settings (#35540) 2026-08-03 12:57:12 -07:00
test_team_endpoints.py test: remove tests that never execute 2026-08-12 10:45:38 -07:00
test_team_model_alias_merge.py chore(oss): litellm oss staging 120626 (#30292) 2026-06-12 09:49:25 -07:00
test_tool_management_endpoints.py fix(tool-management): drop unsupported prisma select kwarg from team lookup (#35293) 2026-07-31 11:42:52 -07:00
test_ui_sso.py feat(teams): apply default organization to new teams from default team settings (#35540) 2026-08-03 12:57:12 -07:00
test_workflow_management_endpoints.py fix(proxy): give proxy_admin_viewer read parity with proxy_admin (#35851) 2026-08-05 18:33:55 +00:00