test(e2e): skip the override strategy cells and describe the 1s cooldown cache

This commit is contained in:
mateo-berri 2026-09-12 17:56:23 -07:00
parent cae009c387
commit 6db93a930f
2 changed files with 29 additions and 11 deletions

View file

@ -6,13 +6,14 @@ way (a 500, a 429, a 401, or a timeout) holding all of the group's shuffle weigh
with an `allowed_fails_policy` of zero for that error class and a short
`cooldown_time`, plus a healthy backup at weight 0. The first call, retries off,
surfaces the failure to the customer as-is and benches the deployment. The proxy
records the bench off the request path, and a sibling replica that checked Redis
for that deployment just before the bench landed keeps sending it traffic until
it looks again, which it does at most every 10s
(litellm.default_redis_batch_cache_expiry). So for REPLICA_PROPAGATION_SECONDS
after the trip every answer has to be either the deployment's own failure or a
200 from the backup, which the proxy names in x-litellm-model-id, and at least
one replica has to have served from the backup by then. From then until shortly
records the bench off the request path, and a sibling replica only sees it on
its next read of the cooldown keys from Redis, which the cooldown cache does at
most every 1s (DEFAULT_COOLDOWN_REDIS_READ_INTERVAL_SECONDS). So for
REPLICA_PROPAGATION_SECONDS after the trip, a window kept far wider than that
so this cell asserts the trip and the recovery rather than how fast siblings
catch up, every answer has to be either the deployment's own failure or a 200
from the backup, which the proxy names in x-litellm-model-id, and at least one
replica has to have served from the backup by then. From then until shortly
before the cooldown can lapse, every call has to land on the backup whichever
replica takes it. Then the test polls until the weighted shuffle opens on the
failing deployment again and the same failure comes back (or, for the 429 pair,

View file

@ -27,10 +27,8 @@ Least-busy reads live traffic, so its group of four equal deployments gets one
long streaming request, opened under least-busy and held unread (its head names
the deployment it landed on), and every short least-busy call sent while it is
in flight must land on one of the other three. The stream itself goes through
least-busy because a proxy process only starts counting in-flight requests once
it has routed a least-busy request, which is what registers the counting
callback, so a stream opened under another strategy would go uncounted in a
process that has never routed one. Three idle deployments rather than one
least-busy because the in-flight counter is the strategy's own callback, so a
stream opened under another strategy would go uncounted. Three idle deployments rather than one
because a process counts in its own memory, reads the shared count from Redis
only on its first look at a group, and releases a call's count in a success
callback that runs some time after the response leaves it, so a process can
@ -43,6 +41,17 @@ would route on its own stale copy, in which nothing is busy. Draining the stream
to its terminator afterwards proves the deployment holding it was healthy the
whole time.
Both the latency-based and the least-busy cell are skipped until LIT-7682 lands.
Since #40229 the per-request override builds its selector without registering
the selector's logging hooks, so an overriding request runs neither the latency
sampler nor the in-flight counter: latency-based picks at random with no
samples, and least-busy picks the first deployment in its list with every count
at zero. Neither failure is guaranteed on a given run (random picks can skip the
slow deployment three times in a row, and which deployment a replica lists first
depends on the order it loaded the group from the DB), so a skip is the honest
bookkeeping this harness asks for: the two cells go back to the gap list instead
of passing by luck, and the fix PR removes the skips as its e2e proof.
The per-request strategy comes in through `router_settings_override`, the same
knob a key or team's `router_settings` feeds, so one long-lived proxy configured
for simple-shuffle serves every strategy.
@ -209,6 +218,10 @@ class TestReliabilityRoutingStrategies:
)
_assert_shuffle_control_lands_on(client, scoped_key, group, capped)
@pytest.mark.skip(
reason="LIT-7682: since #40229 the per-request routing_strategy override runs without the latency sampler, "
"so latency-based has no signal to route on"
)
@pytest.mark.covers("reliability.routing.latency_based.picks_lowest_latency")
def test_latency_based_routes_around_deployment_that_times_out(
self, client: ComplexityRouterClient, resources: ResourceManager, scoped_key: str
@ -235,6 +248,10 @@ class TestReliabilityRoutingStrategies:
f"{control.status_code}: it was benched, so the fast picks above prove nothing"
)
@pytest.mark.skip(
reason="LIT-7682: since #40229 the per-request routing_strategy override runs without the in-flight counter, "
"so least-busy has no signal to route on"
)
@pytest.mark.covers("reliability.routing.least_busy.picks_lowest_traffic")
def test_least_busy_avoids_deployment_with_request_in_flight(
self, client: ComplexityRouterClient, resources: ResourceManager, scoped_key: str