mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-15 23:31:29 +00:00
test(e2e): skip the override strategy cells and describe the 1s cooldown cache
This commit is contained in:
parent
cae009c387
commit
6db93a930f
2 changed files with 29 additions and 11 deletions
|
|
@ -6,13 +6,14 @@ way (a 500, a 429, a 401, or a timeout) holding all of the group's shuffle weigh
|
|||
with an `allowed_fails_policy` of zero for that error class and a short
|
||||
`cooldown_time`, plus a healthy backup at weight 0. The first call, retries off,
|
||||
surfaces the failure to the customer as-is and benches the deployment. The proxy
|
||||
records the bench off the request path, and a sibling replica that checked Redis
|
||||
for that deployment just before the bench landed keeps sending it traffic until
|
||||
it looks again, which it does at most every 10s
|
||||
(litellm.default_redis_batch_cache_expiry). So for REPLICA_PROPAGATION_SECONDS
|
||||
after the trip every answer has to be either the deployment's own failure or a
|
||||
200 from the backup, which the proxy names in x-litellm-model-id, and at least
|
||||
one replica has to have served from the backup by then. From then until shortly
|
||||
records the bench off the request path, and a sibling replica only sees it on
|
||||
its next read of the cooldown keys from Redis, which the cooldown cache does at
|
||||
most every 1s (DEFAULT_COOLDOWN_REDIS_READ_INTERVAL_SECONDS). So for
|
||||
REPLICA_PROPAGATION_SECONDS after the trip, a window kept far wider than that
|
||||
so this cell asserts the trip and the recovery rather than how fast siblings
|
||||
catch up, every answer has to be either the deployment's own failure or a 200
|
||||
from the backup, which the proxy names in x-litellm-model-id, and at least one
|
||||
replica has to have served from the backup by then. From then until shortly
|
||||
before the cooldown can lapse, every call has to land on the backup whichever
|
||||
replica takes it. Then the test polls until the weighted shuffle opens on the
|
||||
failing deployment again and the same failure comes back (or, for the 429 pair,
|
||||
|
|
|
|||
|
|
@ -27,10 +27,8 @@ Least-busy reads live traffic, so its group of four equal deployments gets one
|
|||
long streaming request, opened under least-busy and held unread (its head names
|
||||
the deployment it landed on), and every short least-busy call sent while it is
|
||||
in flight must land on one of the other three. The stream itself goes through
|
||||
least-busy because a proxy process only starts counting in-flight requests once
|
||||
it has routed a least-busy request, which is what registers the counting
|
||||
callback, so a stream opened under another strategy would go uncounted in a
|
||||
process that has never routed one. Three idle deployments rather than one
|
||||
least-busy because the in-flight counter is the strategy's own callback, so a
|
||||
stream opened under another strategy would go uncounted. Three idle deployments rather than one
|
||||
because a process counts in its own memory, reads the shared count from Redis
|
||||
only on its first look at a group, and releases a call's count in a success
|
||||
callback that runs some time after the response leaves it, so a process can
|
||||
|
|
@ -43,6 +41,17 @@ would route on its own stale copy, in which nothing is busy. Draining the stream
|
|||
to its terminator afterwards proves the deployment holding it was healthy the
|
||||
whole time.
|
||||
|
||||
Both the latency-based and the least-busy cell are skipped until LIT-7682 lands.
|
||||
Since #40229 the per-request override builds its selector without registering
|
||||
the selector's logging hooks, so an overriding request runs neither the latency
|
||||
sampler nor the in-flight counter: latency-based picks at random with no
|
||||
samples, and least-busy picks the first deployment in its list with every count
|
||||
at zero. Neither failure is guaranteed on a given run (random picks can skip the
|
||||
slow deployment three times in a row, and which deployment a replica lists first
|
||||
depends on the order it loaded the group from the DB), so a skip is the honest
|
||||
bookkeeping this harness asks for: the two cells go back to the gap list instead
|
||||
of passing by luck, and the fix PR removes the skips as its e2e proof.
|
||||
|
||||
The per-request strategy comes in through `router_settings_override`, the same
|
||||
knob a key or team's `router_settings` feeds, so one long-lived proxy configured
|
||||
for simple-shuffle serves every strategy.
|
||||
|
|
@ -209,6 +218,10 @@ class TestReliabilityRoutingStrategies:
|
|||
)
|
||||
_assert_shuffle_control_lands_on(client, scoped_key, group, capped)
|
||||
|
||||
@pytest.mark.skip(
|
||||
reason="LIT-7682: since #40229 the per-request routing_strategy override runs without the latency sampler, "
|
||||
"so latency-based has no signal to route on"
|
||||
)
|
||||
@pytest.mark.covers("reliability.routing.latency_based.picks_lowest_latency")
|
||||
def test_latency_based_routes_around_deployment_that_times_out(
|
||||
self, client: ComplexityRouterClient, resources: ResourceManager, scoped_key: str
|
||||
|
|
@ -235,6 +248,10 @@ class TestReliabilityRoutingStrategies:
|
|||
f"{control.status_code}: it was benched, so the fast picks above prove nothing"
|
||||
)
|
||||
|
||||
@pytest.mark.skip(
|
||||
reason="LIT-7682: since #40229 the per-request routing_strategy override runs without the in-flight counter, "
|
||||
"so least-busy has no signal to route on"
|
||||
)
|
||||
@pytest.mark.covers("reliability.routing.least_busy.picks_lowest_traffic")
|
||||
def test_least_busy_avoids_deployment_with_request_in_flight(
|
||||
self, client: ComplexityRouterClient, resources: ResourceManager, scoped_key: str
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue