litellm/tests/router_unit_tests/test_router_cooldown_utils.py
Deepanshu Lulla 0580465384
feat(router): add per-deployment allowed_fails_policy and cooldown_time override support (#34416)
* feat(router): add per-deployment allowed_fails_policy and cooldown_time override support

Three bugs fixed in the router cooldown system: (1) deployment-level allowed_fails and
allowed_fails_policy in model_info now take precedence over router-level settings in
_should_cooldown_deployment; (2) failed fallback deployments now get evaluated for
cooldown via _trigger_cooldown_for_failed_deployment, bypassing the Logging dedup gate;
(3) DualCache promotes Redis cooldown entries using default 600s TTL instead of true
remaining cooldown time -- _corrected_active_cooldown now evicts expired entries and
corrects stale in-memory TTLs on backfill. Adds ServiceUnavailableError, BadGatewayError,
and NotFoundError fields to AllowedFailsPolicy and cooldown_time to LiteLLMParamsTypedDict.

* fix(router): gate fallback cooldown trigger on has_logged_async_failure; use only litellm_metadata for deployment ID

* fix(router): use X | Y union syntax to fix UP007 strict lint gate

* test(router_utils): add coverage for _trigger_cooldown_for_failed_deployment and has_logged_async_failure gate

* test(router_utils): cover deployment cooldown override and exception swallow paths

* fix(router): add InternalServerError/ServiceUnavailableError/BadGatewayError/NotFoundError to router-level get_allowed_fails_from_policy

* fix(router): format router.py and add router-level policy tests

* test(router): add CI-visible coverage for per-deployment cooldown policy

Tests for `_get_deployment_cooldown_policy`, `_resolve_allowed_fails_from_policy`,
and `_should_cooldown_based_on_deployment_policy` (cooldown_handlers.py), the
`_corrected_active_cooldown` branches in CooldownCache, and the four new
exception-type branches in `Router.get_allowed_fails_from_policy` (router.py) --
all in `tests/test_litellm/` which the enterprise-routing CI job runs.

* fix(router): use is not None guard for cooldown_time_override in should_cooldown_based_on_allowed_fails_policy

A cooldown_time_override of 0 was previously treated as falsy and silently
fell through to the router-level cooldown_time value. Switched to an explicit
is not None check so that zero is honored as a valid override.

Added a regression test covering the zero case.

* fix(router): honor has_logged_async_failure and metadata for fallback cooldown; support both model_info and litellm_params locations

Manual verification against a live proxy surfaced that the fallback-cooldown-gap
trigger never actually fired: the has_logged_async_failure check read a plain
attribute that Logging never sets (the real flag lives in model_call_details),
and the deployment_id lookup only trusted litellm_metadata, which regular chat
completions never populate (only batch/thread/file endpoints do). Router
overwrites model_info on whichever key is present before every attempt, so
metadata is equally authoritative there, not caller-controlled as previously
assumed. Also let allowed_fails/allowed_fails_policy/cooldown_time be set under
either model_info or litellm_params, each preferring its own canonical location.

* fix(router): fix ContentPolicyViolationError policy shadowing and partial-policy zero-threshold

Two bugs from Greptile review on PR #34416:

- ContentPolicyViolationError subclasses BadRequestError, so listing
  BadRequestError first in _EXCEPTION_POLICY_FIELDS made the isinstance
  check always match BadRequestError for content-policy errors, using the
  wrong allowed_fails threshold. Reordered so the subclass is checked first.

- A deployment with a partial allowed_fails_policy and no deployment-wide
  allowed_fails forced allowed_fails_override=0 for any exception type its
  policy didn't cover, cooling the deployment down on the first unrelated
  failure. Now defers to router-level behavior for uncovered exception
  types instead of forcing an immediate cooldown.

* fix(router): only trust a metadata/litellm_metadata bucket the router itself wrote deployment info into

veria-ai flagged that preferring litellm_metadata whenever present could pick up a
caller-supplied litellm_metadata.model_info.id (preserved via allow_client_pricing_override)
instead of the metadata bucket the router actually populated for a regular completion's
fallback attempt, naming an arbitrary "victim" deployment for cooldown.

Router._update_kwargs_with_deployment() always writes model_info and
deployment_model_name into the same bucket together. Only trust a bucket that
carries deployment_model_name alongside model_info, since that marker is only
ever set by the router itself, not by request-body metadata.

* test(router): add regression coverage for ContentPolicyViolationError policy shadowing

The subclass-ordering fix in commit 38fe4e4490 had no regression test.
Verified the new test fails on the pre-fix ordering (asserts 2, got 10)
before restoring the fix, and confirmed the same behavior through the full
_should_cooldown_deployment call path against a real Router instance.

* fix(router): let explicit allowed_fails_policy entries override the generic 4XX cooldown exclusion

_is_cooldown_required skips cooldown evaluation for any 4XX status outside
{429, 401, 408, 404} by default, since a generic client error is usually not
the deployment's fault. BadRequestError and ContentPolicyViolationError both
carry status 400, so their AllowedFailsPolicy fields (BadRequestErrorAllowedFails,
ContentPolicyViolationErrorAllowedFails, both router-level pre-existing and the
new deployment-level ones) were silently unreachable: an operator could set
them to any value with no effect, since _is_cooldown_required blocked cooldown
evaluation before that policy was ever consulted.

_should_run_cooldown_logic now also checks whether an explicit allowed_fails_policy
entry (deployment-level or router-level) covers the exception's type, and if so,
proceeds with cooldown evaluation regardless of the generic status-code exclusion.
The exclusion remains the default for exception types with no explicit policy.

Verified live against a mock-triggered ContentPolicyViolationError (config-level
mock_response, azure/gpt-4.1-mini deployment) with BadRequestErrorAllowedFails=100
and ContentPolicyViolationErrorAllowedFails=0 on the same deployment: it now cools
down after exactly one ContentPolicyViolationError instead of never cooling down.

* fix(router): use the router-stamped failed_deployment_id for fallback cooldown targeting

Greptile flagged a real gap in the metadata-bucket-based deployment lookup:
for a generic-API-call fallback, the router writes the current attempt into
litellm_metadata, but a stale "metadata" bucket carrying the same
deployment_model_name marker (from an earlier point) would be picked first,
cooling the wrong deployment.

Router already has a more robust, pre-existing mechanism for this exact
problem: _set_failed_deployment_id_on_exception stamps the failing
deployment's id directly onto the exception at the point of failure,
immune to metadata-bucket ambiguity since a caller can't influence it and
it doesn't depend on which bucket the current call type happens to use.
It just wasn't called from _ageneric_api_call_with_fallbacks_helper's
except block, unlike _completion/_acompletion.

Added the missing call there (matching the existing pattern exactly), and
changed _trigger_cooldown_for_failed_deployment to prefer
exception.failed_deployment_id when present, falling back to metadata-bucket
inspection only for call paths that don't stamp it yet.

Verified live: the standard fallback-cooldown-gap scenario (two bad-key
deployments in a fallback chain) still correctly cools down both the
originally-called and fallback deployment.

* fix(router): address human review on per-deployment cooldown overrides

Scope allowed_fails_policy override to deployment-level only (a router-level
policy predates this feature and must keep its existing behavior), exempt
advisor-orchestration failures from the fallback cooldown trigger, keep the
single-deployment model group protection intact against a generic
deployment-level allowed_fails, make cooldown_time precedence consistent
across resolution paths, fix a falsy-zero swallowing bug in the router-level
allowed_fails fallback, and make allowed_fails_policy resolution fall through
to the next matching exception type instead of stopping at the first unset
field.

Also restrict allowed_fails/allowed_fails_policy/cooldown_time to model_info:
litellm_params gets copied into the actual provider request, so a router-only
setting placed there would leak into that request.

* test(router): update test_cooldown_handlers.py for the deployment-policy signature change

Surfaced by the rebase: this mirrored test file (tests/test_litellm/ mirrors
litellm/) predates the router_unit_tests/ coverage added earlier in this PR and
was still calling _should_cooldown_based_on_deployment_policy with its old
4-argument signature and asserting the now-removed litellm_params cooldown_time
location.

* test(router): update test_fallback_event_handlers.py for model_info-only cooldown_time

Another mirrored test file surfaced by the rebase that still asserted the
now-removed litellm_params.cooldown_time location.

* fix(router): match cooldown-duration precedence in the fallback path to the primary path

_trigger_cooldown_for_failed_deployment only checked deployment config before
falling back to the router default, skipping the response Retry-After header
step that Router.deployment_callback_on_failure applies on the primary path.

* fix(router): restore litellm_params.cooldown_time as a pre-existing fallback

cooldown_time already had litellm_params support on Router.deployment_callback_on_failure
before this PR; the earlier model_info-only restriction (aimed at the leak concern
for the genuinely new allowed_fails/allowed_fails_policy fields) incorrectly dropped
that pre-existing capability too. model_info still takes priority when both are set.

* fix(router): keep the fallback-cooldown trigger in sync with #35104's review fixes

Applies the same two fixes landed on the split-out PR #35104 (which #34416
still duplicates until it's rebased onto the merged base): increment the
deployment's per-minute failure counter before evaluating cooldown, and
require the server-stamped failed_deployment_id instead of trusting a
metadata bucket, since neither "metadata" nor "litellm_metadata" can be told
apart from a caller-supplied one without knowing the call's function_name.

* fix(router): freeze the model_info fallback mapping to satisfy the type-discipline gate

* fix(router): defer f-string interpolation in fallback-cooldown debug logs

* fix(router): annotate cooldown-path locals with Final to satisfy the LIT010 budget

* fix(router): suppress reportPrivateUsage for cross-module cooldown helpers

* fix(router): don't cool down deployments for request-scoped 404s on generic API fallbacks

* fix(router): stamp the dynamic client-side-credential deployment id, not the shared static one

* fix(router): keep up with upstream typing modernization and Final-annotation ratchet

* fix(router): don't cool down deployments for a caller-supplied x-litellm-timeout

* fix(router): stamp dynamic client-side-credential id in completion fallback paths too

The generic-API-call helper already stamped the effective (dynamic-if-client-side-credential)
deployment id on exceptions, but the regular _completion/_acompletion exception handlers still
stamped the static shared deployment's id. A tenant using invalid forwarded credentials could
generate repeated failures attributed to, and eventually cooling down, the shared deployment
other tenants rely on. Extracted the stamping logic into one shared helper used by all three
call sites (generic API, sync completion, async completion) so the fix and future changes to it
stay in one place.

* fix(proxy): recognize body-supplied timeout/request_timeout/stream_timeout as caller-controlled

client_side_timeout was only set when the caller used the x-litellm-timeout header, but
Router._get_timeout also resolves the effective timeout from kwargs["timeout"],
kwargs["request_timeout"], and kwargs["stream_timeout"], all settable directly in the
request body (and x-litellm-stream-timeout wasn't marked either). A caller could set any
of those to a near-zero value, force a 408 on every deployment in a fallback chain, and
cool down deployments other tenants rely on without the guard in
_trigger_cooldown_for_failed_deployment recognizing it as caller-controlled. Also strip
any client-forged client_side_timeout from the request body so the marker is always
server-computed.

---------

Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-10 11:02:06 -07:00

604 lines
20 KiB
Python

import sys, os, time
import traceback, asyncio
import pytest
sys.path.insert(
0, os.path.abspath("../..")
) # Adds the parent directory to the system path
import litellm
from litellm import Router
from litellm.router import Deployment, LiteLLM_Params
from litellm.types.router import ModelInfo
from concurrent.futures import ThreadPoolExecutor
from collections import defaultdict
from dotenv import load_dotenv
from unittest.mock import AsyncMock, MagicMock, patch
from litellm.router_utils.cooldown_callbacks import router_cooldown_event_callback
from litellm.router_utils.cooldown_handlers import (
_should_run_cooldown_logic,
_should_cooldown_deployment,
cast_exception_status_to_int,
_is_cooldown_required,
_has_explicit_allowed_fails_policy_for_exception,
)
from litellm.types.router import AllowedFailsPolicy
from litellm.router_utils.router_callbacks.track_deployment_metrics import (
increment_deployment_failures_for_current_minute,
increment_deployment_successes_for_current_minute,
)
import pytest
from unittest.mock import patch
from litellm import Router
from litellm.router_utils.cooldown_handlers import _should_cooldown_deployment
load_dotenv()
@pytest.mark.asyncio
async def test_router_cooldown_event_callback_no_deployment():
"""
Test the router_cooldown_event_callback function
Ensures that the router_cooldown_event_callback function does not raise an error when no deployment is found
In this scenario it should do nothing
"""
# Mock Router instance
mock_router = MagicMock()
mock_router.get_deployment.return_value = None
await router_cooldown_event_callback(
litellm_router_instance=mock_router,
deployment_id="test-deployment",
exception_status="429",
cooldown_time=60.0,
)
# Assert that the router's get_deployment method was called
mock_router.get_deployment.assert_called_once_with(model_id="test-deployment")
@pytest.fixture
def testing_litellm_router():
return Router(
model_list=[
{
"model_name": "gpt-5-mini",
"litellm_params": {"model": "gpt-5-mini"},
"model_id": "test_deployment",
},
{
"model_name": "test_deployment",
"litellm_params": {"model": "openai/test_deployment"},
"model_id": "test_deployment_2",
},
{
"model_name": "test_deployment",
"litellm_params": {"model": "openai/test_deployment-2"},
"model_id": "test_deployment_3",
},
]
)
def test_should_run_cooldown_logic(testing_litellm_router):
testing_litellm_router.disable_cooldowns = True
# don't run cooldown logic if disable_cooldowns is True
assert (
_should_run_cooldown_logic(
testing_litellm_router, "test_deployment", 500, Exception("Test")
)
is False
)
# don't cooldown if deployment is None
testing_litellm_router.disable_cooldowns = False
assert (
_should_run_cooldown_logic(testing_litellm_router, None, 500, Exception("Test"))
is False
)
# don't cooldown if it's a provider default deployment
testing_litellm_router.provider_default_deployment_ids = ["test_deployment"]
assert (
_should_run_cooldown_logic(
testing_litellm_router, "test_deployment", 500, Exception("Test")
)
is False
)
@pytest.fixture
def single_deployment_router():
"""A router with one deployment whose model_info.id is the lookup-able
"dep-1" (unlike `testing_litellm_router`'s top-level "model_id" key, which
is not absorbed into model_info.id and so never resolves via
get_model_info/get_model_group)."""
return Router(
model_list=[
{
"model_name": "gpt-5-mini",
"litellm_params": {"model": "gpt-5-mini"},
"model_info": {"id": "dep-1"},
},
]
)
def test_should_run_cooldown_logic_generic_bad_request_excluded_by_default(
single_deployment_router,
):
"""A generic BadRequestError/ContentPolicyViolationError (400) is excluded from
cooldown evaluation by _is_cooldown_required when no allowed_fails_policy is
configured for that exception type. This is the pre-existing, intentional
default: a client error is usually not the deployment's fault."""
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
assert (
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is False
)
def test_should_run_cooldown_logic_router_level_policy_does_not_override_bad_request_exclusion(
single_deployment_router,
):
"""A router-level allowed_fails_policy is a pre-existing, router-wide setting that
predates the per-deployment override feature, so it must keep its existing behavior
and stay subject to the generic 4XX exclusion. Only an explicit deployment-level
policy (an unambiguous per-exception opt-in for that one deployment) overrides it;
see test_should_run_cooldown_logic_explicit_deployment_level_policy_overrides_content_policy_exclusion."""
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
BadRequestErrorAllowedFails=5
)
assert (
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is False
)
def test_should_run_cooldown_logic_explicit_deployment_level_policy_overrides_content_policy_exclusion(
single_deployment_router,
):
"""Same as the router-level case, but for a deployment-level allowed_fails_policy
entry (this PR's per-deployment feature) targeting ContentPolicyViolationError."""
exc = litellm.ContentPolicyViolationError("flagged content", "openai", "gpt-5-mini")
deployment_dict = single_deployment_router.get_model_info(id="dep-1")
deployment_dict["model_info"]["allowed_fails_policy"] = {
"ContentPolicyViolationErrorAllowedFails": 0
}
assert (
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is True
)
class TestHasExplicitAllowedFailsPolicyForException:
def test_no_policy_anywhere_returns_false(self, single_deployment_router):
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
assert (
_has_explicit_allowed_fails_policy_for_exception(
single_deployment_router, "dep-1", exc
)
is False
)
def test_router_level_policy_for_matching_exception_returns_false(
self, single_deployment_router
):
"""Deliberately scoped to deployment-level only: a router-level policy
predates this feature and must not be treated as an explicit per-exception
opt-in for cooldown-gate purposes."""
exc = litellm.RateLimitError("rate limited", "openai", "gpt-5-mini")
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
RateLimitErrorAllowedFails=3
)
assert (
_has_explicit_allowed_fails_policy_for_exception(
single_deployment_router, "dep-1", exc
)
is False
)
def test_router_level_policy_for_different_exception_returns_false(
self, single_deployment_router
):
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
RateLimitErrorAllowedFails=3
)
assert (
_has_explicit_allowed_fails_policy_for_exception(
single_deployment_router, "dep-1", exc
)
is False
)
def test_deployment_level_policy_for_matching_exception_returns_true(
self, single_deployment_router
):
exc = litellm.ContentPolicyViolationError("flagged", "openai", "gpt-5-mini")
deployment_dict = single_deployment_router.get_model_info(id="dep-1")
deployment_dict["model_info"]["allowed_fails_policy"] = {
"ContentPolicyViolationErrorAllowedFails": 0
}
assert (
_has_explicit_allowed_fails_policy_for_exception(
single_deployment_router, "dep-1", exc
)
is True
)
def test_none_deployment_returns_false(self, single_deployment_router):
exc = litellm.RateLimitError("rate limited", "openai", "gpt-5-mini")
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
RateLimitErrorAllowedFails=3
)
assert (
_has_explicit_allowed_fails_policy_for_exception(
single_deployment_router, None, exc
)
is False
)
def test_should_cooldown_deployment_rate_limit_error(testing_litellm_router):
"""
Test the _should_cooldown_deployment function when a rate limit error occurs
"""
# Test 429 error (rate limit) -> always cooldown a deployment returning 429s
_exception = litellm.exceptions.RateLimitError(
"Rate limit", "openai", "gpt-5-mini"
)
assert (
_should_cooldown_deployment(
testing_litellm_router, "test_deployment", 429, _exception
)
is True
)
def test_should_cooldown_deployment_auth_limit_error(testing_litellm_router):
"""
Test the _should_cooldown_deployment function when an auth limit error occurs
"""
# Test 401 error (auth limit) -> always cooldown a deployment returning 401s
_exception = litellm.exceptions.AuthenticationError(
"Unauthorized", "openai", "gpt-5-mini"
)
assert (
_should_cooldown_deployment(
testing_litellm_router, "test_deployment", 401, _exception
)
is True
)
@pytest.mark.asyncio
async def test_should_cooldown_deployment(testing_litellm_router):
"""
Cooldown a deployment if it fails 60% of requests in 1 minute - DEFAULT threshold is 50%
"""
from litellm._logging import verbose_router_logger
import logging
verbose_router_logger.setLevel(logging.DEBUG)
# Test 429 error (rate limit) -> always cooldown a deployment returning 429s
_exception = litellm.exceptions.RateLimitError(
"Rate limit", "openai", "gpt-5-mini"
)
assert (
_should_cooldown_deployment(
testing_litellm_router, "test_deployment", 429, _exception
)
is True
)
available_deployment = testing_litellm_router.get_available_deployment(
model="test_deployment"
)
print("available_deployment", available_deployment)
assert available_deployment is not None
deployment_id = available_deployment["model_info"]["id"]
print("deployment_id", deployment_id)
# set current success for deployment to 40
for _ in range(40):
increment_deployment_successes_for_current_minute(
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
)
# now we fail 40 requests in a row
tasks = []
for _ in range(41):
tasks.append(
testing_litellm_router.acompletion(
model=deployment_id,
messages=[{"role": "user", "content": "Hello, world!"}],
max_tokens=100,
mock_response="litellm.InternalServerError",
)
)
try:
await asyncio.gather(*tasks)
except Exception:
pass
await asyncio.sleep(1)
# expect this to fail since it's now 51% of requests are failing
assert (
_should_cooldown_deployment(
testing_litellm_router, deployment_id, 500, Exception("Test")
)
is True
)
@pytest.mark.asyncio
async def test_should_cooldown_deployment_allowed_fails_set_on_router():
"""
Test the _should_cooldown_deployment function when Router.allowed_fails is set
"""
# Create a Router instance with a test deployment
router = Router(
model_list=[
{
"model_name": "gpt-5-mini",
"litellm_params": {"model": "gpt-5-mini"},
"model_id": "test_deployment",
},
]
)
# Set up allowed_fails for the test deployment
router.allowed_fails = 100
# should not cooldown when fails are below the allowed limit
for _ in range(100):
assert (
_should_cooldown_deployment(
router, "test_deployment", 500, Exception("Test")
)
is False
)
assert (
_should_cooldown_deployment(router, "test_deployment", 500, Exception("Test"))
is True
)
def test_increment_deployment_successes_for_current_minute_does_not_write_to_redis(
testing_litellm_router,
):
"""
Ensure tracking deployment metrics does not write to redis
Important - If it writes to redis on every request it will seriously impact performance / latency
"""
from litellm.caching.dual_cache import DualCache
from litellm.caching.redis_cache import RedisCache
from litellm.caching.in_memory_cache import InMemoryCache
from litellm.router_utils.router_callbacks.track_deployment_metrics import (
increment_deployment_successes_for_current_minute,
)
# Mock RedisCache
mock_redis_cache = MagicMock(spec=RedisCache)
testing_litellm_router.cache = DualCache(
redis_cache=mock_redis_cache, in_memory_cache=InMemoryCache()
)
# Call the function we're testing
increment_deployment_successes_for_current_minute(
litellm_router_instance=testing_litellm_router, deployment_id="test_deployment"
)
increment_deployment_failures_for_current_minute(
litellm_router_instance=testing_litellm_router, deployment_id="test_deployment"
)
time.sleep(1)
# Assert that no methods were called on the mock_redis_cache
assert not mock_redis_cache.method_calls, "RedisCache methods should not be called"
print(
"in memory cache values=",
testing_litellm_router.cache.in_memory_cache.cache_dict,
)
assert (
testing_litellm_router.cache.in_memory_cache.get_cache(
"test_deployment:successes"
)
is not None
)
def test_cast_exception_status_to_int():
assert cast_exception_status_to_int(200) == 200
assert cast_exception_status_to_int("404") == 404
assert cast_exception_status_to_int("invalid") == 500
@pytest.fixture
def router():
return Router(
model_list=[
{
"model_name": "gpt-5.5",
"litellm_params": {"model": "gpt-5.5"},
"model_info": {
"id": "gpt-4--0",
},
}
]
)
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
)
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
)
def test_should_cooldown_high_traffic_all_fails(mock_failures, mock_successes, router):
# Simulate 10 failures, 0 successes
from litellm.constants import SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD
mock_failures.return_value = SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD + 1
mock_successes.return_value = 0
should_cooldown = _should_cooldown_deployment(
litellm_router_instance=router,
deployment="gpt-4--0",
exception_status=500,
original_exception=Exception("Test error"),
)
assert (
should_cooldown is True
), "Should cooldown when all requests fail with sufficient traffic"
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
)
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
)
def test_no_cooldown_low_traffic(mock_failures, mock_successes, router):
# Simulate 3 failures (below MIN_TRAFFIC_THRESHOLD)
mock_failures.return_value = 3
mock_successes.return_value = 0
should_cooldown = _should_cooldown_deployment(
litellm_router_instance=router,
deployment="gpt-4--0",
exception_status=500,
original_exception=Exception("Test error"),
)
assert (
should_cooldown is False
), "Should not cooldown when traffic is below threshold"
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
)
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
)
def test_cooldown_rate_limit(mock_failures, mock_successes, router):
"""
Don't cooldown single deployment models, for anything besides traffic
"""
mock_failures.return_value = 1
mock_successes.return_value = 0
should_cooldown = _should_cooldown_deployment(
litellm_router_instance=router,
deployment="gpt-4--0",
exception_status=429, # Rate limit error
original_exception=Exception("Rate limit exceeded"),
)
assert (
should_cooldown is False
), "Should not cooldown on rate limit error for single deployment models"
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
)
@patch(
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
)
def test_mixed_success_failure(mock_failures, mock_successes, router):
# Simulate 3 failures, 7 successes
mock_failures.return_value = 3
mock_successes.return_value = 7
should_cooldown = _should_cooldown_deployment(
litellm_router_instance=router,
deployment="gpt-4--0",
exception_status=500,
original_exception=Exception("Test error"),
)
assert (
should_cooldown is False
), "Should not cooldown when failure rate is below threshold"
def test_is_cooldown_required_empty_string_exception_status(testing_litellm_router):
"""
Test that _is_cooldown_required returns False when exception_status is an empty string
"""
result = _is_cooldown_required(
litellm_router_instance=testing_litellm_router,
model_id="test_deployment",
exception_status="",
)
assert (
result is False
), "Should not require cooldown when exception_status is empty string"
def test_should_cooldown_deployment_minimum_request_threshold(testing_litellm_router):
"""
Test that error rate cooldown does NOT trigger on first failure.
Fixes GitHub issue #17418: Error Rate Cooldown Triggers on First Failed Request
The problem: With DEFAULT_FAILURE_THRESHOLD_PERCENT=0.5 (50%), a deployment
gets cooled down after just 1 failed request because 1/1 = 100% > 50%.
The fix: Add a minimum request threshold (DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS)
before applying error rate cooldown.
"""
from litellm.constants import DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS
# Get a deployment that's not a single-deployment model group
# (test_deployment_2 and test_deployment_3 are both for "test_deployment" model)
available_deployment = testing_litellm_router.get_available_deployment(
model="test_deployment"
)
assert available_deployment is not None
deployment_id = available_deployment["model_info"]["id"]
# Simulate only 1 failure (below minimum threshold)
# This should NOT trigger cooldown even though 100% > 50%
increment_deployment_failures_for_current_minute(
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
)
_exception = litellm.exceptions.InternalServerError(
"Internal error", "openai", "gpt-5-mini"
)
# With only 1 request, should NOT cooldown (below minimum threshold)
should_cooldown = _should_cooldown_deployment(
testing_litellm_router, deployment_id, 500, _exception
)
assert (
should_cooldown is False
), f"Should NOT cooldown with only 1 failed request (below minimum threshold of {DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS})"
# Now add more failures to reach the minimum threshold
for _ in range(DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS - 1):
increment_deployment_failures_for_current_minute(
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
)
# Now with enough requests (all failures), it SHOULD trigger cooldown
should_cooldown = _should_cooldown_deployment(
testing_litellm_router, deployment_id, 500, _exception
)
assert (
should_cooldown is True
), f"Should cooldown when we have {DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS} failed requests (100% failure rate)"