mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
* feat(router): add per-deployment allowed_fails_policy and cooldown_time override support Three bugs fixed in the router cooldown system: (1) deployment-level allowed_fails and allowed_fails_policy in model_info now take precedence over router-level settings in _should_cooldown_deployment; (2) failed fallback deployments now get evaluated for cooldown via _trigger_cooldown_for_failed_deployment, bypassing the Logging dedup gate; (3) DualCache promotes Redis cooldown entries using default 600s TTL instead of true remaining cooldown time -- _corrected_active_cooldown now evicts expired entries and corrects stale in-memory TTLs on backfill. Adds ServiceUnavailableError, BadGatewayError, and NotFoundError fields to AllowedFailsPolicy and cooldown_time to LiteLLMParamsTypedDict. * fix(router): gate fallback cooldown trigger on has_logged_async_failure; use only litellm_metadata for deployment ID * fix(router): use X | Y union syntax to fix UP007 strict lint gate * test(router_utils): add coverage for _trigger_cooldown_for_failed_deployment and has_logged_async_failure gate * test(router_utils): cover deployment cooldown override and exception swallow paths * fix(router): add InternalServerError/ServiceUnavailableError/BadGatewayError/NotFoundError to router-level get_allowed_fails_from_policy * fix(router): format router.py and add router-level policy tests * test(router): add CI-visible coverage for per-deployment cooldown policy Tests for `_get_deployment_cooldown_policy`, `_resolve_allowed_fails_from_policy`, and `_should_cooldown_based_on_deployment_policy` (cooldown_handlers.py), the `_corrected_active_cooldown` branches in CooldownCache, and the four new exception-type branches in `Router.get_allowed_fails_from_policy` (router.py) -- all in `tests/test_litellm/` which the enterprise-routing CI job runs. * fix(router): use is not None guard for cooldown_time_override in should_cooldown_based_on_allowed_fails_policy A cooldown_time_override of 0 was previously treated as falsy and silently fell through to the router-level cooldown_time value. Switched to an explicit is not None check so that zero is honored as a valid override. Added a regression test covering the zero case. * fix(router): honor has_logged_async_failure and metadata for fallback cooldown; support both model_info and litellm_params locations Manual verification against a live proxy surfaced that the fallback-cooldown-gap trigger never actually fired: the has_logged_async_failure check read a plain attribute that Logging never sets (the real flag lives in model_call_details), and the deployment_id lookup only trusted litellm_metadata, which regular chat completions never populate (only batch/thread/file endpoints do). Router overwrites model_info on whichever key is present before every attempt, so metadata is equally authoritative there, not caller-controlled as previously assumed. Also let allowed_fails/allowed_fails_policy/cooldown_time be set under either model_info or litellm_params, each preferring its own canonical location. * fix(router): fix ContentPolicyViolationError policy shadowing and partial-policy zero-threshold Two bugs from Greptile review on PR #34416: - ContentPolicyViolationError subclasses BadRequestError, so listing BadRequestError first in _EXCEPTION_POLICY_FIELDS made the isinstance check always match BadRequestError for content-policy errors, using the wrong allowed_fails threshold. Reordered so the subclass is checked first. - A deployment with a partial allowed_fails_policy and no deployment-wide allowed_fails forced allowed_fails_override=0 for any exception type its policy didn't cover, cooling the deployment down on the first unrelated failure. Now defers to router-level behavior for uncovered exception types instead of forcing an immediate cooldown. * fix(router): only trust a metadata/litellm_metadata bucket the router itself wrote deployment info into veria-ai flagged that preferring litellm_metadata whenever present could pick up a caller-supplied litellm_metadata.model_info.id (preserved via allow_client_pricing_override) instead of the metadata bucket the router actually populated for a regular completion's fallback attempt, naming an arbitrary "victim" deployment for cooldown. Router._update_kwargs_with_deployment() always writes model_info and deployment_model_name into the same bucket together. Only trust a bucket that carries deployment_model_name alongside model_info, since that marker is only ever set by the router itself, not by request-body metadata. * test(router): add regression coverage for ContentPolicyViolationError policy shadowing The subclass-ordering fix in commit 38fe4e4490 had no regression test. Verified the new test fails on the pre-fix ordering (asserts 2, got 10) before restoring the fix, and confirmed the same behavior through the full _should_cooldown_deployment call path against a real Router instance. * fix(router): let explicit allowed_fails_policy entries override the generic 4XX cooldown exclusion _is_cooldown_required skips cooldown evaluation for any 4XX status outside {429, 401, 408, 404} by default, since a generic client error is usually not the deployment's fault. BadRequestError and ContentPolicyViolationError both carry status 400, so their AllowedFailsPolicy fields (BadRequestErrorAllowedFails, ContentPolicyViolationErrorAllowedFails, both router-level pre-existing and the new deployment-level ones) were silently unreachable: an operator could set them to any value with no effect, since _is_cooldown_required blocked cooldown evaluation before that policy was ever consulted. _should_run_cooldown_logic now also checks whether an explicit allowed_fails_policy entry (deployment-level or router-level) covers the exception's type, and if so, proceeds with cooldown evaluation regardless of the generic status-code exclusion. The exclusion remains the default for exception types with no explicit policy. Verified live against a mock-triggered ContentPolicyViolationError (config-level mock_response, azure/gpt-4.1-mini deployment) with BadRequestErrorAllowedFails=100 and ContentPolicyViolationErrorAllowedFails=0 on the same deployment: it now cools down after exactly one ContentPolicyViolationError instead of never cooling down. * fix(router): use the router-stamped failed_deployment_id for fallback cooldown targeting Greptile flagged a real gap in the metadata-bucket-based deployment lookup: for a generic-API-call fallback, the router writes the current attempt into litellm_metadata, but a stale "metadata" bucket carrying the same deployment_model_name marker (from an earlier point) would be picked first, cooling the wrong deployment. Router already has a more robust, pre-existing mechanism for this exact problem: _set_failed_deployment_id_on_exception stamps the failing deployment's id directly onto the exception at the point of failure, immune to metadata-bucket ambiguity since a caller can't influence it and it doesn't depend on which bucket the current call type happens to use. It just wasn't called from _ageneric_api_call_with_fallbacks_helper's except block, unlike _completion/_acompletion. Added the missing call there (matching the existing pattern exactly), and changed _trigger_cooldown_for_failed_deployment to prefer exception.failed_deployment_id when present, falling back to metadata-bucket inspection only for call paths that don't stamp it yet. Verified live: the standard fallback-cooldown-gap scenario (two bad-key deployments in a fallback chain) still correctly cools down both the originally-called and fallback deployment. * fix(router): address human review on per-deployment cooldown overrides Scope allowed_fails_policy override to deployment-level only (a router-level policy predates this feature and must keep its existing behavior), exempt advisor-orchestration failures from the fallback cooldown trigger, keep the single-deployment model group protection intact against a generic deployment-level allowed_fails, make cooldown_time precedence consistent across resolution paths, fix a falsy-zero swallowing bug in the router-level allowed_fails fallback, and make allowed_fails_policy resolution fall through to the next matching exception type instead of stopping at the first unset field. Also restrict allowed_fails/allowed_fails_policy/cooldown_time to model_info: litellm_params gets copied into the actual provider request, so a router-only setting placed there would leak into that request. * test(router): update test_cooldown_handlers.py for the deployment-policy signature change Surfaced by the rebase: this mirrored test file (tests/test_litellm/ mirrors litellm/) predates the router_unit_tests/ coverage added earlier in this PR and was still calling _should_cooldown_based_on_deployment_policy with its old 4-argument signature and asserting the now-removed litellm_params cooldown_time location. * test(router): update test_fallback_event_handlers.py for model_info-only cooldown_time Another mirrored test file surfaced by the rebase that still asserted the now-removed litellm_params.cooldown_time location. * fix(router): match cooldown-duration precedence in the fallback path to the primary path _trigger_cooldown_for_failed_deployment only checked deployment config before falling back to the router default, skipping the response Retry-After header step that Router.deployment_callback_on_failure applies on the primary path. * fix(router): restore litellm_params.cooldown_time as a pre-existing fallback cooldown_time already had litellm_params support on Router.deployment_callback_on_failure before this PR; the earlier model_info-only restriction (aimed at the leak concern for the genuinely new allowed_fails/allowed_fails_policy fields) incorrectly dropped that pre-existing capability too. model_info still takes priority when both are set. * fix(router): keep the fallback-cooldown trigger in sync with #35104's review fixes Applies the same two fixes landed on the split-out PR #35104 (which #34416 still duplicates until it's rebased onto the merged base): increment the deployment's per-minute failure counter before evaluating cooldown, and require the server-stamped failed_deployment_id instead of trusting a metadata bucket, since neither "metadata" nor "litellm_metadata" can be told apart from a caller-supplied one without knowing the call's function_name. * fix(router): freeze the model_info fallback mapping to satisfy the type-discipline gate * fix(router): defer f-string interpolation in fallback-cooldown debug logs * fix(router): annotate cooldown-path locals with Final to satisfy the LIT010 budget * fix(router): suppress reportPrivateUsage for cross-module cooldown helpers * fix(router): don't cool down deployments for request-scoped 404s on generic API fallbacks * fix(router): stamp the dynamic client-side-credential deployment id, not the shared static one * fix(router): keep up with upstream typing modernization and Final-annotation ratchet * fix(router): don't cool down deployments for a caller-supplied x-litellm-timeout * fix(router): stamp dynamic client-side-credential id in completion fallback paths too The generic-API-call helper already stamped the effective (dynamic-if-client-side-credential) deployment id on exceptions, but the regular _completion/_acompletion exception handlers still stamped the static shared deployment's id. A tenant using invalid forwarded credentials could generate repeated failures attributed to, and eventually cooling down, the shared deployment other tenants rely on. Extracted the stamping logic into one shared helper used by all three call sites (generic API, sync completion, async completion) so the fix and future changes to it stay in one place. * fix(proxy): recognize body-supplied timeout/request_timeout/stream_timeout as caller-controlled client_side_timeout was only set when the caller used the x-litellm-timeout header, but Router._get_timeout also resolves the effective timeout from kwargs["timeout"], kwargs["request_timeout"], and kwargs["stream_timeout"], all settable directly in the request body (and x-litellm-stream-timeout wasn't marked either). A caller could set any of those to a near-zero value, force a 408 on every deployment in a fallback chain, and cool down deployments other tenants rely on without the guard in _trigger_cooldown_for_failed_deployment recognizing it as caller-controlled. Also strip any client-forged client_side_timeout from the request body so the marker is always server-computed. --------- Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
604 lines
20 KiB
Python
604 lines
20 KiB
Python
import sys, os, time
|
|
import traceback, asyncio
|
|
import pytest
|
|
|
|
sys.path.insert(
|
|
0, os.path.abspath("../..")
|
|
) # Adds the parent directory to the system path
|
|
import litellm
|
|
from litellm import Router
|
|
from litellm.router import Deployment, LiteLLM_Params
|
|
from litellm.types.router import ModelInfo
|
|
from concurrent.futures import ThreadPoolExecutor
|
|
from collections import defaultdict
|
|
from dotenv import load_dotenv
|
|
from unittest.mock import AsyncMock, MagicMock, patch
|
|
from litellm.router_utils.cooldown_callbacks import router_cooldown_event_callback
|
|
from litellm.router_utils.cooldown_handlers import (
|
|
_should_run_cooldown_logic,
|
|
_should_cooldown_deployment,
|
|
cast_exception_status_to_int,
|
|
_is_cooldown_required,
|
|
_has_explicit_allowed_fails_policy_for_exception,
|
|
)
|
|
from litellm.types.router import AllowedFailsPolicy
|
|
from litellm.router_utils.router_callbacks.track_deployment_metrics import (
|
|
increment_deployment_failures_for_current_minute,
|
|
increment_deployment_successes_for_current_minute,
|
|
)
|
|
|
|
import pytest
|
|
from unittest.mock import patch
|
|
from litellm import Router
|
|
from litellm.router_utils.cooldown_handlers import _should_cooldown_deployment
|
|
|
|
load_dotenv()
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_cooldown_event_callback_no_deployment():
|
|
"""
|
|
Test the router_cooldown_event_callback function
|
|
|
|
Ensures that the router_cooldown_event_callback function does not raise an error when no deployment is found
|
|
|
|
In this scenario it should do nothing
|
|
"""
|
|
# Mock Router instance
|
|
mock_router = MagicMock()
|
|
mock_router.get_deployment.return_value = None
|
|
|
|
await router_cooldown_event_callback(
|
|
litellm_router_instance=mock_router,
|
|
deployment_id="test-deployment",
|
|
exception_status="429",
|
|
cooldown_time=60.0,
|
|
)
|
|
|
|
# Assert that the router's get_deployment method was called
|
|
mock_router.get_deployment.assert_called_once_with(model_id="test-deployment")
|
|
|
|
|
|
@pytest.fixture
|
|
def testing_litellm_router():
|
|
return Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {"model": "gpt-5-mini"},
|
|
"model_id": "test_deployment",
|
|
},
|
|
{
|
|
"model_name": "test_deployment",
|
|
"litellm_params": {"model": "openai/test_deployment"},
|
|
"model_id": "test_deployment_2",
|
|
},
|
|
{
|
|
"model_name": "test_deployment",
|
|
"litellm_params": {"model": "openai/test_deployment-2"},
|
|
"model_id": "test_deployment_3",
|
|
},
|
|
]
|
|
)
|
|
|
|
|
|
def test_should_run_cooldown_logic(testing_litellm_router):
|
|
testing_litellm_router.disable_cooldowns = True
|
|
# don't run cooldown logic if disable_cooldowns is True
|
|
assert (
|
|
_should_run_cooldown_logic(
|
|
testing_litellm_router, "test_deployment", 500, Exception("Test")
|
|
)
|
|
is False
|
|
)
|
|
|
|
# don't cooldown if deployment is None
|
|
testing_litellm_router.disable_cooldowns = False
|
|
assert (
|
|
_should_run_cooldown_logic(testing_litellm_router, None, 500, Exception("Test"))
|
|
is False
|
|
)
|
|
|
|
# don't cooldown if it's a provider default deployment
|
|
testing_litellm_router.provider_default_deployment_ids = ["test_deployment"]
|
|
assert (
|
|
_should_run_cooldown_logic(
|
|
testing_litellm_router, "test_deployment", 500, Exception("Test")
|
|
)
|
|
is False
|
|
)
|
|
|
|
|
|
@pytest.fixture
|
|
def single_deployment_router():
|
|
"""A router with one deployment whose model_info.id is the lookup-able
|
|
"dep-1" (unlike `testing_litellm_router`'s top-level "model_id" key, which
|
|
is not absorbed into model_info.id and so never resolves via
|
|
get_model_info/get_model_group)."""
|
|
return Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {"model": "gpt-5-mini"},
|
|
"model_info": {"id": "dep-1"},
|
|
},
|
|
]
|
|
)
|
|
|
|
|
|
def test_should_run_cooldown_logic_generic_bad_request_excluded_by_default(
|
|
single_deployment_router,
|
|
):
|
|
"""A generic BadRequestError/ContentPolicyViolationError (400) is excluded from
|
|
cooldown evaluation by _is_cooldown_required when no allowed_fails_policy is
|
|
configured for that exception type. This is the pre-existing, intentional
|
|
default: a client error is usually not the deployment's fault."""
|
|
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
|
|
assert (
|
|
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is False
|
|
)
|
|
|
|
|
|
def test_should_run_cooldown_logic_router_level_policy_does_not_override_bad_request_exclusion(
|
|
single_deployment_router,
|
|
):
|
|
"""A router-level allowed_fails_policy is a pre-existing, router-wide setting that
|
|
predates the per-deployment override feature, so it must keep its existing behavior
|
|
and stay subject to the generic 4XX exclusion. Only an explicit deployment-level
|
|
policy (an unambiguous per-exception opt-in for that one deployment) overrides it;
|
|
see test_should_run_cooldown_logic_explicit_deployment_level_policy_overrides_content_policy_exclusion."""
|
|
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
|
|
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
|
|
BadRequestErrorAllowedFails=5
|
|
)
|
|
assert (
|
|
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is False
|
|
)
|
|
|
|
|
|
def test_should_run_cooldown_logic_explicit_deployment_level_policy_overrides_content_policy_exclusion(
|
|
single_deployment_router,
|
|
):
|
|
"""Same as the router-level case, but for a deployment-level allowed_fails_policy
|
|
entry (this PR's per-deployment feature) targeting ContentPolicyViolationError."""
|
|
exc = litellm.ContentPolicyViolationError("flagged content", "openai", "gpt-5-mini")
|
|
deployment_dict = single_deployment_router.get_model_info(id="dep-1")
|
|
deployment_dict["model_info"]["allowed_fails_policy"] = {
|
|
"ContentPolicyViolationErrorAllowedFails": 0
|
|
}
|
|
assert (
|
|
_should_run_cooldown_logic(single_deployment_router, "dep-1", 400, exc) is True
|
|
)
|
|
|
|
|
|
class TestHasExplicitAllowedFailsPolicyForException:
|
|
def test_no_policy_anywhere_returns_false(self, single_deployment_router):
|
|
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
|
|
assert (
|
|
_has_explicit_allowed_fails_policy_for_exception(
|
|
single_deployment_router, "dep-1", exc
|
|
)
|
|
is False
|
|
)
|
|
|
|
def test_router_level_policy_for_matching_exception_returns_false(
|
|
self, single_deployment_router
|
|
):
|
|
"""Deliberately scoped to deployment-level only: a router-level policy
|
|
predates this feature and must not be treated as an explicit per-exception
|
|
opt-in for cooldown-gate purposes."""
|
|
exc = litellm.RateLimitError("rate limited", "openai", "gpt-5-mini")
|
|
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
|
|
RateLimitErrorAllowedFails=3
|
|
)
|
|
assert (
|
|
_has_explicit_allowed_fails_policy_for_exception(
|
|
single_deployment_router, "dep-1", exc
|
|
)
|
|
is False
|
|
)
|
|
|
|
def test_router_level_policy_for_different_exception_returns_false(
|
|
self, single_deployment_router
|
|
):
|
|
exc = litellm.BadRequestError("bad request", "openai", "gpt-5-mini")
|
|
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
|
|
RateLimitErrorAllowedFails=3
|
|
)
|
|
assert (
|
|
_has_explicit_allowed_fails_policy_for_exception(
|
|
single_deployment_router, "dep-1", exc
|
|
)
|
|
is False
|
|
)
|
|
|
|
def test_deployment_level_policy_for_matching_exception_returns_true(
|
|
self, single_deployment_router
|
|
):
|
|
exc = litellm.ContentPolicyViolationError("flagged", "openai", "gpt-5-mini")
|
|
deployment_dict = single_deployment_router.get_model_info(id="dep-1")
|
|
deployment_dict["model_info"]["allowed_fails_policy"] = {
|
|
"ContentPolicyViolationErrorAllowedFails": 0
|
|
}
|
|
assert (
|
|
_has_explicit_allowed_fails_policy_for_exception(
|
|
single_deployment_router, "dep-1", exc
|
|
)
|
|
is True
|
|
)
|
|
|
|
def test_none_deployment_returns_false(self, single_deployment_router):
|
|
exc = litellm.RateLimitError("rate limited", "openai", "gpt-5-mini")
|
|
single_deployment_router.allowed_fails_policy = AllowedFailsPolicy(
|
|
RateLimitErrorAllowedFails=3
|
|
)
|
|
assert (
|
|
_has_explicit_allowed_fails_policy_for_exception(
|
|
single_deployment_router, None, exc
|
|
)
|
|
is False
|
|
)
|
|
|
|
|
|
def test_should_cooldown_deployment_rate_limit_error(testing_litellm_router):
|
|
"""
|
|
Test the _should_cooldown_deployment function when a rate limit error occurs
|
|
"""
|
|
# Test 429 error (rate limit) -> always cooldown a deployment returning 429s
|
|
_exception = litellm.exceptions.RateLimitError(
|
|
"Rate limit", "openai", "gpt-5-mini"
|
|
)
|
|
assert (
|
|
_should_cooldown_deployment(
|
|
testing_litellm_router, "test_deployment", 429, _exception
|
|
)
|
|
is True
|
|
)
|
|
|
|
|
|
def test_should_cooldown_deployment_auth_limit_error(testing_litellm_router):
|
|
"""
|
|
Test the _should_cooldown_deployment function when an auth limit error occurs
|
|
"""
|
|
# Test 401 error (auth limit) -> always cooldown a deployment returning 401s
|
|
_exception = litellm.exceptions.AuthenticationError(
|
|
"Unauthorized", "openai", "gpt-5-mini"
|
|
)
|
|
assert (
|
|
_should_cooldown_deployment(
|
|
testing_litellm_router, "test_deployment", 401, _exception
|
|
)
|
|
is True
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_should_cooldown_deployment(testing_litellm_router):
|
|
"""
|
|
Cooldown a deployment if it fails 60% of requests in 1 minute - DEFAULT threshold is 50%
|
|
"""
|
|
from litellm._logging import verbose_router_logger
|
|
import logging
|
|
|
|
verbose_router_logger.setLevel(logging.DEBUG)
|
|
|
|
# Test 429 error (rate limit) -> always cooldown a deployment returning 429s
|
|
_exception = litellm.exceptions.RateLimitError(
|
|
"Rate limit", "openai", "gpt-5-mini"
|
|
)
|
|
assert (
|
|
_should_cooldown_deployment(
|
|
testing_litellm_router, "test_deployment", 429, _exception
|
|
)
|
|
is True
|
|
)
|
|
|
|
available_deployment = testing_litellm_router.get_available_deployment(
|
|
model="test_deployment"
|
|
)
|
|
print("available_deployment", available_deployment)
|
|
assert available_deployment is not None
|
|
|
|
deployment_id = available_deployment["model_info"]["id"]
|
|
print("deployment_id", deployment_id)
|
|
|
|
# set current success for deployment to 40
|
|
for _ in range(40):
|
|
increment_deployment_successes_for_current_minute(
|
|
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
|
|
)
|
|
|
|
# now we fail 40 requests in a row
|
|
tasks = []
|
|
for _ in range(41):
|
|
tasks.append(
|
|
testing_litellm_router.acompletion(
|
|
model=deployment_id,
|
|
messages=[{"role": "user", "content": "Hello, world!"}],
|
|
max_tokens=100,
|
|
mock_response="litellm.InternalServerError",
|
|
)
|
|
)
|
|
try:
|
|
await asyncio.gather(*tasks)
|
|
except Exception:
|
|
pass
|
|
|
|
await asyncio.sleep(1)
|
|
|
|
# expect this to fail since it's now 51% of requests are failing
|
|
assert (
|
|
_should_cooldown_deployment(
|
|
testing_litellm_router, deployment_id, 500, Exception("Test")
|
|
)
|
|
is True
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_should_cooldown_deployment_allowed_fails_set_on_router():
|
|
"""
|
|
Test the _should_cooldown_deployment function when Router.allowed_fails is set
|
|
"""
|
|
# Create a Router instance with a test deployment
|
|
router = Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {"model": "gpt-5-mini"},
|
|
"model_id": "test_deployment",
|
|
},
|
|
]
|
|
)
|
|
|
|
# Set up allowed_fails for the test deployment
|
|
router.allowed_fails = 100
|
|
|
|
# should not cooldown when fails are below the allowed limit
|
|
for _ in range(100):
|
|
assert (
|
|
_should_cooldown_deployment(
|
|
router, "test_deployment", 500, Exception("Test")
|
|
)
|
|
is False
|
|
)
|
|
|
|
assert (
|
|
_should_cooldown_deployment(router, "test_deployment", 500, Exception("Test"))
|
|
is True
|
|
)
|
|
|
|
|
|
def test_increment_deployment_successes_for_current_minute_does_not_write_to_redis(
|
|
testing_litellm_router,
|
|
):
|
|
"""
|
|
Ensure tracking deployment metrics does not write to redis
|
|
|
|
Important - If it writes to redis on every request it will seriously impact performance / latency
|
|
"""
|
|
from litellm.caching.dual_cache import DualCache
|
|
from litellm.caching.redis_cache import RedisCache
|
|
from litellm.caching.in_memory_cache import InMemoryCache
|
|
from litellm.router_utils.router_callbacks.track_deployment_metrics import (
|
|
increment_deployment_successes_for_current_minute,
|
|
)
|
|
|
|
# Mock RedisCache
|
|
mock_redis_cache = MagicMock(spec=RedisCache)
|
|
|
|
testing_litellm_router.cache = DualCache(
|
|
redis_cache=mock_redis_cache, in_memory_cache=InMemoryCache()
|
|
)
|
|
|
|
# Call the function we're testing
|
|
increment_deployment_successes_for_current_minute(
|
|
litellm_router_instance=testing_litellm_router, deployment_id="test_deployment"
|
|
)
|
|
|
|
increment_deployment_failures_for_current_minute(
|
|
litellm_router_instance=testing_litellm_router, deployment_id="test_deployment"
|
|
)
|
|
|
|
time.sleep(1)
|
|
|
|
# Assert that no methods were called on the mock_redis_cache
|
|
assert not mock_redis_cache.method_calls, "RedisCache methods should not be called"
|
|
|
|
print(
|
|
"in memory cache values=",
|
|
testing_litellm_router.cache.in_memory_cache.cache_dict,
|
|
)
|
|
assert (
|
|
testing_litellm_router.cache.in_memory_cache.get_cache(
|
|
"test_deployment:successes"
|
|
)
|
|
is not None
|
|
)
|
|
|
|
|
|
def test_cast_exception_status_to_int():
|
|
assert cast_exception_status_to_int(200) == 200
|
|
assert cast_exception_status_to_int("404") == 404
|
|
assert cast_exception_status_to_int("invalid") == 500
|
|
|
|
|
|
@pytest.fixture
|
|
def router():
|
|
return Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "gpt-5.5",
|
|
"litellm_params": {"model": "gpt-5.5"},
|
|
"model_info": {
|
|
"id": "gpt-4--0",
|
|
},
|
|
}
|
|
]
|
|
)
|
|
|
|
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
|
|
)
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
|
|
)
|
|
def test_should_cooldown_high_traffic_all_fails(mock_failures, mock_successes, router):
|
|
# Simulate 10 failures, 0 successes
|
|
from litellm.constants import SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD
|
|
|
|
mock_failures.return_value = SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD + 1
|
|
mock_successes.return_value = 0
|
|
|
|
should_cooldown = _should_cooldown_deployment(
|
|
litellm_router_instance=router,
|
|
deployment="gpt-4--0",
|
|
exception_status=500,
|
|
original_exception=Exception("Test error"),
|
|
)
|
|
|
|
assert (
|
|
should_cooldown is True
|
|
), "Should cooldown when all requests fail with sufficient traffic"
|
|
|
|
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
|
|
)
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
|
|
)
|
|
def test_no_cooldown_low_traffic(mock_failures, mock_successes, router):
|
|
# Simulate 3 failures (below MIN_TRAFFIC_THRESHOLD)
|
|
mock_failures.return_value = 3
|
|
mock_successes.return_value = 0
|
|
|
|
should_cooldown = _should_cooldown_deployment(
|
|
litellm_router_instance=router,
|
|
deployment="gpt-4--0",
|
|
exception_status=500,
|
|
original_exception=Exception("Test error"),
|
|
)
|
|
|
|
assert (
|
|
should_cooldown is False
|
|
), "Should not cooldown when traffic is below threshold"
|
|
|
|
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
|
|
)
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
|
|
)
|
|
def test_cooldown_rate_limit(mock_failures, mock_successes, router):
|
|
"""
|
|
Don't cooldown single deployment models, for anything besides traffic
|
|
"""
|
|
mock_failures.return_value = 1
|
|
mock_successes.return_value = 0
|
|
|
|
should_cooldown = _should_cooldown_deployment(
|
|
litellm_router_instance=router,
|
|
deployment="gpt-4--0",
|
|
exception_status=429, # Rate limit error
|
|
original_exception=Exception("Rate limit exceeded"),
|
|
)
|
|
|
|
assert (
|
|
should_cooldown is False
|
|
), "Should not cooldown on rate limit error for single deployment models"
|
|
|
|
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_successes_for_current_minute"
|
|
)
|
|
@patch(
|
|
"litellm.router_utils.cooldown_handlers.get_deployment_failures_for_current_minute"
|
|
)
|
|
def test_mixed_success_failure(mock_failures, mock_successes, router):
|
|
# Simulate 3 failures, 7 successes
|
|
mock_failures.return_value = 3
|
|
mock_successes.return_value = 7
|
|
|
|
should_cooldown = _should_cooldown_deployment(
|
|
litellm_router_instance=router,
|
|
deployment="gpt-4--0",
|
|
exception_status=500,
|
|
original_exception=Exception("Test error"),
|
|
)
|
|
|
|
assert (
|
|
should_cooldown is False
|
|
), "Should not cooldown when failure rate is below threshold"
|
|
|
|
|
|
def test_is_cooldown_required_empty_string_exception_status(testing_litellm_router):
|
|
"""
|
|
Test that _is_cooldown_required returns False when exception_status is an empty string
|
|
"""
|
|
result = _is_cooldown_required(
|
|
litellm_router_instance=testing_litellm_router,
|
|
model_id="test_deployment",
|
|
exception_status="",
|
|
)
|
|
|
|
assert (
|
|
result is False
|
|
), "Should not require cooldown when exception_status is empty string"
|
|
|
|
|
|
def test_should_cooldown_deployment_minimum_request_threshold(testing_litellm_router):
|
|
"""
|
|
Test that error rate cooldown does NOT trigger on first failure.
|
|
|
|
Fixes GitHub issue #17418: Error Rate Cooldown Triggers on First Failed Request
|
|
|
|
The problem: With DEFAULT_FAILURE_THRESHOLD_PERCENT=0.5 (50%), a deployment
|
|
gets cooled down after just 1 failed request because 1/1 = 100% > 50%.
|
|
|
|
The fix: Add a minimum request threshold (DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS)
|
|
before applying error rate cooldown.
|
|
"""
|
|
from litellm.constants import DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS
|
|
|
|
# Get a deployment that's not a single-deployment model group
|
|
# (test_deployment_2 and test_deployment_3 are both for "test_deployment" model)
|
|
available_deployment = testing_litellm_router.get_available_deployment(
|
|
model="test_deployment"
|
|
)
|
|
assert available_deployment is not None
|
|
deployment_id = available_deployment["model_info"]["id"]
|
|
|
|
# Simulate only 1 failure (below minimum threshold)
|
|
# This should NOT trigger cooldown even though 100% > 50%
|
|
increment_deployment_failures_for_current_minute(
|
|
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
|
|
)
|
|
|
|
_exception = litellm.exceptions.InternalServerError(
|
|
"Internal error", "openai", "gpt-5-mini"
|
|
)
|
|
|
|
# With only 1 request, should NOT cooldown (below minimum threshold)
|
|
should_cooldown = _should_cooldown_deployment(
|
|
testing_litellm_router, deployment_id, 500, _exception
|
|
)
|
|
assert (
|
|
should_cooldown is False
|
|
), f"Should NOT cooldown with only 1 failed request (below minimum threshold of {DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS})"
|
|
|
|
# Now add more failures to reach the minimum threshold
|
|
for _ in range(DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS - 1):
|
|
increment_deployment_failures_for_current_minute(
|
|
litellm_router_instance=testing_litellm_router, deployment_id=deployment_id
|
|
)
|
|
|
|
# Now with enough requests (all failures), it SHOULD trigger cooldown
|
|
should_cooldown = _should_cooldown_deployment(
|
|
testing_litellm_router, deployment_id, 500, _exception
|
|
)
|
|
assert (
|
|
should_cooldown is True
|
|
), f"Should cooldown when we have {DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS} failed requests (100% failure rate)"
|