litellm/litellm/router_utils
devin-ai-integration[bot] 96f58fac53
fix(router): don't cool down parent deployment on advisor sub-call failure (#33792)
* fix(router): don't cool down parent deployment on advisor sub-call failure

Advisor orchestration issues a sub-call to a different provider/credentials than the selected deployment. When that sub-call fails (e.g. a 401 because no advisor API key is configured), the exception propagates up and the router's deployment_callback_on_failure attributes it to the healthy parent deployment's model_info.id, cooling it down and rejecting unrelated callers to the same model group.

Tag advisor sub-call failures on the exception and skip cooldown for them in deployment_callback_on_failure. The exception is tagged rather than wrapped so its type is preserved and retry/fallback classification and the client-facing error are unchanged. Genuine executor/deployment failures are untagged and still cool down as before.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(router): tag advisor orchestration failures via provider-neutral util

Address review on LIT-4565: move the cooldown-exemption marker into
litellm/router_utils/cooldown_handlers.py so the router imports it at
module top instead of an in-function anthropic import, and extend the
exemption to AdvisorMaxIterationsError so a max-iterations orchestration
failure no longer cools down the healthy executor deployment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-25 10:17:13 -07:00
..
pre_call_checks fix(router): take the lowest minimum across a model group, not the highest 2026-07-16 19:53:10 -07:00
router_callbacks build(deps-dev): bump black to 26.3.1 and apply formatting (#28525) 2026-05-21 17:24:18 -07:00
add_retry_fallback_headers.py feat(router): add separate ITPM/OTPM deployment rate limits (#31952) 2026-07-05 21:58:35 +05:30
batch_utils.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
client_initalization_utils.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
clientside_credential_handler.py fix(proxy): extend banned-params + admin-clear lists for NVIDIA Riva (VERIA-493) (#31742) 2026-06-30 15:30:08 -07:00
common_utils.py fix(proxy): route master key to team-scoped models (#32926) 2026-07-14 10:11:56 -07:00
cooldown_cache.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cooldown_callbacks.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cooldown_handlers.py fix(router): don't cool down parent deployment on advisor sub-call failure (#33792) 2026-07-25 10:17:13 -07:00
fallback_event_handlers.py fix(router): mask provider credentials embedded in fallback error messages (#32083) 2026-07-03 18:48:06 -07:00
get_retry_from_policy.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
handle_error.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
health_state_cache.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
pattern_match_deployments.py fix(router): keep team wildcard routers fresh and prioritize them over global patterns 2026-07-17 18:19:02 -07:00
prompt_caching_cache.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
response_headers.py LiteLLM Minor Fixes & Improvements (11/26/2024) (#6913) 2024-11-28 00:01:38 +05:30
search_api_router.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00