litellm/tests/test_litellm/router_utils
yassin 8d972eefc7 feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use
Replace the per-deployment asyncio.Semaphore with MaxParallelRequestsLimit, which admits a call synchronously or raises the router's RateLimitError (429) right away. Nothing waits for a slot any more, so the max_parallel_requests_queue_size and default_max_parallel_requests_queue_size settings from the earlier commits are dropped along with their proxy validation, dashboard control and generated schema entries. The rpm/tpm derivation of the cap is unchanged. Every router endpoint family now enters the slot through one _deployment_slot context, and the provider coroutine is only created once the slot is held

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:14:09 +00:00
..
pre_call_checks Merge pull request #41174 from BerriAI/litellm_tier_model_affinity 2026-09-15 09:54:53 -07:00
test_add_retry_fallback_headers.py feat(proxy): expose complexity routing headers (#40792) 2026-09-11 17:14:40 -07:00
test_auto_router_model_naming.py feat(router): apply entitlement limits to forecast classifiers 2026-09-15 16:17:00 -07:00
test_auto_router_tuning_baseline.py feat(complexity_router): rebalance heuristic weights in the dashboard and grade custom dimensions by match count (#40205) 2026-09-08 13:28:57 -07:00
test_client_initalization_utils.py feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use 2026-09-17 18:14:09 +00:00
test_cooldown_cache.py fix(router): give cooldowns their own cache so siblings see a bench in ~1s (#40025) 2026-09-08 10:11:20 -07:00
test_cooldown_handlers.py test(router): annotate return types of team cooldown test helpers 2026-09-13 09:50:12 +00:00
test_fallback_event_handlers.py fix: re-check budget on router fallback targets 2026-09-16 19:06:38 +09:00
test_get_retry_from_policy.py refactor(router): resolve retry policy by exception MRO and add DefaultRetries 2026-09-04 16:09:01 -07:00
test_health_check_allowed_fails_integration.py fix(router): count allowed_fails in the shared router cache so multi-worker proxies bench a deployment fleet-wide (#40224) 2026-09-08 13:54:48 -07:00
test_health_state_cache.py fix(router): treat a breaker-refused Redis read as a miss in the health state cache 2026-09-10 18:16:59 -07:00
test_pattern_match_deployments.py fix(proxy): make the invalid-model 403 path cheap under a burst of rejections (#39892) 2026-09-05 11:53:04 -07:00
test_reasoning_effort_capability.py fix(cost-map): stop advertising reasoning_effort max on the azure gpt-6-astra rows 2026-09-05 22:31:32 -07:00
test_router_health_check_routing.py feat(health): opt-in model-group allowlist for background health checks and health-check routing (#38539) 2026-08-27 12:25:56 -07:00
test_router_interactions_endpoints.py Gemini managed agents support (#28270) 2026-05-19 16:02:03 -07:00
test_router_utils_common_utils.py fix(router): keep a model's own provider prefix for generic SDK calls 2026-09-04 23:32:26 -07:00