From a26da3783cfd1df99af7a5d4ee705c10f8051881 Mon Sep 17 00:00:00 2001 From: Deepanshu Date: Tue, 25 Aug 2026 16:48:01 -0400 Subject: [PATCH] docs(rate-limiting): correct the apply_to_models fallback-escape note Update the module docstring now that the fallback bypass is fixed instead of merely documented as an accepted limitation. --- .../hooks/global_tag_rate_limits_hook.py | 21 +++++++++---------- 1 file changed, 10 insertions(+), 11 deletions(-) diff --git a/litellm/proxy/hooks/global_tag_rate_limits_hook.py b/litellm/proxy/hooks/global_tag_rate_limits_hook.py index 4a36d8a4731..5af7e4d8714 100644 --- a/litellm/proxy/hooks/global_tag_rate_limits_hook.py +++ b/litellm/proxy/hooks/global_tag_rate_limits_hook.py @@ -36,17 +36,16 @@ how its bucket is shared: track whichever model actually ends up serving a request needs `model_info.tag_rate_limits` instead; but (2) if this hook's own admission *rejects* the request, - `common_request_processing.py` catches that rejection and retries the - whole pre-call pipeline (this hook included) against - `litellm_settings.fallbacks`/`router_settings.fallbacks` configured for - the original model, with `data["model"]` mutated to the fallback target -- - a fresh, correct evaluation of `apply_to_models` against that new model. - If that fallback model is NOT also in `apply_to_models`, this is a real - escape hatch: the rejected request is transparently admitted anyway. - List every model that should share the cap (the whole chain, not just its - primary member) in `apply_to_models` to close this -- a fallback target - that's also listed re-hits the same, already-exhausted shared bucket and - is correctly rejected too. + `common_request_processing.py` would otherwise catch that rejection and + retry the whole pre-call pipeline against + `litellm_settings.fallbacks`/`router_settings.fallbacks`, with + `data["model"]` mutated to the fallback target -- silently admitting the + request via a model outside `apply_to_models`, defeating the cap. A + rejection from an `apply_to_models`-scoped entry carries + `detail["cross_model_scope"] = True` for exactly this reason: + `_pre_call_with_fallbacks` checks that marker and re-raises immediately + instead of trying any fallback, so this bypass is closed regardless of + whether the fallback chain is also listed in `apply_to_models`. - `scope_by_key_hash` (already exists on `TagRateLimitEntry`): whether the keys an entry applies to share one bucket, or each gets its own.