mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
_completion_streaming_iterator.stream_with_fallbacks() re-entered the fallback chain via function_with_fallbacks(), which tries the original group first. The retry returns a fresh stream that succeeds at creation time (HTTP 200) and only fails while being iterated; each nested Router.completion() wraps its own stream in this same iterator, so every retry fails again during iteration and re-enters the chain — recursing until the stack runs out (measured: 478 requests to the failing group, ~116 s, then InternalServerError, with a healthy fallback configured). Mirror the async twin: hand the MidStreamFallbackError to async_function_with_fallbacks_common_utils() via run_async_function, which cools down the failed deployment and walks the fallback list directly. Sync tests that patched function_with_fallbacks are updated to the new re-entry point, plus a regression test asserting the triggering error reaches the common utils and the original group is not re-run. Fixes #43945 |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| test_enforce_model_rate_limits.py | ||
| test_io_token_rate_limits.py | ||
| test_router.py | ||