feat(router): add order-based fallback so higher order deployments are tried on failure

When a request to an order=1 deployment fails, the router now
automatically tries order=2, order=3, etc. before falling through to
external fallbacks. Works for all error types (429, 404, connection
errors). Requires enable_pre_call_checks=True.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
Sameer Kankute 2026-03-26 12:07:37 +05:30 • committed by Yuneng Jiang
parent 040c7b53d3
commit 08d1303c3c
No known key found for this signature in database
2 changed files with 38 additions and 2 deletions

View file

@ -324,7 +324,7 @@ model_list:
litellm_params:
model: azure/gpt-4-fallback
api_key: os.environ/AZURE_API_KEY_2
order: 2 # 👈 Used when order=1 is unavailable
order: 2 # 👈 Used when order=1 fails
router_settings:
enable_pre_call_checks: true # 👈 Required for 'order' to work
@ -334,6 +334,42 @@ router_settings:
The `order` parameter requires `enable_pre_call_checks: true` in `router_settings`.
:::
### How order-based fallback works
When a request to an `order=1` deployment fails (connection error, 404, 429, etc.), the router automatically tries `order=2` deployments, then `order=3`, and so on. Each order level gets its own set of retries before escalating to the next.
If all order levels are exhausted, the router falls through to any configured [model-level fallbacks](#fallbacks).
```yaml
model_list:
- model_name: gpt-4
litellm_params:
model: azure/gpt-4-primary
api_key: os.environ/AZURE_API_KEY
order: 1
- model_name: gpt-4
litellm_params:
model: azure/gpt-4-secondary
api_key: os.environ/AZURE_API_KEY_2
order: 2
- model_name: gpt-4-fallback
litellm_params:
model: openai/gpt-4
api_key: os.environ/OPENAI_API_KEY
router_settings:
enable_pre_call_checks: true
fallbacks:
- gpt-4:
- gpt-4-fallback # tried after all order levels fail
```
:::important
The `order` parameter requires `enable_pre_call_checks: true` in `router_settings`.
:::
If `order=1` deployment is unavailable (e.g., rate-limited), the router falls back to `order=2` deployments.
### When You'll See Load Balancing in Action

View file

@ -889,7 +889,7 @@ model_list:
litellm_params:
model: azure/gpt-4-fallback
api_key: os.environ/AZURE_API_KEY_2
order: 2 # 👈 Used when order=1 is unavailable
order: 2 # 👈 Tried when order=1 fails
router_settings:
enable_pre_call_checks: true # 👈 Required for 'order' to work