mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-06 08:16:43 +00:00
The SLO measures how many gateway replicas happen to be warm rather than the request path. Clearing the floor needs roughly 5-7 replicas at ~10-14 RPS each; stage idles at one and reactive HPA scale-up lands minutes into a ~3 minute test. It failed both of its assertions on consecutive days: 93.3% errors at an inflated 264 RPS, where the failing requests never reached a pod and closed-loop RPS rose because they failed fast, then 16.7 RPS with zero failures. The covers marker and registry row stay put, so the cell returns to the gap list rather than disappearing from the denominator. |
||
|---|---|---|
| .. | ||
| conftest.py | ||
| load_client.py | ||
| load_constants.py | ||
| locust_load.py | ||
| locustfile.py | ||
| session_anomaly.py | ||
| test_chat_completions_throughput_e2e.py | ||
| test_session_anomaly.py | ||
| test_weekly_session_anomaly_e2e.py | ||
| weekly_anomaly_config.yml | ||