docs benchmark

This commit is contained in:
Ishaan Jaff 2025-01-14 10:48:43 -08:00
parent eb2770fee2
commit 8c016d0184

View file

@ -18,13 +18,13 @@ model_list:
api_key: "test"
```
## 1 Instance LiteLLM Proxy
### 1 Instance LiteLLM Proxy
In these tests the median latency of directly calling the fake-openai-endpoint is 60ms.
| Metric | Litellm Proxy (1 Instance) |
|--------|------------------------|
| RPS | 500 |
| RPS | 475 |
| Median Latency (ms) | 100 |
| Latency overhead added by LiteLLM Proxy | 40ms |
@ -35,8 +35,9 @@ In these tests the median latency of directly calling the fake-openai-endpoint i
<Image img={require('../img/instances_vs_rps.png')} /> -->
#### Key Findings
- Single instance: 500 RPS @ 100ms latency
- 4 LiteLLM instances: 1000 RPS @ 100ms latency
- Single instance: 475 RPS @ 100ms latency
- 2 LiteLLM instances: 950 RPS @ 100ms latency
- 4 LiteLLM instances: 1900 RPS @ 100ms latency
### 2 Instances
@ -45,9 +46,16 @@ In these tests the median latency of directly calling the fake-openai-endpoint i
| Metric | Litellm Proxy (2 Instances) |
|--------|------------------------|
| Median Latency (ms) | 100 |
| RPS | 500 |
| RPS | 950 |
## Machine Spec used for testing
Each machine deploying LiteLLM had the following specs:
- 2 CPU
- 4GB RAM
## Logging Callbacks