diff --git a/docs/my-website/docs/benchmarks.md b/docs/my-website/docs/benchmarks.md index 7939f835d18..c445ff303a1 100644 --- a/docs/my-website/docs/benchmarks.md +++ b/docs/my-website/docs/benchmarks.md @@ -18,13 +18,13 @@ model_list: api_key: "test" ``` -## 1 Instance LiteLLM Proxy +### 1 Instance LiteLLM Proxy In these tests the median latency of directly calling the fake-openai-endpoint is 60ms. | Metric | Litellm Proxy (1 Instance) | |--------|------------------------| -| RPS | 500 | +| RPS | 475 | | Median Latency (ms) | 100 | | Latency overhead added by LiteLLM Proxy | 40ms | @@ -35,8 +35,9 @@ In these tests the median latency of directly calling the fake-openai-endpoint i --> #### Key Findings -- Single instance: 500 RPS @ 100ms latency -- 4 LiteLLM instances: 1000 RPS @ 100ms latency +- Single instance: 475 RPS @ 100ms latency +- 2 LiteLLM instances: 950 RPS @ 100ms latency +- 4 LiteLLM instances: 1900 RPS @ 100ms latency ### 2 Instances @@ -45,9 +46,16 @@ In these tests the median latency of directly calling the fake-openai-endpoint i | Metric | Litellm Proxy (2 Instances) | |--------|------------------------| | Median Latency (ms) | 100 | -| RPS | 500 | +| RPS | 950 | +## Machine Spec used for testing + +Each machine deploying LiteLLM had the following specs: + +- 2 CPU +- 4GB RAM + ## Logging Callbacks