mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-09 03:18:44 +00:00
Merge pull request #2667 from BerriAI/litellm_update_gunicorn_instructions
[Docs] update gunicorn instructions - Uvicorn perf is significantly better on K8s
This commit is contained in:
commit
f646a4612b
3 changed files with 8 additions and 19 deletions
|
|
@ -70,5 +70,5 @@ EXPOSE 4000/tcp
|
|||
ENTRYPOINT ["litellm"]
|
||||
|
||||
# Append "--detailed_debug" to the end of CMD to view detailed debug logs
|
||||
# CMD ["--port", "4000", "--config", "./proxy_server_config.yaml", "--run_gunicorn", "--detailed_debug"]
|
||||
CMD ["--port", "4000", "--config", "./proxy_server_config.yaml", "--run_gunicorn", "--num_workers", "4"]
|
||||
# CMD ["--port", "4000", "--config", "./proxy_server_config.yaml"]
|
||||
CMD ["--port", "4000", "--config", "./proxy_server_config.yaml"]
|
||||
|
|
|
|||
|
|
@ -72,5 +72,5 @@ EXPOSE 4000/tcp
|
|||
ENTRYPOINT ["litellm"]
|
||||
|
||||
# Append "--detailed_debug" to the end of CMD to view detailed debug logs
|
||||
# CMD ["--port", "4000","--run_gunicorn", "--detailed_debug"]
|
||||
CMD ["--port", "4000", "--run_gunicorn"]
|
||||
# CMD ["--port", "4000", "--detailed_debug"]
|
||||
CMD ["--port", "4000"]
|
||||
|
|
|
|||
|
|
@ -103,7 +103,10 @@ RUN chmod +x entrypoint.sh
|
|||
EXPOSE 4000/tcp
|
||||
|
||||
# Override the CMD instruction with your desired command and arguments
|
||||
CMD ["--port", "4000", "--config", "config.yaml", "--detailed_debug", "--run_gunicorn"]
|
||||
# WARNING: FOR PROD DO NOT USE `--detailed_debug` it slows down response times, instead use the following CMD
|
||||
# CMD ["--port", "4000", "--config", "config.yaml"]
|
||||
|
||||
CMD ["--port", "4000", "--config", "config.yaml", "--detailed_debug"]
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
|
|
@ -478,20 +481,6 @@ ghcr.io/berriai/litellm-database:main-latest --config your_config.yaml
|
|||
### 1. Switch of debug logs in production
|
||||
don't use [`--detailed-debug`, `--debug`](https://docs.litellm.ai/docs/proxy/debugging#detailed-debug) or `litellm.set_verbose=True`. We found using debug logs can add 5-10% latency per LLM API call
|
||||
|
||||
### 2. Use `run_gunicorn` and `num_workers`
|
||||
|
||||
Example setting `--run_gunicorn` and `--num_workers`
|
||||
```shell
|
||||
docker run ghcr.io/berriai/litellm-database:main-latest --run_gunicorn --num_workers 4
|
||||
```
|
||||
|
||||
Why `Gunicorn`?
|
||||
- Gunicorn takes care of running multiple instances of your web application
|
||||
- Gunicorn is ideal for running litellm proxy on cluster of machines with Kubernetes
|
||||
|
||||
Why `num_workers`?
|
||||
Setting `num_workers` to the number of CPUs available ensures optimal utilization of system resources by matching the number of worker processes to the available CPU cores.
|
||||
|
||||
|
||||
## Advanced Deployment Settings
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue