mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-09 03:18:44 +00:00
Benchmark results comparing 4 scenarios: - Scenario 1: Baseline uvicorn (1 worker) — 19 req/s - Scenario 2: Standard gunicorn (4 workers) — 76-78 req/s - Scenario 3: fast-litellm gunicorn (4 workers) — 79 req/s Key findings: - Multi-worker is the biggest win (4x throughput) - fast-litellm gives ~8% better P50 latency - fast-litellm has much lower run-to-run variance (0.9% vs 4.0% CoV) - Remote DB overhead (~1s/req) dominates, masking proxy-level gains - All scenarios: 0 failures across 6000 requests Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
14 lines
411 B
Python
14 lines
411 B
Python
"""
|
|
Gunicorn wrapper for standard LiteLLM proxy (no fast-litellm).
|
|
|
|
Usage:
|
|
CONFIG_FILE_PATH=benchmark_config.yaml \
|
|
gunicorn standard_app:app --preload -w 4 -k uvicorn.workers.UvicornWorker -b 0.0.0.0:4000
|
|
"""
|
|
|
|
import os
|
|
|
|
# Set config file path before anything else loads
|
|
os.environ.setdefault("CONFIG_FILE_PATH", "benchmark_config.yaml")
|
|
|
|
from litellm.proxy.proxy_server import app # noqa: F401, E402
|