litellm/standard_app.py
Cursor Agent e2baaaddb0
Add benchmark results: fast-litellm vs standard LiteLLM proxy
Benchmark results comparing 4 scenarios:
- Scenario 1: Baseline uvicorn (1 worker) — 19 req/s
- Scenario 2: Standard gunicorn (4 workers) — 76-78 req/s
- Scenario 3: fast-litellm gunicorn (4 workers) — 79 req/s

Key findings:
- Multi-worker is the biggest win (4x throughput)
- fast-litellm gives ~8% better P50 latency
- fast-litellm has much lower run-to-run variance (0.9% vs 4.0% CoV)
- Remote DB overhead (~1s/req) dominates, masking proxy-level gains
- All scenarios: 0 failures across 6000 requests

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-16 16:39:57 +00:00

14 lines
411 B
Python

"""
Gunicorn wrapper for standard LiteLLM proxy (no fast-litellm).
Usage:
CONFIG_FILE_PATH=benchmark_config.yaml \
gunicorn standard_app:app --preload -w 4 -k uvicorn.workers.UvicornWorker -b 0.0.0.0:4000
"""
import os
# Set config file path before anything else loads
os.environ.setdefault("CONFIG_FILE_PATH", "benchmark_config.yaml")
from litellm.proxy.proxy_server import app # noqa: F401, E402