mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
Update /realtime benchmarks to show current performance only
- Removed before/after comparison, showing only current metrics - Clarified that benchmarks are e2e latency against fake realtime endpoint - Simplified table format for better readability Co-authored-by: ishaan <ishaan@berri.ai>
This commit is contained in:
parent
e5e20954a5
commit
5b7458234b
1 changed files with 8 additions and 14 deletions
|
|
@ -50,17 +50,17 @@ In these tests the baseline latency characteristics are measured against a fake-
|
|||
|
||||
## `/realtime` API Benchmarks
|
||||
|
||||
LiteLLM's `/realtime` endpoint has been optimized for low-latency WebSocket connections, achieving significant performance improvements through removal of redundant encodings, SSL context reuse, and caching of formatting strings.
|
||||
End-to-end latency benchmarks for the `/realtime` endpoint tested against a fake realtime endpoint.
|
||||
|
||||
### Performance Metrics
|
||||
|
||||
| Metric | Before | After | Improvement |
|
||||
| --------------- | --------- | --------- | -------------------------- |
|
||||
| Median latency | 2,200 ms | **59 ms** | **−97% (~37× faster)** |
|
||||
| p95 latency | 8,500 ms | **67 ms** | **−99% (~127× faster)** |
|
||||
| p99 latency | 18,000 ms | **99 ms** | **−99% (~182× faster)** |
|
||||
| Average latency | 3,214 ms | **63 ms** | **−98% (~51× faster)** |
|
||||
| RPS | 165 | **1,207** | **+631% (~7.3× increase)** |
|
||||
| Metric | Value |
|
||||
| --------------- | ---------- |
|
||||
| Median latency | 59 ms |
|
||||
| p95 latency | 67 ms |
|
||||
| p99 latency | 99 ms |
|
||||
| Average latency | 63 ms |
|
||||
| RPS | 1,207 |
|
||||
|
||||
### Test Setup
|
||||
|
||||
|
|
@ -70,12 +70,6 @@ LiteLLM's `/realtime` endpoint has been optimized for low-latency WebSocket conn
|
|||
| **System** | 4 vCPUs, 8 GB RAM, 4 workers, 4 instances |
|
||||
| **Database** | PostgreSQL (Redis unused) |
|
||||
|
||||
### Key Optimizations
|
||||
|
||||
- Removed redundant encodings on the hot path
|
||||
- Reused shared SSL contexts to prevent excessive memory allocation
|
||||
- Cached formatting strings that were being regenerated twice per request
|
||||
|
||||
## Machine Spec used for testing
|
||||
|
||||
Each machine deploying LiteLLM had the following specs:
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue