chore(gatewaybench): narrow scope to 4 gateways, simplify chart to overhead bars

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
This commit is contained in:
Devin AI 2026-07-18 23:11:47 +00:00
parent 343979cd31
commit ca97083d03
11 changed files with 69 additions and 182 deletions

View file

@ -4,11 +4,11 @@ A reproducible benchmark for **AI-gateway overhead**: the latency, throughput ce
CursorBench measures model *quality*. The right question for a gateway is not quality but overhead, and specifically the overhead that a coding agent feels on every turn: how much longer until the first token, how evenly the stream flows, how tool calls behave, and how the whole thing holds up under concurrency. GatewayBench measures that, and it measures it the same way for every gateway so the numbers are comparable.
![GatewayBench throughput vs memory](analyze/throughput_vs_memory.png)
Gateways compared: **LiteLLM (Rust)**, **LiteLLM (Python v1)**, **Portkey**, and **Bifrost**.
![GatewayBench streaming TTFT overhead vs memory](analyze/ttft_vs_memory.png)
![GatewayBench overhead comparison](analyze/overhead_comparison.png)
> The charts above use **illustrative placeholder data** to show the design. They are regenerated from real `results/*.jsonl` once the harness has run; see [Generating the charts](#generating-the-charts).
> The chart above uses **illustrative placeholder data** to show the design. It is regenerated from real `results/*.jsonl` once the harness has run; see [Generating the chart](#generating-the-chart).
## The core idea: isolate the gateway from the provider
@ -38,7 +38,7 @@ The mock and the load driver must never be the bottleneck. A Python mock or clos
- the driver is **open-loop** (constant arrival rate), which avoids coordinated omission and keeps the p99/p99.9 tail honest
- runs are deterministic: fixed seeds, fixed payloads, a discarded warmup window
Only the harness is Rust. The gateways run as their real selves: LiteLLM (Python), Bifrost (Go), Portkey (the OSS JS gateway).
Only the harness is Rust. The gateways run as their real selves: LiteLLM Rust, LiteLLM Python v1, Bifrost (Go), and Portkey (the OSS JS gateway).
## Metrics
@ -70,16 +70,13 @@ Load and cost
Since LiteLLM publishes this, transparency is the whole point. Every gateway config, image, and pinned version lives in `gateways/` and is meant to be challenged. Rules
- all four gateways run on the **same single machine** so the numbers are directly comparable; a run split across different hosts would not be
- each gateway runs in its **recommended production config**, not a strawman, and we run two configs per gateway: *bare passthrough* (no logging, DB, cache, or rate limit) to isolate pure proxy cost, and *realistic prod* (key auth plus a logging or spend callback on) to match what people actually run
- equal resources for every gateway (same container CPU and memory limits), so a result is never "who got more cores"
- load driver, gateway, and mock run on separate pinned cores so the driver never steals the gateway's CPU
- fixed seeds and payloads, a discarded warmup window, multiple runs with median-of-runs and variance reported
Category differences we state honestly rather than hide
- **LiteLLM SDK** is an in-process library with no HTTP hop, so it is a different product category from the network proxies. It is reported but kept off the same latency axis
- **OpenRouter** is SaaS only and cannot be pointed at the mock. Any number for it includes their infrastructure, the network round trip to their datacenter, and a real upstream, so it is not comparable to the self-hosted numbers. It is either excluded from the isolated-overhead charts or reported separately as an end-to-end measurement, always clearly labeled
- **Portkey** is benchmarked as the open-source self-hostable gateway, not the SaaS
Portkey is benchmarked as the open-source self-hostable gateway, not the SaaS.
## Layout
@ -102,7 +99,7 @@ gatewaybench/ cargo workspace
throughput/ RPS ceiling + CPU/RSS cost per 1M requests
tail-latency/ p99/p99.9 added latency under concurrency
gateways/ per-gateway configs (bare + prod), pinned versions
litellm-proxy/ bifrost/ portkey/ openrouter/
litellm-rust/ litellm-python/ bifrost/ portkey/
xtask/ `cargo xtask bench` sweeps the matrix -> results/*.jsonl
analyze/ chart generation for RESULTS.md
```
@ -123,9 +120,9 @@ cargo run --release -p ttft
cargo xtask bench
```
## Generating the charts
## Generating the chart
The hero charts are produced from the benchmark results by `analyze/make_hero_charts.py`. Point `_DATA` at real rows from `results/*.jsonl` and regenerate
The overhead comparison chart is produced from the benchmark results by `analyze/make_hero_charts.py`. Point `_DATA` at real rows from `results/*.jsonl` and regenerate
```bash
python analyze/make_hero_charts.py

View file

@ -1,12 +1,13 @@
"""Generate the GatewayBench hero charts (CursorBench-styled scatter plots).
"""Generate the GatewayBench hero chart: a simple overhead comparison.
The numbers here are ILLUSTRATIVE PLACEHOLDERS to convey the chart design; they
are not measured. Replace `_DATA` with real rows from results/*.jsonl once the
harness has run, then regenerate:
Two horizontal-bar panels comparing gateways head-to-head on overhead. The
numbers here are ILLUSTRATIVE PLACEHOLDERS to convey the design; they are not
measured. Replace `_DATA` with real rows from results/*.jsonl once the harness
has run, then regenerate:
python analyze/make_hero_charts.py
Outputs analyze/throughput_vs_memory.png and analyze/ttft_vs_memory.png.
Outputs analyze/overhead_comparison.png.
"""
from __future__ import annotations
@ -18,204 +19,93 @@ import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib import font_manager
_OUT_DIR = Path(__file__).resolve().parent
_INK = "#1f2933"
_MUTED = "#6b7280"
_GRID = "#e5e7eb"
_DARK = "#26272b"
_HILITE = "#f2530b"
@dataclass(frozen=True)
class Point:
class Gateway:
label: str
added_latency_ms: float
memory_mb: float
throughput_rps: float
ttft_overhead_ms: float
color: str
note: str = ""
is_litellm: bool
# ILLUSTRATIVE PLACEHOLDER DATA -- not measured.
_DATA: tuple[Point, ...] = (
Point("LiteLLM Proxy", 450, 1800, 12.0, "#0b7285"),
Point("Bifrost (Go)", 180, 4200, 3.0, "#e8590c"),
Point("Portkey (OSS)", 260, 2600, 7.0, "#ae3ec9"),
Point(
"LiteLLM SDK",
120,
6000,
1.5,
"#2f9e44",
note="in-process, no HTTP hop",
),
_DATA: tuple[Gateway, ...] = (
Gateway("LiteLLM (Rust)", 2.0, 90, True),
Gateway("Bifrost", 3.0, 180, False),
Gateway("Portkey", 7.0, 260, False),
Gateway("LiteLLM (Python v1)", 12.0, 450, True),
)
def _style_axes(ax: "plt.Axes", title: str, ylabel: str, xlabel: str) -> None:
ax.set_title(title, loc="left", fontsize=15, fontweight="bold", color=_INK, pad=18)
ax.set_ylabel(ylabel, fontsize=11, color=_MUTED)
ax.set_xlabel(xlabel, fontsize=11, color=_MUTED)
for spine in ("top", "right"):
def _panel(ax: "plt.Axes", values: list[float], labels: list[str], colors: list[str], xlabel: str) -> None:
y = range(len(labels))
ax.barh(list(y), values, color=colors, height=0.62, zorder=3)
ax.set_yticks(list(y))
ax.set_yticklabels(labels, fontsize=11, color=_INK)
ax.invert_yaxis()
ax.set_xlabel(xlabel, fontsize=11, color=_MUTED, labelpad=10)
for spine in ("top", "right", "left"):
ax.spines[spine].set_visible(False)
for spine in ("left", "bottom"):
ax.spines[spine].set_color(_GRID)
ax.tick_params(colors=_MUTED, labelsize=10)
ax.grid(True, color=_GRID, linewidth=0.8, zorder=0)
ax.spines["bottom"].set_color(_GRID)
ax.tick_params(axis="x", colors=_MUTED, labelsize=10)
ax.tick_params(axis="y", length=0)
ax.grid(True, axis="x", color=_GRID, linewidth=0.8, zorder=0)
ax.set_axisbelow(True)
for i, v in enumerate(values):
ax.text(v + max(values) * 0.02, i, f"{v:g}", va="center", fontsize=10, color=_MUTED)
def _annotate(ax: "plt.Axes", p: Point, dx: float, dy: float) -> None:
text = p.label if not p.note else f"{p.label}"
ax.annotate(
text,
xy=(p.memory_mb, _y(ax, p)),
xytext=(p.memory_mb + dx, _y(ax, p) + dy),
fontsize=11,
def make_chart() -> Path:
ordered = sorted(_DATA, key=lambda g: g.added_latency_ms)
labels = [g.label for g in ordered]
colors = [_HILITE if g.is_litellm else _DARK for g in ordered]
fig, (ax_l, ax_r) = plt.subplots(1, 2, figsize=(11.5, 5.2), dpi=130)
fig.patch.set_facecolor("white")
fig.suptitle(
"Gateway overhead, lower is better",
x=0.06,
y=1.02,
ha="left",
fontsize=16,
fontweight="bold",
color=_INK,
fontweight="medium",
)
if p.note:
ax.annotate(
p.note,
xy=(p.memory_mb, _y(ax, p)),
xytext=(p.memory_mb + dx, _y(ax, p) + dy - _note_gap(ax)),
fontsize=8.5,
color=_MUTED,
style="italic",
)
_MODE = {"metric": "throughput_rps"}
def _y(ax: "plt.Axes", p: Point) -> float:
return getattr(p, _MODE["metric"])
def _note_gap(ax: "plt.Axes") -> float:
lo, hi = ax.get_ylim()
return (hi - lo) * 0.045
def _frontier(ax: "plt.Axes", pts: list[Point], color: str) -> None:
xs = [p.memory_mb for p in pts]
ys = [_y(ax, p) for p in pts]
ax.plot(xs, ys, color=color, linewidth=1.6, alpha=0.55, zorder=2)
def _scatter(pts: tuple[Point, ...]) -> None:
for p in pts:
ax = plt.gca()
ax.scatter([p.memory_mb], [_y(ax, p)], s=90, color=p.color, zorder=3, edgecolors="white", linewidths=1.2)
def make_throughput_chart() -> Path:
_MODE["metric"] = "throughput_rps"
fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130)
fig.patch.set_facecolor("white")
ax.set_facecolor("white")
ax.set_xlim(80, 520)
ax.set_ylim(1000, 6800)
_style_axes(
ax,
"GatewayBench: throughput vs memory",
"sustained throughput (req/s) higher is better ->",
"peak resident memory (MB) <- lower is better",
_panel(ax_l, [g.added_latency_ms for g in ordered], labels, colors, "p99 added latency (ms)")
_panel(
ax_r,
[g.memory_mb for g in ordered],
["" for _ in ordered],
colors,
"peak memory (MB)",
)
_scatter(_DATA)
frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb)
_frontier(ax, frontier_pts, "#0b7285")
offsets = {
"LiteLLM Proxy": (10, 120),
"Bifrost (Go)": (10, 180),
"Portkey (OSS)": (14, 150),
"LiteLLM SDK": (-40, -320),
}
for p in _DATA:
dx, dy = offsets[p.label]
_annotate(ax, p, dx, dy)
ax.annotate(
"efficiency frontier",
xy=(120, 6000),
xytext=(300, 6300),
fontsize=12,
color=_MUTED,
fontweight="bold",
)
ax.annotate(
fig.text(
0.99,
-0.04,
"ILLUSTRATIVE PLACEHOLDER DATA -- not measured",
xy=(0, 0),
xytext=(0.99, -0.13),
xycoords="axes fraction",
ha="right",
fontsize=8.5,
color="#b91c1c",
style="italic",
)
out = _OUT_DIR / "throughput_vs_memory.png"
fig.tight_layout()
fig.savefig(out, bbox_inches="tight", facecolor="white")
plt.close(fig)
return out
def make_ttft_chart() -> Path:
_MODE["metric"] = "ttft_overhead_ms"
fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130)
fig.patch.set_facecolor("white")
ax.set_facecolor("white")
ax.set_xlim(80, 520)
ax.set_ylim(0, 14)
_style_axes(
ax,
"GatewayBench: streaming TTFT overhead vs memory",
"p99 added time-to-first-chunk (ms) <- lower is better",
"peak resident memory (MB) <- lower is better",
)
_scatter(_DATA)
frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb)
_frontier(ax, frontier_pts, "#0b7285")
offsets = {
"LiteLLM Proxy": (10, 0.4),
"Bifrost (Go)": (10, 0.4),
"Portkey (OSS)": (10, 0.4),
"LiteLLM SDK": (-30, 0.6),
}
for p in _DATA:
dx, dy = offsets[p.label]
_annotate(ax, p, dx, dy)
ax.annotate(
"ideal: fast + lean (bottom-left)",
xy=(120, 1.5),
xytext=(150, 11.5),
fontsize=11,
color=_MUTED,
fontweight="bold",
)
ax.annotate(
"ILLUSTRATIVE PLACEHOLDER DATA -- not measured",
xy=(0, 0),
xytext=(0.99, -0.13),
xycoords="axes fraction",
ha="right",
fontsize=8.5,
color="#b91c1c",
style="italic",
)
out = _OUT_DIR / "ttft_vs_memory.png"
fig.tight_layout()
out = _OUT_DIR / "overhead_comparison.png"
fig.tight_layout(rect=(0, 0, 1, 0.98))
fig.savefig(out, bbox_inches="tight", facecolor="white")
plt.close(fig)
return out
def main() -> None:
_ = font_manager # ensure font cache import for consistent rendering
t = make_throughput_chart()
s = make_ttft_chart()
print(f"wrote {t}")
print(f"wrote {s}")
out = make_chart()
print(f"wrote {out}")
if __name__ == "__main__":

Binary file not shown.

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 85 KiB

View file

@ -1 +0,0 @@
Placeholder for LiteLLM proxy benchmark configuration.

View file

@ -0,0 +1 @@
Placeholder for LiteLLM (Python v1) proxy benchmark configuration (bare + prod).

View file

@ -0,0 +1 @@
Placeholder for LiteLLM (Rust) gateway benchmark configuration (bare + prod).

View file

@ -1 +0,0 @@
Placeholder for OpenRouter benchmark configuration.