diff --git a/gatewaybench/README.md b/gatewaybench/README.md index 5fdcd22ad11..2ac2915a6e1 100644 --- a/gatewaybench/README.md +++ b/gatewaybench/README.md @@ -4,11 +4,11 @@ A reproducible benchmark for **AI-gateway overhead**: the latency, throughput ce CursorBench measures model *quality*. The right question for a gateway is not quality but overhead, and specifically the overhead that a coding agent feels on every turn: how much longer until the first token, how evenly the stream flows, how tool calls behave, and how the whole thing holds up under concurrency. GatewayBench measures that, and it measures it the same way for every gateway so the numbers are comparable. -![GatewayBench throughput vs memory](analyze/throughput_vs_memory.png) +Gateways compared: **LiteLLM (Rust)**, **LiteLLM (Python v1)**, **Portkey**, and **Bifrost**. -![GatewayBench streaming TTFT overhead vs memory](analyze/ttft_vs_memory.png) +![GatewayBench overhead comparison](analyze/overhead_comparison.png) -> The charts above use **illustrative placeholder data** to show the design. They are regenerated from real `results/*.jsonl` once the harness has run; see [Generating the charts](#generating-the-charts). +> The chart above uses **illustrative placeholder data** to show the design. It is regenerated from real `results/*.jsonl` once the harness has run; see [Generating the chart](#generating-the-chart). ## The core idea: isolate the gateway from the provider @@ -38,7 +38,7 @@ The mock and the load driver must never be the bottleneck. A Python mock or clos - the driver is **open-loop** (constant arrival rate), which avoids coordinated omission and keeps the p99/p99.9 tail honest - runs are deterministic: fixed seeds, fixed payloads, a discarded warmup window -Only the harness is Rust. The gateways run as their real selves: LiteLLM (Python), Bifrost (Go), Portkey (the OSS JS gateway). +Only the harness is Rust. The gateways run as their real selves: LiteLLM Rust, LiteLLM Python v1, Bifrost (Go), and Portkey (the OSS JS gateway). ## Metrics @@ -70,16 +70,13 @@ Load and cost Since LiteLLM publishes this, transparency is the whole point. Every gateway config, image, and pinned version lives in `gateways/` and is meant to be challenged. Rules +- all four gateways run on the **same single machine** so the numbers are directly comparable; a run split across different hosts would not be - each gateway runs in its **recommended production config**, not a strawman, and we run two configs per gateway: *bare passthrough* (no logging, DB, cache, or rate limit) to isolate pure proxy cost, and *realistic prod* (key auth plus a logging or spend callback on) to match what people actually run - equal resources for every gateway (same container CPU and memory limits), so a result is never "who got more cores" - load driver, gateway, and mock run on separate pinned cores so the driver never steals the gateway's CPU - fixed seeds and payloads, a discarded warmup window, multiple runs with median-of-runs and variance reported -Category differences we state honestly rather than hide - -- **LiteLLM SDK** is an in-process library with no HTTP hop, so it is a different product category from the network proxies. It is reported but kept off the same latency axis -- **OpenRouter** is SaaS only and cannot be pointed at the mock. Any number for it includes their infrastructure, the network round trip to their datacenter, and a real upstream, so it is not comparable to the self-hosted numbers. It is either excluded from the isolated-overhead charts or reported separately as an end-to-end measurement, always clearly labeled -- **Portkey** is benchmarked as the open-source self-hostable gateway, not the SaaS +Portkey is benchmarked as the open-source self-hostable gateway, not the SaaS. ## Layout @@ -102,7 +99,7 @@ gatewaybench/ cargo workspace throughput/ RPS ceiling + CPU/RSS cost per 1M requests tail-latency/ p99/p99.9 added latency under concurrency gateways/ per-gateway configs (bare + prod), pinned versions - litellm-proxy/ bifrost/ portkey/ openrouter/ + litellm-rust/ litellm-python/ bifrost/ portkey/ xtask/ `cargo xtask bench` sweeps the matrix -> results/*.jsonl analyze/ chart generation for RESULTS.md ``` @@ -123,9 +120,9 @@ cargo run --release -p ttft cargo xtask bench ``` -## Generating the charts +## Generating the chart -The hero charts are produced from the benchmark results by `analyze/make_hero_charts.py`. Point `_DATA` at real rows from `results/*.jsonl` and regenerate +The overhead comparison chart is produced from the benchmark results by `analyze/make_hero_charts.py`. Point `_DATA` at real rows from `results/*.jsonl` and regenerate ```bash python analyze/make_hero_charts.py diff --git a/gatewaybench/analyze/make_hero_charts.py b/gatewaybench/analyze/make_hero_charts.py index c3c0ca30bd4..609ac2bd5c6 100644 --- a/gatewaybench/analyze/make_hero_charts.py +++ b/gatewaybench/analyze/make_hero_charts.py @@ -1,12 +1,13 @@ -"""Generate the GatewayBench hero charts (CursorBench-styled scatter plots). +"""Generate the GatewayBench hero chart: a simple overhead comparison. -The numbers here are ILLUSTRATIVE PLACEHOLDERS to convey the chart design; they -are not measured. Replace `_DATA` with real rows from results/*.jsonl once the -harness has run, then regenerate: +Two horizontal-bar panels comparing gateways head-to-head on overhead. The +numbers here are ILLUSTRATIVE PLACEHOLDERS to convey the design; they are not +measured. Replace `_DATA` with real rows from results/*.jsonl once the harness +has run, then regenerate: python analyze/make_hero_charts.py -Outputs analyze/throughput_vs_memory.png and analyze/ttft_vs_memory.png. +Outputs analyze/overhead_comparison.png. """ from __future__ import annotations @@ -18,204 +19,93 @@ import matplotlib matplotlib.use("Agg") import matplotlib.pyplot as plt -from matplotlib import font_manager _OUT_DIR = Path(__file__).resolve().parent _INK = "#1f2933" _MUTED = "#6b7280" _GRID = "#e5e7eb" +_DARK = "#26272b" +_HILITE = "#f2530b" @dataclass(frozen=True) -class Point: +class Gateway: label: str + added_latency_ms: float memory_mb: float - throughput_rps: float - ttft_overhead_ms: float - color: str - note: str = "" + is_litellm: bool # ILLUSTRATIVE PLACEHOLDER DATA -- not measured. -_DATA: tuple[Point, ...] = ( - Point("LiteLLM Proxy", 450, 1800, 12.0, "#0b7285"), - Point("Bifrost (Go)", 180, 4200, 3.0, "#e8590c"), - Point("Portkey (OSS)", 260, 2600, 7.0, "#ae3ec9"), - Point( - "LiteLLM SDK", - 120, - 6000, - 1.5, - "#2f9e44", - note="in-process, no HTTP hop", - ), +_DATA: tuple[Gateway, ...] = ( + Gateway("LiteLLM (Rust)", 2.0, 90, True), + Gateway("Bifrost", 3.0, 180, False), + Gateway("Portkey", 7.0, 260, False), + Gateway("LiteLLM (Python v1)", 12.0, 450, True), ) -def _style_axes(ax: "plt.Axes", title: str, ylabel: str, xlabel: str) -> None: - ax.set_title(title, loc="left", fontsize=15, fontweight="bold", color=_INK, pad=18) - ax.set_ylabel(ylabel, fontsize=11, color=_MUTED) - ax.set_xlabel(xlabel, fontsize=11, color=_MUTED) - for spine in ("top", "right"): +def _panel(ax: "plt.Axes", values: list[float], labels: list[str], colors: list[str], xlabel: str) -> None: + y = range(len(labels)) + ax.barh(list(y), values, color=colors, height=0.62, zorder=3) + ax.set_yticks(list(y)) + ax.set_yticklabels(labels, fontsize=11, color=_INK) + ax.invert_yaxis() + ax.set_xlabel(xlabel, fontsize=11, color=_MUTED, labelpad=10) + for spine in ("top", "right", "left"): ax.spines[spine].set_visible(False) - for spine in ("left", "bottom"): - ax.spines[spine].set_color(_GRID) - ax.tick_params(colors=_MUTED, labelsize=10) - ax.grid(True, color=_GRID, linewidth=0.8, zorder=0) + ax.spines["bottom"].set_color(_GRID) + ax.tick_params(axis="x", colors=_MUTED, labelsize=10) + ax.tick_params(axis="y", length=0) + ax.grid(True, axis="x", color=_GRID, linewidth=0.8, zorder=0) ax.set_axisbelow(True) + for i, v in enumerate(values): + ax.text(v + max(values) * 0.02, i, f"{v:g}", va="center", fontsize=10, color=_MUTED) -def _annotate(ax: "plt.Axes", p: Point, dx: float, dy: float) -> None: - text = p.label if not p.note else f"{p.label}" - ax.annotate( - text, - xy=(p.memory_mb, _y(ax, p)), - xytext=(p.memory_mb + dx, _y(ax, p) + dy), - fontsize=11, +def make_chart() -> Path: + ordered = sorted(_DATA, key=lambda g: g.added_latency_ms) + labels = [g.label for g in ordered] + colors = [_HILITE if g.is_litellm else _DARK for g in ordered] + + fig, (ax_l, ax_r) = plt.subplots(1, 2, figsize=(11.5, 5.2), dpi=130) + fig.patch.set_facecolor("white") + fig.suptitle( + "Gateway overhead, lower is better", + x=0.06, + y=1.02, + ha="left", + fontsize=16, + fontweight="bold", color=_INK, - fontweight="medium", ) - if p.note: - ax.annotate( - p.note, - xy=(p.memory_mb, _y(ax, p)), - xytext=(p.memory_mb + dx, _y(ax, p) + dy - _note_gap(ax)), - fontsize=8.5, - color=_MUTED, - style="italic", - ) - - -_MODE = {"metric": "throughput_rps"} - - -def _y(ax: "plt.Axes", p: Point) -> float: - return getattr(p, _MODE["metric"]) - - -def _note_gap(ax: "plt.Axes") -> float: - lo, hi = ax.get_ylim() - return (hi - lo) * 0.045 - - -def _frontier(ax: "plt.Axes", pts: list[Point], color: str) -> None: - xs = [p.memory_mb for p in pts] - ys = [_y(ax, p) for p in pts] - ax.plot(xs, ys, color=color, linewidth=1.6, alpha=0.55, zorder=2) - - -def _scatter(pts: tuple[Point, ...]) -> None: - for p in pts: - ax = plt.gca() - ax.scatter([p.memory_mb], [_y(ax, p)], s=90, color=p.color, zorder=3, edgecolors="white", linewidths=1.2) - - -def make_throughput_chart() -> Path: - _MODE["metric"] = "throughput_rps" - fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130) - fig.patch.set_facecolor("white") - ax.set_facecolor("white") - ax.set_xlim(80, 520) - ax.set_ylim(1000, 6800) - _style_axes( - ax, - "GatewayBench: throughput vs memory", - "sustained throughput (req/s) higher is better ->", - "peak resident memory (MB) <- lower is better", + _panel(ax_l, [g.added_latency_ms for g in ordered], labels, colors, "p99 added latency (ms)") + _panel( + ax_r, + [g.memory_mb for g in ordered], + ["" for _ in ordered], + colors, + "peak memory (MB)", ) - _scatter(_DATA) - frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb) - _frontier(ax, frontier_pts, "#0b7285") - offsets = { - "LiteLLM Proxy": (10, 120), - "Bifrost (Go)": (10, 180), - "Portkey (OSS)": (14, 150), - "LiteLLM SDK": (-40, -320), - } - for p in _DATA: - dx, dy = offsets[p.label] - _annotate(ax, p, dx, dy) - ax.annotate( - "efficiency frontier", - xy=(120, 6000), - xytext=(300, 6300), - fontsize=12, - color=_MUTED, - fontweight="bold", - ) - ax.annotate( + fig.text( + 0.99, + -0.04, "ILLUSTRATIVE PLACEHOLDER DATA -- not measured", - xy=(0, 0), - xytext=(0.99, -0.13), - xycoords="axes fraction", ha="right", fontsize=8.5, color="#b91c1c", style="italic", ) - out = _OUT_DIR / "throughput_vs_memory.png" - fig.tight_layout() - fig.savefig(out, bbox_inches="tight", facecolor="white") - plt.close(fig) - return out - - -def make_ttft_chart() -> Path: - _MODE["metric"] = "ttft_overhead_ms" - fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130) - fig.patch.set_facecolor("white") - ax.set_facecolor("white") - ax.set_xlim(80, 520) - ax.set_ylim(0, 14) - _style_axes( - ax, - "GatewayBench: streaming TTFT overhead vs memory", - "p99 added time-to-first-chunk (ms) <- lower is better", - "peak resident memory (MB) <- lower is better", - ) - _scatter(_DATA) - frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb) - _frontier(ax, frontier_pts, "#0b7285") - offsets = { - "LiteLLM Proxy": (10, 0.4), - "Bifrost (Go)": (10, 0.4), - "Portkey (OSS)": (10, 0.4), - "LiteLLM SDK": (-30, 0.6), - } - for p in _DATA: - dx, dy = offsets[p.label] - _annotate(ax, p, dx, dy) - ax.annotate( - "ideal: fast + lean (bottom-left)", - xy=(120, 1.5), - xytext=(150, 11.5), - fontsize=11, - color=_MUTED, - fontweight="bold", - ) - ax.annotate( - "ILLUSTRATIVE PLACEHOLDER DATA -- not measured", - xy=(0, 0), - xytext=(0.99, -0.13), - xycoords="axes fraction", - ha="right", - fontsize=8.5, - color="#b91c1c", - style="italic", - ) - out = _OUT_DIR / "ttft_vs_memory.png" - fig.tight_layout() + out = _OUT_DIR / "overhead_comparison.png" + fig.tight_layout(rect=(0, 0, 1, 0.98)) fig.savefig(out, bbox_inches="tight", facecolor="white") plt.close(fig) return out def main() -> None: - _ = font_manager # ensure font cache import for consistent rendering - t = make_throughput_chart() - s = make_ttft_chart() - print(f"wrote {t}") - print(f"wrote {s}") + out = make_chart() + print(f"wrote {out}") if __name__ == "__main__": diff --git a/gatewaybench/analyze/overhead_comparison.png b/gatewaybench/analyze/overhead_comparison.png new file mode 100644 index 00000000000..9fe994ac7de Binary files /dev/null and b/gatewaybench/analyze/overhead_comparison.png differ diff --git a/gatewaybench/analyze/throughput_vs_memory.png b/gatewaybench/analyze/throughput_vs_memory.png deleted file mode 100644 index af2cf537734..00000000000 Binary files a/gatewaybench/analyze/throughput_vs_memory.png and /dev/null differ diff --git a/gatewaybench/analyze/ttft_vs_memory.png b/gatewaybench/analyze/ttft_vs_memory.png deleted file mode 100644 index aca6c82f015..00000000000 Binary files a/gatewaybench/analyze/ttft_vs_memory.png and /dev/null differ diff --git a/gatewaybench/gateways/litellm-proxy/README.md b/gatewaybench/gateways/litellm-proxy/README.md deleted file mode 100644 index 6ea3db2b315..00000000000 --- a/gatewaybench/gateways/litellm-proxy/README.md +++ /dev/null @@ -1 +0,0 @@ -Placeholder for LiteLLM proxy benchmark configuration. diff --git a/gatewaybench/gateways/litellm-proxy/.gitkeep b/gatewaybench/gateways/litellm-python/.gitkeep similarity index 100% rename from gatewaybench/gateways/litellm-proxy/.gitkeep rename to gatewaybench/gateways/litellm-python/.gitkeep diff --git a/gatewaybench/gateways/litellm-python/README.md b/gatewaybench/gateways/litellm-python/README.md new file mode 100644 index 00000000000..54039936e12 --- /dev/null +++ b/gatewaybench/gateways/litellm-python/README.md @@ -0,0 +1 @@ +Placeholder for LiteLLM (Python v1) proxy benchmark configuration (bare + prod). diff --git a/gatewaybench/gateways/openrouter/.gitkeep b/gatewaybench/gateways/litellm-rust/.gitkeep similarity index 100% rename from gatewaybench/gateways/openrouter/.gitkeep rename to gatewaybench/gateways/litellm-rust/.gitkeep diff --git a/gatewaybench/gateways/litellm-rust/README.md b/gatewaybench/gateways/litellm-rust/README.md new file mode 100644 index 00000000000..6ed456bccad --- /dev/null +++ b/gatewaybench/gateways/litellm-rust/README.md @@ -0,0 +1 @@ +Placeholder for LiteLLM (Rust) gateway benchmark configuration (bare + prod). diff --git a/gatewaybench/gateways/openrouter/README.md b/gatewaybench/gateways/openrouter/README.md deleted file mode 100644 index dc79b80a95c..00000000000 --- a/gatewaybench/gateways/openrouter/README.md +++ /dev/null @@ -1 +0,0 @@ -Placeholder for OpenRouter benchmark configuration.