feat(gatewaybench): scaffold AI-gateway overhead benchmark

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
This commit is contained in:
Devin AI 2026-07-18 23:04:07 +00:00
parent 366ec6f487
commit 343979cd31
47 changed files with 2595 additions and 0 deletions

2
gatewaybench/.gitignore vendored Normal file
View file

@ -0,0 +1,2 @@
/target
results/*.jsonl

1921
gatewaybench/Cargo.lock generated Normal file

File diff suppressed because it is too large Load diff

3
gatewaybench/Cargo.toml Normal file
View file

@ -0,0 +1,3 @@
[workspace]
resolver = "2"
members = ["crates/*", "tests/*", "xtask"]

136
gatewaybench/README.md Normal file
View file

@ -0,0 +1,136 @@
# GatewayBench
A reproducible benchmark for **AI-gateway overhead**: the latency, throughput ceiling, and resource cost a gateway adds *on top of* the upstream LLM.
CursorBench measures model *quality*. The right question for a gateway is not quality but overhead, and specifically the overhead that a coding agent feels on every turn: how much longer until the first token, how evenly the stream flows, how tool calls behave, and how the whole thing holds up under concurrency. GatewayBench measures that, and it measures it the same way for every gateway so the numbers are comparable.
![GatewayBench throughput vs memory](analyze/throughput_vs_memory.png)
![GatewayBench streaming TTFT overhead vs memory](analyze/ttft_vs_memory.png)
> The charts above use **illustrative placeholder data** to show the design. They are regenerated from real `results/*.jsonl` once the harness has run; see [Generating the charts](#generating-the-charts).
## The core idea: isolate the gateway from the provider
Real provider latency is hundreds of milliseconds to several seconds with large variance. Gateway overhead is microseconds to low-single-digit milliseconds. Benchmarking against a live provider drowns the signal in provider noise and nothing reproduces.
So every self-hostable gateway points at a **local mock upstream** we control (`crates/mock-upstream`): an OpenAI-compatible server with deterministic, configurable behavior
- fixed, configurable time-to-first-byte and inter-token delay, so streaming is realistic without real-world variance
- configurable request size (prompt tokens) and response size (completion tokens)
- correct SSE streaming for `stream: true`, a normal JSON body otherwise
- a valid `usage` block, so spend-tracking and logging code paths actually run
Overhead is then
```
overhead = latency(client -> gateway -> mock) - latency(client -> mock directly)
```
measured through the same client over the same loopback. The direct-to-mock path is the zero point.
## Why Rust
The mock and the load driver must never be the bottleneck. A Python mock or closed-loop Python driver caps throughput well below where these gateways saturate, and it cannot measure tail latency honestly. The harness is a Rust (tokio) workspace so that
- the mock and driver stay far above the gateways' saturation point
- latency is captured losslessly with `hdrhistogram`
- the driver is **open-loop** (constant arrival rate), which avoids coordinated omission and keeps the p99/p99.9 tail honest
- runs are deterministic: fixed seeds, fixed payloads, a discarded warmup window
Only the harness is Rust. The gateways run as their real selves: LiteLLM (Python), Bifrost (Go), Portkey (the OSS JS gateway).
## Metrics
Coding agents stream, call tools, and send large contexts, so the streaming and payload metrics carry the most weight.
Streaming, what an agent feels every turn
- **TTFT overhead** - added time to the first chunk. The single most perceived metric
- **Inter-chunk latency (ITL) and jitter** - added gap between SSE chunks (p50/p99/max and standard deviation). This exposes gateways that buffer or re-chunk the stream; some coalesce the whole response and re-emit it, which is catastrophic for an agent
- **Chunk fidelity** - whether chunks pass through 1:1 or get coalesced (chunk-count ratio versus the mock)
- **Time-to-last-token** - added total stream duration
Tool calls, core to coding agents
- **Tool-call TTFT and total overhead** versus plain text; streamed `tool_calls` argument deltas are a heavier transform path than text
- **Tool-call correctness under streaming** - whether the gateway reassembles partial JSON argument deltas without corrupting or reordering them
Payload scaling, agents send whole files
- **Large-prompt overhead** - overhead as a function of input size (1k / 10k / 100k tokens), where JSON parse and serialize cost dominates
Load and cost
- **Throughput ceiling** - max sustained req/s before the latency knee or errors
- **Tail latency** - p99 / p99.9 added latency under concurrency
- **Resource footprint** - peak RSS and CPU at a fixed req/s, from which we derive cost per 1M requests
## Fairness
Since LiteLLM publishes this, transparency is the whole point. Every gateway config, image, and pinned version lives in `gateways/` and is meant to be challenged. Rules
- each gateway runs in its **recommended production config**, not a strawman, and we run two configs per gateway: *bare passthrough* (no logging, DB, cache, or rate limit) to isolate pure proxy cost, and *realistic prod* (key auth plus a logging or spend callback on) to match what people actually run
- equal resources for every gateway (same container CPU and memory limits), so a result is never "who got more cores"
- load driver, gateway, and mock run on separate pinned cores so the driver never steals the gateway's CPU
- fixed seeds and payloads, a discarded warmup window, multiple runs with median-of-runs and variance reported
Category differences we state honestly rather than hide
- **LiteLLM SDK** is an in-process library with no HTTP hop, so it is a different product category from the network proxies. It is reported but kept off the same latency axis
- **OpenRouter** is SaaS only and cannot be pointed at the mock. Any number for it includes their infrastructure, the network round trip to their datacenter, and a real upstream, so it is not comparable to the self-hosted numbers. It is either excluded from the isolated-overhead charts or reported separately as an end-to-end measurement, always clearly labeled
- **Portkey** is benchmarked as the open-source self-hostable gateway, not the SaaS
## Layout
Each folder under `tests/` is one self-contained test: its own crate, its own driver invocation, its own assertion.
```
gatewaybench/ cargo workspace
crates/
gwbench-core/ shared types: Scenario, tagged-union Outcome,
hdrhistogram-backed LatencySummary, ResourceSample,
JSONL BenchResult writer
mock-upstream/ axum mock: configurable TTFT, inter-token gap, size, SSE, usage
driver/ open-loop (constant-arrival-rate) load engine
tests/ each subdir = ONE test (bin + README describing it)
ttft/ added time-to-first-chunk
inter-chunk-latency/ added inter-chunk gap + jitter; buffering detection
chunk-fidelity/ 1:1 passthrough vs coalescing
tool-call-latency/ tool-call TTFT + streamed-arg reassembly correctness
large-prompt/ overhead vs input size (1k/10k/100k tokens)
throughput/ RPS ceiling + CPU/RSS cost per 1M requests
tail-latency/ p99/p99.9 added latency under concurrency
gateways/ per-gateway configs (bare + prod), pinned versions
litellm-proxy/ bifrost/ portkey/ openrouter/
xtask/ `cargo xtask bench` sweeps the matrix -> results/*.jsonl
analyze/ chart generation for RESULTS.md
```
## Quickstart
```bash
# build the workspace
cargo build --workspace --release
# run the mock upstream (defaults to :8080)
GWBENCH_MOCK_PORT=8080 cargo run --release -p mock-upstream
# run a single test against a target (e.g. a gateway on :4000, or the mock directly for the baseline)
cargo run --release -p ttft
# sweep the full matrix and write results/*.jsonl
cargo xtask bench
```
## Generating the charts
The hero charts are produced from the benchmark results by `analyze/make_hero_charts.py`. Point `_DATA` at real rows from `results/*.jsonl` and regenerate
```bash
python analyze/make_hero_charts.py
```
## Status
Scaffold: the workspace compiles, the domain types are real, and the mock serves a valid non-streaming response. The measurement internals (open-loop driver, streaming mock knobs, cgroup resource sampling, per-test assertions, the `xtask` matrix sweep) land in follow-ups. The charts show placeholder data until then.

View file

View file

@ -0,0 +1,222 @@
"""Generate the GatewayBench hero charts (CursorBench-styled scatter plots).
The numbers here are ILLUSTRATIVE PLACEHOLDERS to convey the chart design; they
are not measured. Replace `_DATA` with real rows from results/*.jsonl once the
harness has run, then regenerate:
python analyze/make_hero_charts.py
Outputs analyze/throughput_vs_memory.png and analyze/ttft_vs_memory.png.
"""
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib import font_manager
_OUT_DIR = Path(__file__).resolve().parent
_INK = "#1f2933"
_MUTED = "#6b7280"
_GRID = "#e5e7eb"
@dataclass(frozen=True)
class Point:
label: str
memory_mb: float
throughput_rps: float
ttft_overhead_ms: float
color: str
note: str = ""
# ILLUSTRATIVE PLACEHOLDER DATA -- not measured.
_DATA: tuple[Point, ...] = (
Point("LiteLLM Proxy", 450, 1800, 12.0, "#0b7285"),
Point("Bifrost (Go)", 180, 4200, 3.0, "#e8590c"),
Point("Portkey (OSS)", 260, 2600, 7.0, "#ae3ec9"),
Point(
"LiteLLM SDK",
120,
6000,
1.5,
"#2f9e44",
note="in-process, no HTTP hop",
),
)
def _style_axes(ax: "plt.Axes", title: str, ylabel: str, xlabel: str) -> None:
ax.set_title(title, loc="left", fontsize=15, fontweight="bold", color=_INK, pad=18)
ax.set_ylabel(ylabel, fontsize=11, color=_MUTED)
ax.set_xlabel(xlabel, fontsize=11, color=_MUTED)
for spine in ("top", "right"):
ax.spines[spine].set_visible(False)
for spine in ("left", "bottom"):
ax.spines[spine].set_color(_GRID)
ax.tick_params(colors=_MUTED, labelsize=10)
ax.grid(True, color=_GRID, linewidth=0.8, zorder=0)
ax.set_axisbelow(True)
def _annotate(ax: "plt.Axes", p: Point, dx: float, dy: float) -> None:
text = p.label if not p.note else f"{p.label}"
ax.annotate(
text,
xy=(p.memory_mb, _y(ax, p)),
xytext=(p.memory_mb + dx, _y(ax, p) + dy),
fontsize=11,
color=_INK,
fontweight="medium",
)
if p.note:
ax.annotate(
p.note,
xy=(p.memory_mb, _y(ax, p)),
xytext=(p.memory_mb + dx, _y(ax, p) + dy - _note_gap(ax)),
fontsize=8.5,
color=_MUTED,
style="italic",
)
_MODE = {"metric": "throughput_rps"}
def _y(ax: "plt.Axes", p: Point) -> float:
return getattr(p, _MODE["metric"])
def _note_gap(ax: "plt.Axes") -> float:
lo, hi = ax.get_ylim()
return (hi - lo) * 0.045
def _frontier(ax: "plt.Axes", pts: list[Point], color: str) -> None:
xs = [p.memory_mb for p in pts]
ys = [_y(ax, p) for p in pts]
ax.plot(xs, ys, color=color, linewidth=1.6, alpha=0.55, zorder=2)
def _scatter(pts: tuple[Point, ...]) -> None:
for p in pts:
ax = plt.gca()
ax.scatter([p.memory_mb], [_y(ax, p)], s=90, color=p.color, zorder=3, edgecolors="white", linewidths=1.2)
def make_throughput_chart() -> Path:
_MODE["metric"] = "throughput_rps"
fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130)
fig.patch.set_facecolor("white")
ax.set_facecolor("white")
ax.set_xlim(80, 520)
ax.set_ylim(1000, 6800)
_style_axes(
ax,
"GatewayBench: throughput vs memory",
"sustained throughput (req/s) higher is better ->",
"peak resident memory (MB) <- lower is better",
)
_scatter(_DATA)
frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb)
_frontier(ax, frontier_pts, "#0b7285")
offsets = {
"LiteLLM Proxy": (10, 120),
"Bifrost (Go)": (10, 180),
"Portkey (OSS)": (14, 150),
"LiteLLM SDK": (-40, -320),
}
for p in _DATA:
dx, dy = offsets[p.label]
_annotate(ax, p, dx, dy)
ax.annotate(
"efficiency frontier",
xy=(120, 6000),
xytext=(300, 6300),
fontsize=12,
color=_MUTED,
fontweight="bold",
)
ax.annotate(
"ILLUSTRATIVE PLACEHOLDER DATA -- not measured",
xy=(0, 0),
xytext=(0.99, -0.13),
xycoords="axes fraction",
ha="right",
fontsize=8.5,
color="#b91c1c",
style="italic",
)
out = _OUT_DIR / "throughput_vs_memory.png"
fig.tight_layout()
fig.savefig(out, bbox_inches="tight", facecolor="white")
plt.close(fig)
return out
def make_ttft_chart() -> Path:
_MODE["metric"] = "ttft_overhead_ms"
fig, ax = plt.subplots(figsize=(9.2, 6.4), dpi=130)
fig.patch.set_facecolor("white")
ax.set_facecolor("white")
ax.set_xlim(80, 520)
ax.set_ylim(0, 14)
_style_axes(
ax,
"GatewayBench: streaming TTFT overhead vs memory",
"p99 added time-to-first-chunk (ms) <- lower is better",
"peak resident memory (MB) <- lower is better",
)
_scatter(_DATA)
frontier_pts = sorted(_DATA, key=lambda p: p.memory_mb)
_frontier(ax, frontier_pts, "#0b7285")
offsets = {
"LiteLLM Proxy": (10, 0.4),
"Bifrost (Go)": (10, 0.4),
"Portkey (OSS)": (10, 0.4),
"LiteLLM SDK": (-30, 0.6),
}
for p in _DATA:
dx, dy = offsets[p.label]
_annotate(ax, p, dx, dy)
ax.annotate(
"ideal: fast + lean (bottom-left)",
xy=(120, 1.5),
xytext=(150, 11.5),
fontsize=11,
color=_MUTED,
fontweight="bold",
)
ax.annotate(
"ILLUSTRATIVE PLACEHOLDER DATA -- not measured",
xy=(0, 0),
xytext=(0.99, -0.13),
xycoords="axes fraction",
ha="right",
fontsize=8.5,
color="#b91c1c",
style="italic",
)
out = _OUT_DIR / "ttft_vs_memory.png"
fig.tight_layout()
fig.savefig(out, bbox_inches="tight", facecolor="white")
plt.close(fig)
return out
def main() -> None:
_ = font_manager # ensure font cache import for consistent rendering
t = make_throughput_chart()
s = make_ttft_chart()
print(f"wrote {t}")
print(f"wrote {s}")
if __name__ == "__main__":
main()

Binary file not shown.

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 85 KiB

View file

@ -0,0 +1,10 @@
[package]
name = "driver"
version = "0.1.0"
edition = "2021"
[dependencies]
gwbench-core = { version = "0.1.0", path = "../gwbench-core" }
hdrhistogram = "7.6.0"
reqwest = "0.13.4"
tokio = { version = "1.53.0", features = ["rt-multi-thread", "macros"] }

View file

@ -0,0 +1,27 @@
#![forbid(unsafe_code)]
use std::fmt;
use gwbench_core::{LatencySummary, Scenario};
use hdrhistogram::Histogram;
#[derive(Debug)]
pub struct DriverError {
message: String,
}
impl fmt::Display for DriverError {
fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result {
formatter.write_str(&self.message)
}
}
impl std::error::Error for DriverError {}
/// Runs a stub open-loop constant-arrival-rate engine against `target`.
pub async fn run(_scenario: &Scenario, _target: &str) -> Result<LatencySummary, DriverError> {
let histogram = Histogram::<u64>::new(3).map_err(|error| DriverError {
message: error.to_string(),
})?;
Ok(LatencySummary::from_histogram(&histogram))
}

View file

@ -0,0 +1,6 @@
#![forbid(unsafe_code)]
#[tokio::main]
async fn main() {
println!("driver stub: open-loop load engine is not implemented");
}

View file

@ -0,0 +1,10 @@
[package]
name = "gwbench-core"
version = "0.1.0"
edition = "2021"
[dependencies]
hdrhistogram = "7.6.0"
serde = { version = "1.0.228", features = ["derive"] }
serde_json = "1.0.150"
thiserror = "2.0.19"

View file

@ -0,0 +1,122 @@
#![forbid(unsafe_code)]
use std::io::Write;
use std::time::Duration;
use hdrhistogram::Histogram;
use serde::{Deserialize, Serialize};
use thiserror::Error;
#[derive(Debug, Error)]
pub enum CoreError {
#[error("failed to serialize benchmark result: {0}")]
Serialization(#[from] serde_json::Error),
#[error("failed to write benchmark result: {0}")]
Io(#[from] std::io::Error),
}
mod duration_millis {
use serde::{Deserialize, Deserializer, Serializer};
use std::time::Duration;
pub fn serialize<S>(duration: &Duration, serializer: S) -> Result<S::Ok, S::Error>
where
S: Serializer,
{
serializer.serialize_u64(duration.as_millis() as u64)
}
pub fn deserialize<'de, D>(deserializer: D) -> Result<Duration, D::Error>
where
D: Deserializer<'de>,
{
let millis = u64::deserialize(deserializer)?;
Ok(Duration::from_millis(millis))
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Scenario {
pub name: String,
pub target_rps: u32,
pub concurrency: u32,
pub streaming: bool,
pub prompt_tokens: u32,
pub completion_tokens: u32,
#[serde(with = "duration_millis")]
pub duration: Duration,
#[serde(with = "duration_millis")]
pub warmup: Duration,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(tag = "kind")]
pub enum Outcome {
Success { status: u16 },
Timeout,
TransportError { message: String },
UpstreamError { status: u16 },
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct LatencySummary {
pub count: u64,
pub p50_ms: f64,
pub p90_ms: f64,
pub p99_ms: f64,
pub p999_ms: f64,
pub max_ms: f64,
}
impl LatencySummary {
/// Converts histogram values recorded in microseconds to millisecond percentiles.
pub fn from_histogram(histogram: &Histogram<u64>) -> Self {
let micros_to_millis = |micros: u64| micros as f64 / 1_000.0;
Self {
count: histogram.len(),
p50_ms: micros_to_millis(histogram.value_at_quantile(0.50)),
p90_ms: micros_to_millis(histogram.value_at_quantile(0.90)),
p99_ms: micros_to_millis(histogram.value_at_quantile(0.99)),
p999_ms: micros_to_millis(histogram.value_at_quantile(0.999)),
max_ms: micros_to_millis(histogram.max()),
}
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ResourceSample {
pub rss_bytes: u64,
pub cpu_millis: u64,
pub ts: u64,
}
impl ResourceSample {
/// Stub resource sampler; real `/proc` and cgroup sampling will land later.
pub fn sample() -> Result<Self, CoreError> {
let ts = match std::time::SystemTime::now().duration_since(std::time::UNIX_EPOCH) {
Ok(duration) => duration.as_millis() as u64,
Err(_) => 0,
};
Ok(Self {
rss_bytes: 0,
cpu_millis: 0,
ts,
})
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct BenchResult {
pub gateway: String,
pub scenario: Scenario,
pub latency: LatencySummary,
pub resources: ResourceSample,
}
impl BenchResult {
pub fn append_json_line<W: Write>(&self, writer: &mut W) -> Result<(), CoreError> {
serde_json::to_writer(&mut *writer, self)?;
writer.write_all(b"\n")?;
Ok(())
}
}

View file

@ -0,0 +1,9 @@
[package]
name = "mock-upstream"
version = "0.1.0"
edition = "2021"
[dependencies]
axum = "0.8.9"
serde_json = "1.0.150"
tokio = { version = "1.53.0", features = ["rt-multi-thread", "macros", "net"] }

View file

@ -0,0 +1,37 @@
#![forbid(unsafe_code)]
use axum::{routing::post, Json, Router};
use serde_json::{json, Value};
async fn chat_completions() -> Json<Value> {
Json(json!({
"id": "chatcmpl-gwbench",
"object": "chat.completion",
"created": 0,
"model": "mock-upstream",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "stub response"},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 1,
"completion_tokens": 2,
"total_tokens": 3
}
}))
}
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let port = std::env::var("GWBENCH_MOCK_PORT")
.ok()
.and_then(|value| value.parse::<u16>().ok())
.unwrap_or(8080);
let app = Router::new().route("/v1/chat/completions", post(chat_completions));
let listener = tokio::net::TcpListener::bind(("0.0.0.0", port)).await?;
println!("mock upstream listening on {listener:?}");
// Streaming and configurable response timing are intentionally deferred.
axum::serve(listener, app).await?;
Ok(())
}

View file

View file

@ -0,0 +1 @@
Placeholder for Bifrost benchmark configuration.

View file

@ -0,0 +1 @@
Placeholder for LiteLLM proxy benchmark configuration.

View file

@ -0,0 +1 @@
Placeholder for OpenRouter benchmark configuration.

View file

View file

@ -0,0 +1 @@
Placeholder for Portkey benchmark configuration.

View file

@ -0,0 +1,2 @@
[toolchain]
channel = "stable"

View file

@ -0,0 +1,6 @@
[package]
name = "test-chunk-fidelity"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures whether a gateway passes server-sent event chunks one-to-one or coalesces them. It reports the gateway-to-upstream chunk-count ratio.

View file

@ -0,0 +1,3 @@
fn main() {
println!("chunk-fidelity: SSE chunk-count ratio vs upstream");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-inter-chunk-latency"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures added inter-chunk gap and jitter at p50, p99, and max. It detects gateways that buffer or re-chunk server-sent events.

View file

@ -0,0 +1,3 @@
fn main() {
println!("inter-chunk-latency: added inter-chunk gap and jitter");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-large-prompt"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures gateway overhead as input size grows. It compares scenarios using 1k, 10k, and 100k input tokens.

View file

@ -0,0 +1,3 @@
fn main() {
println!("large-prompt: overhead across 1k, 10k, and 100k token inputs");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-tail-latency"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures p99 and p99.9 added latency under concurrency. It is intended to expose gateway behavior in the latency tail.

View file

@ -0,0 +1,3 @@
fn main() {
println!("tail-latency: p99 and p99.9 added latency under concurrency");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-throughput"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures the maximum sustained RPS before the latency knee. It also measures CPU and RSS cost per one million requests.

View file

@ -0,0 +1,3 @@
fn main() {
println!("throughput: sustained RPS latency knee and resource cost");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-tool-call-latency"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures tool-call streaming TTFT and total overhead. It also checks correctness when reassembling streamed `tool_call` argument deltas.

View file

@ -0,0 +1,3 @@
fn main() {
println!("tool-call-latency: tool-call streaming TTFT and total overhead");
}

View file

@ -0,0 +1,6 @@
[package]
name = "test-ttft"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1 @@
This test measures added time-to-first-chunk (streaming TTFT) overhead compared with a direct-to-mock baseline. It will report gateway overhead for the first streamed response chunk.

View file

@ -0,0 +1,3 @@
fn main() {
println!("ttft: added time-to-first-chunk overhead vs direct-to-mock baseline");
}

View file

@ -0,0 +1,6 @@
[package]
name = "xtask"
version = "0.1.0"
edition = "2021"
[dependencies]

View file

@ -0,0 +1,8 @@
#![forbid(unsafe_code)]
fn main() {
match std::env::args().nth(1).as_deref() {
Some("bench") => println!("would sweep the matrix"),
_ => println!("usage: cargo xtask bench"),
}
}