test(e2e): add scripted-provider cost calculation suite

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This commit is contained in:
kerry 2026-09-15 23:03:56 +00:00
parent a6127d2363
commit 61ac4f5739
12 changed files with 1980 additions and 2 deletions

View file

@ -21,6 +21,7 @@ Each subdirectory under `tests/e2e/` is one suite, scoped to an endpoint family
- `load/` - performance-category tests, kept OUT of the main suite: throughput/load SLO tests are a different testing category from functional e2e (variance-driven, historically flaky) and live outside this suite until re-implemented as their own pipeline (LIT-5163); do not add a live load test that runs in the default collection. What lives here: the weekly session-anomaly test (`test_weekly_session_anomaly_e2e.py`, Claude Code-shaped multi-turn sessions against real providers with ceilings on error rate, cache read/write, turn time, and spend; marked `weekly` and deselected unless `E2E_WEEKLY_ANOMALY` is set, driven by `.github/workflows/weekly_load_anomaly.yml`), the Redis chaos test (`test_redis_chaos_e2e.py`, locust load against mock deployments split round robin over `/chat/completions` and `/v1/messages`, one endpoint per simulated user, with `CLIENT PAUSE ALL` on the proxy's Redis mid-run to simulate it being down outright, asserting zero failed requests on every endpoint, budgeting RSS and CPU-per-request as ratios against the same run's healthy phase, and holding p50/p90/p99 latency and log-bytes-per-request to flat ceilings (a ratio cannot bound those two: an open breaker skips Redis instead of waiting on it, so the chaos phase can measure cheaper than baseline while still being far slower than a user should see); needs a proxy booted from `gateway/redis_chaos_ci_config.yml` on the same host with `E2E_PROXY_PID` and `E2E_PROXY_LOG` set, marked `redis_chaos`, deselected unless `E2E_REDIS_CHAOS` is set and excluded from the per-PR selector like the rest of `load/`, driven by `.github/workflows/test-e2e-redis-chaos.yml` and by the Buildkite `e2e-redis-chaos` step in project-releaser, which runs the proxy, Postgres and Valkey co-located with pytest in one pod and sets the opt-in), and markerless harness unit tests for the locust, process-usage, and session-anomaly aggregation logic
- `other/` - the holding-pen suite for the `other.*` registry cluster with no home of its own yet: the master-key auth gate, JWT auth (access tokens issued by a real Keycloak realm, `idp.py` plus `idp_realm.json`, whose JWKS the proxy's `JWT_PUBLIC_KEY_URL` points at; see CONTRIBUTING.md for the start command and config block), and the process-lifecycle health probes (liveness, public readiness, authenticated readiness diagnostics). Promote a cluster out once it is large/stable enough for its own suite
- `gateway/` - proxy configuration only (`litellm-config.yml`); no tests
- `cost_calculation/` - cost accounting against a dedicated proxy whose whole model cost map is the test-owned `tests/e2e/cost_map.json` (loaded via `LITELLM_MODEL_COST_MAP_URL`), with provider calls answered by the scripted-provider sidecar in `scripted_provider.py`; asserts literal rate arithmetic on scripted usage across every provider wire and pricing component, deselected unless `E2E_COST_MAP_STACK` is set
- `claude_code/` - the Claude Code compatibility matrix: drives the real `claude` CLI (and HTTP probes) against a proxy for each feature x provider cell, reporting tagged-union outcomes via the `compat_result` fixture; ships its own driver/builder/publisher plus `_*_unit_tests/` trees. The HTTP probes ride the shared transport (`ProxyClient.count_tokens` / `ProxyClient.messages`); the CLI-driving path stays bespoke
- `ui/` - the Admin UI browser suite: Playwright in TypeScript, driving the dashboard served by a live proxy on port 4000 (seeded postgres + mock LLM upstream; see its `run_e2e.sh`). It is a self-contained npm package with its own lockfile and does not use the Python harness, pytest markers, or the shared transport; the Python rules in this file (typed models, `Result` unions, basedpyright zero-error gate) do not apply inside it. Its only Python file, `fixtures/mock_llm_server/server.py`, is excluded from the e2e basedpyright gate via the root `pyrightconfig.json`
@ -221,7 +222,7 @@ other.<area>.<case>.<assertion>
```
## Hard Rules
- no unit tests of any kind under `tests/e2e`. a product feature is proven end to end against a live proxy, never with a unit test, and the harness itself is not unit-tested here either. no monkeypatching or mock tests. if a contributor asks you to write an end to end test, do NOT stage a unit test with it; if you find a product gap, call it out in the PR description
- no unit tests of any kind under `tests/e2e`. a product feature is proven end to end against a live proxy, never with a unit test, and the harness itself is not unit-tested here either. no monkeypatching or mock tests; the one carve-out is a scripted upstream served through a real HTTP sidecar (the cost_calculation suite's scripted provider), allowed because provider-response-shape coverage needs a controlled usage payload and every hop from the proxy's upstream call to the spend row still executes for real. if a contributor asks you to write an end to end test, do NOT stage a unit test with it; if you find a product gap, call it out in the PR description
- use model management endpoints to create new models for a test. this could be in a conftest / inline for each test. ask the user what they want.

View file

@ -22,9 +22,9 @@ from typing import Final
import pytest
import requests
from e2e_config import (
CONTROL_PLANE_BASE_URL,
COST_MAP_OPT_IN_ENV,
FIXTURE_DIR,
FIXTURE_MODE_RAW,
MANAGED_FILES_OPT_IN_ENV,
@ -53,6 +53,7 @@ OPT_IN_MARKERS: Final = MappingProxyType(
"managed_files": MANAGED_FILES_OPT_IN_ENV,
"prompt_caching_stack": PROMPT_CACHING_OPT_IN_ENV,
"redis_chaos": REDIS_CHAOS_OPT_IN_ENV,
"cost_map_stack": COST_MAP_OPT_IN_ENV,
}
)
@ -120,6 +121,12 @@ def pytest_configure(config: pytest.Config) -> None:
"redis_chaos: load test that pauses the proxy's Redis outright mid-run; needs a proxy booted from "
"gateway/redis_chaos_ci_config.yml on the same host, and is deselected unless E2E_REDIS_CHAOS is set",
)
config.addinivalue_line(
"markers",
"cost_map_stack: needs a proxy whose whole cost map is tests/e2e/cost_map.json "
"(LITELLM_MODEL_COST_MAP_URL) plus a scripted-provider sidecar; deselected unless "
"E2E_COST_MAP_STACK is set",
)
def pytest_sessionstart(session: pytest.Session) -> None:

View file

@ -0,0 +1,139 @@
"""Cost-calculation suite fixtures.
Runs against a dedicated proxy whose whole model cost map is the test-owned
``tests/e2e/cost_map.json`` (LITELLM_MODEL_COST_MAP_URL), so every deployment
bills at rates the test asserts literal arithmetic on. Provider calls are
answered by the scripted-provider sidecar (``scripted_provider.py``), registered
per scenario over its control API.
Deselected unless E2E_COST_MAP_STACK is set (marker `cost_map_stack`).
"""
from __future__ import annotations
import importlib.util
import sys
from collections.abc import Callable
from dataclasses import dataclass
from pathlib import Path
from types import ModuleType
from typing import Final, Protocol, cast
import pytest
from cost_matrix import Case, FrontierModel
from e2e_config import COST_MAP_PROXY_URL
from lifecycle import ResourceManager
from models import LiteLLMParamsBody, ModelInfoBody, ModelNewBody
from proxy_client import ProxyClient, build_proxy_client
from scripted_client import ScenarioHandle, delete_scenario, register_scenario
from scripted_provider import Scenario
def _load_cost_rows() -> ModuleType:
"""Load quota_management/spend_tracking/cost_rows.py by path (the e2e tree
has no package layout), the same trick the mcp suite uses for
logging/datadog_reader.py."""
path = (
Path(__file__).resolve().parent.parent
/ "quota_management"
/ "spend_tracking"
/ "cost_rows.py"
)
name = "e2e_spend_tracking_cost_rows"
spec = importlib.util.spec_from_file_location(name, path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
sys.modules[name] = module
spec.loader.exec_module(module)
return module
class SpendCostBreakdown(Protocol):
input_cost: float | None
output_cost: float | None
cache_read_cost: float | None
cache_creation_cost: float | None
reasoning_cost: float | None
tool_usage_cost: float | None
total_cost: float | None
service_tier: str | None
def model_dump(self) -> dict[str, object]: ...
class SpendRowMetadata(Protocol):
cost_breakdown: SpendCostBreakdown | None
class SpendCostRow(Protocol):
"""The slice of spend_tracking.cost_rows.CostRow this suite reads."""
spend: float | None
prompt_tokens: int | None
completion_tokens: int | None
metadata: SpendRowMetadata | None
@property
def breakdown(self) -> SpendCostBreakdown: ...
class CostRowsModule(Protocol):
"""cost_rows.py loaded by path has no importable name for basedpyright, so
its surface is declared here and reached through a single cast."""
approx_equal: Callable[[float, float], bool]
assert_total_is_sum_of_components: Callable[[SpendCostRow], None]
poll_cost_row_where: Callable[
[ProxyClient, str, Callable[[SpendCostRow], bool]], SpendCostRow | None
]
cost_rows: Final[CostRowsModule] = cast(CostRowsModule, _load_cost_rows())
@dataclass(frozen=True, slots=True)
class CostCalcClient:
"""The suite's client: a ProxyClient pointed at the cost-map proxy pod."""
proxy: ProxyClient
@pytest.fixture(scope="session")
def client() -> CostCalcClient:
proxy = build_proxy_client(
base_url=COST_MAP_PROXY_URL,
control_plane_base_url=COST_MAP_PROXY_URL,
replica_urls=(COST_MAP_PROXY_URL,),
)
return CostCalcClient(proxy=proxy)
def register_scenario_deployment(
client: CostCalcClient,
resources: ResourceManager,
model: FrontierModel,
case: Case,
marker: str,
) -> tuple[str, ScenarioHandle]:
"""Register the case's scenario on the sidecar plus a deployment pointed at
it; both are torn down by ``resources``. Returns the callable model_name."""
scenario: Scenario = case.scenario(
scenario_id=f"sc-{marker}", model=model, text=f"scripted answer {marker}"
)
handle = register_scenario(scenario)
resources.defer(lambda: delete_scenario(handle))
model_name = f"{model.model_name}-{marker}"
model_id = client.proxy.register_model(
ModelNewBody(
model_name=model_name,
litellm_params=LiteLLMParamsBody(
model=model.litellm_model,
api_key="sk-scripted-provider",
api_base=handle.api_base(),
),
model_info=ModelInfoBody(),
)
)
resources.defer(lambda: client.proxy.delete_model(model_id))
return model_name, handle

View file

@ -0,0 +1,458 @@
"""The cost-calculation matrix: frontier model set, the pricing-component cases
each model runs, and the expected-cost arithmetic.
Rates come from ``tests/e2e/cost_map.json``, which the proxy under test loads as
its ENTIRE model cost map (LITELLM_MODEL_COST_MAP_URL), so an entry's rates are
exactly what the proxy bills and nothing in the suite depends on the bundled
map. Each model's rates are a distinct multiple of a shared base set, so a
component billed at the wrong model's rate (or the wrong case's rate) can never
coincidentally match.
Case applicability is pricing-field-gated AND wire-gated: a case runs for a
model only when the entry carries the rate the case exercises and the wire can
report the token kind that rate prices. When the wire cannot report a kind
(e.g. Anthropic has no reasoning-token field, Responses reports no cache
creation), the case is absent from the matrix rather than silently zero.
"""
from __future__ import annotations
import json
from dataclasses import dataclass
from pathlib import Path
from typing import Final, Literal
from pydantic import BaseModel, ConfigDict, TypeAdapter
from scripted_provider import Scenario, ScriptedOutput, ScriptedUsage, Wire
COST_MAP_PATH: Final = Path(__file__).resolve().parent.parent / "cost_map.json"
class SearchContextCostPerQuery(BaseModel):
model_config = ConfigDict(frozen=True)
search_context_size_low: float | None = None
search_context_size_medium: float | None = None
search_context_size_high: float | None = None
class CostMapEntry(BaseModel):
"""The pricing fields of a cost-map entry the matrix reads. Shaped like a
``model_prices_and_context_window.json`` entry; unmodelled keys are ignored."""
model_config = ConfigDict(frozen=True, extra="ignore")
litellm_provider: str
mode: str
input_cost_per_token: float | None = None
output_cost_per_token: float | None = None
cache_read_input_token_cost: float | None = None
cache_creation_input_token_cost: float | None = None
cache_creation_input_token_cost_above_1hr: float | None = None
output_cost_per_reasoning_token: float | None = None
input_cost_per_audio_token: float | None = None
output_cost_per_audio_token: float | None = None
input_cost_per_token_above_200k_tokens: float | None = None
output_cost_per_token_above_200k_tokens: float | None = None
input_cost_per_token_flex: float | None = None
output_cost_per_token_flex: float | None = None
input_cost_per_token_priority: float | None = None
output_cost_per_token_priority: float | None = None
search_context_cost_per_query: SearchContextCostPerQuery | None = None
web_search_billing_unit: str | None = None
_COST_MAP_ADAPTER: Final = TypeAdapter(dict[str, CostMapEntry])
_COST_MAP: Final[dict[str, CostMapEntry]] = _COST_MAP_ADAPTER.validate_python(
json.loads(COST_MAP_PATH.read_text())
)
TIER_THRESHOLD_TOKENS: Final = 200_000
@dataclass(frozen=True, slots=True)
class FrontierModel:
"""One deployment under test: the model_name the suite registers, the
provider-prefixed litellm model string, the wire the scripted upstream
speaks, its cost-map key, and the sibling map model the response_model
override case reports."""
model_name: str
litellm_model: str
wire: Wire
map_key: str
override_model: str
@property
def rates(self) -> CostMapEntry:
return _COST_MAP[self.map_key]
@property
def override_rates(self) -> CostMapEntry:
return _COST_MAP[self.override_map_key]
@property
def override_map_key(self) -> str:
return _OVERRIDE_MAP_KEYS[self.override_model]
@property
def provider(self) -> str:
return self.rates.litellm_provider
@property
def api_key(self) -> str:
# The scripted upstream ignores auth; a fixed bogus key proves the suite
# spends zero real provider calls.
return "sk-scripted-provider"
# Response-model override targets: emit a sibling's bare provider-facing name so
# the biller's provider-prefixed lookup lands on that sibling's map key.
_OVERRIDE_MODELS: Final[dict[str, str]] = {
"gpt-5.6": "gpt-5.4-mini",
"gpt-5.5-pro": "gpt-5.3-codex",
"gpt-5.3-codex": "gpt-5.5-pro",
"gpt-5.4-mini": "gpt-5.6",
"claude-opus-5": "claude-sonnet-5",
"claude-sonnet-5": "claude-opus-5",
"claude-haiku-4-5": "claude-sonnet-5",
"gemini/gemini-3.8-flash": "gemini-3.1-pro-preview",
"gemini/gemini-3.1-pro-preview": "gemini-3.8-flash",
"together_ai/moonshotai/Kimi-K3": "zai-org/GLM-5.3",
"together_ai/zai-org/GLM-5.3": "moonshotai/Kimi-K3",
"fireworks_ai/kimi-k3": "qwen3p8-max",
"fireworks_ai/qwen3p8-max": "kimi-k3",
"fireworks_ai/deepseek-v4p1-flash": "kimi-k3",
}
_OVERRIDE_MAP_KEYS: Final[dict[str, str]] = {
"gpt-5.4-mini": "gpt-5.4-mini",
"gpt-5.6": "gpt-5.6",
"gpt-5.3-codex": "gpt-5.3-codex",
"gpt-5.5-pro": "gpt-5.5-pro",
"claude-sonnet-5": "claude-sonnet-5",
"claude-opus-5": "claude-opus-5",
"gemini-3.1-pro-preview": "gemini/gemini-3.1-pro-preview",
"gemini-3.8-flash": "gemini/gemini-3.8-flash",
"zai-org/GLM-5.3": "together_ai/zai-org/GLM-5.3",
"moonshotai/Kimi-K3": "together_ai/moonshotai/Kimi-K3",
"qwen3p8-max": "fireworks_ai/qwen3p8-max",
"kimi-k3": "fireworks_ai/kimi-k3",
}
_FRONTIER_SPECS: Final[tuple[tuple[str, str, Wire], ...]] = (
("gpt-5.6", "openai/gpt-5.6", "openai_chat"),
("gpt-5.5-pro", "openai/gpt-5.5-pro", "openai_responses"),
("gpt-5.3-codex", "openai/gpt-5.3-codex", "openai_responses"),
("gpt-5.4-mini", "openai/gpt-5.4-mini", "openai_chat"),
("claude-opus-5", "anthropic/claude-opus-5", "anthropic_messages"),
("claude-sonnet-5", "anthropic/claude-sonnet-5", "anthropic_messages"),
("claude-haiku-4-5", "anthropic/claude-haiku-4-5", "anthropic_messages"),
("gemini/gemini-3.8-flash", "gemini/gemini-3.8-flash", "gemini_generate"),
("gemini/gemini-3.1-pro-preview", "gemini/gemini-3.1-pro-preview", "gemini_generate"),
("together_ai/moonshotai/Kimi-K3", "together_ai/moonshotai/Kimi-K3", "together_chat"),
("together_ai/zai-org/GLM-5.3", "together_ai/zai-org/GLM-5.3", "together_chat"),
("fireworks_ai/kimi-k3", "fireworks_ai/kimi-k3", "fireworks_chat"),
("fireworks_ai/qwen3p8-max", "fireworks_ai/qwen3p8-max", "fireworks_chat"),
("fireworks_ai/deepseek-v4p1-flash", "fireworks_ai/deepseek-v4p1-flash", "fireworks_chat"),
)
def _frontier() -> tuple[FrontierModel, ...]:
return tuple(
FrontierModel(
model_name=f"cc-{map_key.replace('/', '-').lower()}",
litellm_model=litellm_model,
wire=wire,
map_key=map_key,
override_model=_OVERRIDE_MODELS[map_key],
)
for map_key, litellm_model, wire in _FRONTIER_SPECS
)
FRONTIER_MODELS: Final[tuple[FrontierModel, ...]] = _frontier()
# Token kinds each wire can report, gating which pricing cases apply.
_WIRE_CAPS: Final[dict[str, frozenset[str]]] = {
"openai_chat": frozenset(
{
"cache_read", "cache_write_5m", "cache_write_1h", "reasoning", "audio",
"web_search", "response_model", "absent_usage",
}
),
"openai_responses": frozenset({"cache_read", "reasoning", "web_search", "response_model", "absent_usage"}),
# Product gap: litellm hard-indexes message_delta["usage"] in
# anthropic/chat/handler.py, so a usage-absent anthropic stream raises
# KeyError; the real wire always carries it, so the case cannot be
# represented.
"anthropic_messages": frozenset({"cache_read", "cache_write_5m", "cache_write_1h", "web_search", "response_model"}),
# Product gap: the gemini transform sets ModelResponse.model from the
# request and drops the provider's modelVersion, so a response-model
# override can never be priced on this wire.
"gemini_generate": frozenset({"cache_read", "reasoning", "audio", "web_search", "absent_usage"}),
"together_chat": frozenset(
{
"cache_read", "cache_write_5m", "cache_write_1h", "reasoning", "audio",
"web_search", "response_model", "absent_usage",
}
),
"fireworks_chat": frozenset(
{
"cache_read", "cache_write_5m", "cache_write_1h", "reasoning", "audio",
"web_search", "response_model", "absent_usage",
}
),
}
CaseName = Literal[
"basic",
"cache_read",
"cache_write_5m",
"cache_write_1h",
"reasoning",
"audio",
"tiered",
"service_tier_flex",
"service_tier_priority",
"web_search",
"stream",
"stream_no_usage",
"response_model_override",
]
@dataclass(frozen=True, slots=True)
class Case:
name: CaseName
usage: ScriptedUsage
stream: bool = False
stream_usage: Literal["final_chunk", "absent"] = "final_chunk"
service_tier: Literal["flex", "priority"] | None = None
# For web_search the wire's reported call count is not always what gets
# billed: chat-completions surfaces only expose url_citation annotations, so
# the biller floors to one call; responses/messages/gemini report a real
# count.
billed_web_search_calls: int = 0
response_model_override: bool = False
exact_spend: bool = True
# stream_usage=absent on a wire with no proxy-side token recount means the
# bill is exactly zero; asserted as such rather than skipped.
expect_zero_bill: bool = False
def scenario(self, scenario_id: str, model: FrontierModel, text: str) -> Scenario:
return Scenario(
scenario_id=scenario_id,
wire=model.wire,
usage=self.usage,
output=ScriptedOutput(
text=text,
response_model=model.override_model if self.response_model_override else None,
),
stream_usage=self.stream_usage,
service_tier=self.service_tier,
)
_BASIC_USAGE: Final = ScriptedUsage(fresh_input_tokens=120, output_tokens=40)
def _web_search_case(model: FrontierModel) -> Case:
counts_exactly = model.wire in ("openai_responses", "anthropic_messages", "gemini_generate")
return Case(
name="web_search",
usage=ScriptedUsage(fresh_input_tokens=100, output_tokens=30, web_search_calls=3),
billed_web_search_calls=3 if counts_exactly else 1,
)
def cases_for(model: FrontierModel) -> tuple[Case, ...]:
rates = model.rates
caps = _WIRE_CAPS[model.wire]
cases: list[Case] = [Case(name="basic", usage=_BASIC_USAGE)]
if rates.cache_read_input_token_cost is not None and "cache_read" in caps:
cases.append(
Case(name="cache_read", usage=ScriptedUsage(fresh_input_tokens=100, cache_read_tokens=50, output_tokens=30))
)
if rates.cache_creation_input_token_cost is not None and "cache_write_5m" in caps:
cases.append(
Case(
name="cache_write_5m",
usage=ScriptedUsage(fresh_input_tokens=90, cache_write_5m_tokens=60, output_tokens=30),
)
)
if (
rates.cache_creation_input_token_cost_above_1hr is not None
and rates.cache_creation_input_token_cost is not None
and "cache_write_1h" in caps
):
cases.append(
Case(
name="cache_write_1h",
usage=ScriptedUsage(
fresh_input_tokens=90,
cache_write_5m_tokens=20,
cache_write_1h_tokens=40,
output_tokens=30,
),
)
)
if rates.output_cost_per_reasoning_token is not None and "reasoning" in caps:
cases.append(
Case(
name="reasoning",
usage=ScriptedUsage(fresh_input_tokens=100, output_tokens=30, reasoning_tokens=70),
)
)
if (
rates.input_cost_per_audio_token is not None
and rates.output_cost_per_audio_token is not None
and "audio" in caps
):
cases.append(
Case(
name="audio",
usage=ScriptedUsage(
fresh_input_tokens=100, audio_input_tokens=25, output_tokens=30, audio_output_tokens=15
),
)
)
if (
rates.input_cost_per_token_above_200k_tokens is not None
and rates.output_cost_per_token_above_200k_tokens is not None
):
cases.append(
Case(
name="tiered",
usage=ScriptedUsage(
fresh_input_tokens=TIER_THRESHOLD_TOKENS + 1, output_tokens=30
),
)
)
if rates.input_cost_per_token_flex is not None and rates.output_cost_per_token_flex is not None:
cases.append(
Case(name="service_tier_flex", usage=_BASIC_USAGE, service_tier="flex")
)
if rates.input_cost_per_token_priority is not None and rates.output_cost_per_token_priority is not None:
cases.append(
Case(name="service_tier_priority", usage=_BASIC_USAGE, service_tier="priority")
)
if rates.search_context_cost_per_query is not None and "web_search" in caps:
cases.append(_web_search_case(model))
cases.append(Case(name="stream", usage=_BASIC_USAGE, stream=True))
if "absent_usage" in caps:
cases.append(
Case(
name="stream_no_usage",
usage=_BASIC_USAGE,
stream=True,
stream_usage="absent",
exact_spend=False,
# The responses surface bills only provider-reported usage;
# with no usage in the stream the spend row is zero. Other
# wires recount tokens proxy-side and bill a nonzero amount.
expect_zero_bill=model.wire == "openai_responses",
)
)
if "response_model" in caps:
cases.append(Case(name="response_model_override", usage=_BASIC_USAGE, response_model_override=True))
return tuple(cases)
@dataclass(frozen=True, slots=True)
class ExpectedCost:
"""The expected bill split the way the spend row's cost_breakdown reports
it: the gross input component (cache reads/writes folded in), the output
component, and the tool-usage component."""
input_cost: float
output_cost: float
tool_cost: float
@property
def total(self) -> float:
return self.input_cost + self.output_cost + self.tool_cost
def expected_breakdown(model: FrontierModel, case: Case) -> ExpectedCost:
"""Literal arithmetic on the test-map rates over the scripted token counts.
Input = fresh*in + read*read + 5m*create + 1h*create_1h + audio_in*audio_in;
output = text*out + reasoning*reasoning + audio_out*audio_out; plus the
billed web-search calls at the medium search-context rate. Above-threshold
swaps every input/output rate to its ``_above_200k_tokens`` variant when
total prompt tokens exceed the threshold; a service tier swaps input/output
to the tier's variants, falling back to the base rate when a variant is
unset -- mirroring _get_token_base_cost in litellm's cost calculator.
"""
rates = model.override_rates if case.response_model_override else model.rates
u = case.usage
prompt_tokens = (
u.fresh_input_tokens + u.cache_read_tokens + u.cache_write_5m_tokens
+ u.cache_write_1h_tokens + u.audio_input_tokens
)
tiered = prompt_tokens > TIER_THRESHOLD_TOKENS
in_rate = rates.input_cost_per_token or 0.0
out_rate = rates.output_cost_per_token or 0.0
if case.service_tier == "flex":
in_rate = rates.input_cost_per_token_flex or in_rate
out_rate = rates.output_cost_per_token_flex or out_rate
if case.service_tier == "priority":
in_rate = rates.input_cost_per_token_priority or in_rate
out_rate = rates.output_cost_per_token_priority or out_rate
if tiered:
in_rate = rates.input_cost_per_token_above_200k_tokens or in_rate
out_rate = rates.output_cost_per_token_above_200k_tokens or out_rate
input_cost = (
u.fresh_input_tokens * in_rate
+ u.cache_read_tokens * (rates.cache_read_input_token_cost or 0.0)
+ u.cache_write_5m_tokens * (rates.cache_creation_input_token_cost or 0.0)
+ u.cache_write_1h_tokens * (rates.cache_creation_input_token_cost_above_1hr or 0.0)
+ u.audio_input_tokens * (rates.input_cost_per_audio_token or 0.0)
)
output_cost = (
u.output_tokens * out_rate
+ u.reasoning_tokens * (rates.output_cost_per_reasoning_token or out_rate)
+ u.audio_output_tokens * (rates.output_cost_per_audio_token or out_rate)
)
search = rates.search_context_cost_per_query
tool_cost = case.billed_web_search_calls * (
search.search_context_size_medium if search and search.search_context_size_medium else 0.0
)
return ExpectedCost(input_cost=input_cost, output_cost=output_cost, tool_cost=tool_cost)
def expected_cost(model: FrontierModel, case: Case) -> float:
return expected_breakdown(model, case).total
def expected_token_columns(model: FrontierModel, case: Case) -> tuple[int, int]:
"""(prompt_tokens, completion_tokens) the spend row should carry, per the
wire's normalization: Anthropic folds cache read/write into prompt_tokens,
everyone else reports the totals the wire emitted."""
u = case.usage
if model.wire == "anthropic_messages":
return (
u.fresh_input_tokens + u.cache_read_tokens + u.cache_write_5m_tokens + u.cache_write_1h_tokens,
u.output_tokens,
)
if model.wire == "gemini_generate":
return (
u.fresh_input_tokens + u.cache_read_tokens + u.audio_input_tokens,
u.output_tokens + u.reasoning_tokens + u.audio_output_tokens,
)
if model.wire == "openai_responses":
return (
u.fresh_input_tokens + u.cache_read_tokens,
u.output_tokens + u.reasoning_tokens,
)
return (
u.fresh_input_tokens
+ u.cache_read_tokens
+ u.cache_write_5m_tokens
+ u.cache_write_1h_tokens
+ u.audio_input_tokens,
u.output_tokens + u.reasoning_tokens + u.audio_output_tokens,
)

View file

@ -0,0 +1,70 @@
"""Client side of the scripted-provider sidecar: register scenarios over its
control API through the shared transport helpers and get back a handle whose
``api_base`` is what a /model/new deployment should register for the proxy to
reach the scripted wire."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Final
from e2e_config import SCRIPTED_PROVIDER_CONTROL_URL, SCRIPTED_PROVIDER_PROXY_BASE
from e2e_http import URL, NoBody, unwrap, post
from e2e_http import delete as http_delete
from scripted_provider import (
Scenario,
ScenarioDeleted,
ScenarioRegistered,
Wire,
)
@dataclass(frozen=True, slots=True)
class ScenarioHandle:
scenario_id: str
wire: Wire
proxy_base: str
def api_base(self) -> str:
return f"{self.proxy_base}/{self.scenario_id}/{self._mount()}"
def _mount(self) -> str:
return {
"openai_chat": "openai",
"openai_responses": "openai",
"anthropic_messages": "anthropic",
"gemini_generate": "gemini",
"together_chat": "together",
"fireworks_chat": "fireworks",
}[self.wire]
def register_scenario(scenario: Scenario) -> ScenarioHandle:
"""POST the scenario to the sidecar's control API and return its handle."""
result = unwrap(
post(
URL(f"{SCRIPTED_PROVIDER_CONTROL_URL}/_scenarios"),
headers=NoBody(),
json=scenario,
response_type=ScenarioRegistered,
)
)
return ScenarioHandle(
scenario_id=result.scenario_id,
wire=scenario.wire,
proxy_base=SCRIPTED_PROVIDER_PROXY_BASE,
)
def delete_scenario(handle: ScenarioHandle) -> None:
unwrap(
http_delete(
URL(f"{SCRIPTED_PROVIDER_CONTROL_URL}/_scenarios/{handle.scenario_id}"),
headers=NoBody(),
json=NoBody(),
response_type=ScenarioDeleted,
)
)
CONTROL_URL: Final = SCRIPTED_PROVIDER_CONTROL_URL

View file

@ -0,0 +1,631 @@
"""Scripted provider sidecar for the cost-calculation e2e suite.
A standalone process (``python -m cost_calculation.scripted_provider``) that
pretends to be an LLM provider for the proxy under test. The suite registers a
Scenario over a small control API; the provider wire routes then answer the
proxy's upstream calls with the scripted usage figures, in the exact wire shape
the real provider would emit (OpenAI chat completions, OpenAI Responses,
Anthropic Messages, Gemini generateContent, or the OpenAI-compatible Together /
Fireworks surfaces). Because the usage is scripted, expected spend is literal
arithmetic on the test cost map's rates, with no dependency on what a real
provider would report.
Layout on one port:
- ``GET /health`` liveness
- ``POST /_scenarios`` register a Scenario JSON, returns its id
- ``DELETE /_scenarios/<id>`` remove it
- ``POST /<id>/<mount>/<provider path>`` provider wire; mount is one of
``openai``, ``anthropic``, ``gemini``, ``together``, ``fireworks`` and the
remainder is whatever path the provider client appends (``chat/completions``,
``responses``, ``v1/messages``, ``models/<m>:generateContent`` ...)
A request carrying ``"stream": true`` (or the ``:streamGenerateContent`` Gemini
verb) gets an SSE answer; ``stream_usage`` on the Scenario decides whether the
final stream chunk carries usage or the provider reports none.
"""
from __future__ import annotations
import json
import sys
import threading
import time
from dataclasses import dataclass
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from typing import Final, Literal
from urllib.parse import urlsplit
from pydantic import BaseModel, ConfigDict, TypeAdapter, ValidationError
Wire = Literal[
"openai_chat",
"openai_responses",
"anthropic_messages",
"gemini_generate",
"together_chat",
"fireworks_chat",
]
_WIRE_MOUNTS: Final[dict[str, str]] = {
"openai_chat": "openai",
"openai_responses": "openai",
"anthropic_messages": "anthropic",
"gemini_generate": "gemini",
"together_chat": "together",
"fireworks_chat": "fireworks",
}
StreamUsage = Literal["final_chunk", "absent"]
ServiceTier = Literal["flex", "priority"]
class ScriptedUsage(BaseModel):
"""Physical token counts the scripted response reports. ``fresh_input_tokens``
is the uncached, never-written, non-audio input count; ``output_tokens`` is
the non-reasoning, non-audio output count. Renderers add the cached, written,
audio, and reasoning counts into the wire's total fields the way the real
provider does (inside prompt_tokens for OpenAI/Gemini, as uncached-only
input_tokens for Anthropic)."""
model_config = ConfigDict(frozen=True)
fresh_input_tokens: int = 0
output_tokens: int = 0
cache_read_tokens: int = 0
cache_write_5m_tokens: int = 0
cache_write_1h_tokens: int = 0
reasoning_tokens: int = 0
audio_input_tokens: int = 0
audio_output_tokens: int = 0
web_search_calls: int = 0
class ScriptedOutput(BaseModel):
model_config = ConfigDict(frozen=True)
text: str
finish_reason: str = "stop"
# When set, emitted verbatim as the response's model field, letting a test
# prove the biller prices the provider-reported model.
response_model: str | None = None
# OpenAI-compatible providers can report a provider-computed cost; emitted as
# the top-level "cost" field on the together/fireworks wire.
provider_cost: float | None = None
class Scenario(BaseModel):
model_config = ConfigDict(frozen=True)
scenario_id: str
wire: Wire
usage: ScriptedUsage
output: ScriptedOutput
stream_usage: StreamUsage = "final_chunk"
service_tier: ServiceTier | None = None
@property
def mount(self) -> str:
return _WIRE_MOUNTS[self.wire]
class ScenarioRegistered(BaseModel):
scenario_id: str
class ScenarioDeleted(BaseModel):
deleted: bool
class HealthStatus(BaseModel):
status: str
@dataclass(frozen=True, slots=True)
class RenderedResponse:
status_code: int
content_type: str
body: bytes
def _json_bytes(payload: dict[str, object]) -> bytes:
return json.dumps(payload).encode("utf-8")
def _sse(events: tuple[tuple[str | None, dict[str, object] | str], ...]) -> bytes:
frames: list[str] = []
for event_name, data in events:
head = f"event: {event_name}\n" if event_name is not None else ""
payload = data if isinstance(data, str) else json.dumps(data)
frames.append(f"{head}data: {payload}\n\n")
return "".join(frames).encode("utf-8")
# ---------- per-wire usage shapes ----------
def _openai_usage(u: ScriptedUsage) -> dict[str, object]:
prompt_tokens = (
u.fresh_input_tokens
+ u.cache_read_tokens
+ u.cache_write_5m_tokens
+ u.cache_write_1h_tokens
+ u.audio_input_tokens
)
completion_tokens = u.output_tokens + u.reasoning_tokens + u.audio_output_tokens
prompt_details: dict[str, object] = {}
if u.cache_read_tokens:
prompt_details["cached_tokens"] = u.cache_read_tokens
if u.cache_write_5m_tokens or u.cache_write_1h_tokens:
prompt_details["cache_write_tokens"] = u.cache_write_5m_tokens + u.cache_write_1h_tokens
prompt_details["cache_creation_token_details"] = {
"ephemeral_5m_input_tokens": u.cache_write_5m_tokens,
"ephemeral_1h_input_tokens": u.cache_write_1h_tokens,
}
if u.audio_input_tokens:
prompt_details["audio_tokens"] = u.audio_input_tokens
completion_details: dict[str, object] = {}
if u.reasoning_tokens:
completion_details["reasoning_tokens"] = u.reasoning_tokens
if u.audio_output_tokens:
completion_details["audio_tokens"] = u.audio_output_tokens
usage: dict[str, object] = {
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"total_tokens": prompt_tokens + completion_tokens,
}
if prompt_details:
usage["prompt_tokens_details"] = prompt_details
if completion_details:
usage["completion_tokens_details"] = completion_details
return usage
def _anthropic_usage(u: ScriptedUsage) -> dict[str, object]:
# Anthropic reports uncached-only input_tokens; cache reads and writes ride
# top-level fields, with the 5m/1h write split under cache_creation.
usage: dict[str, object] = {
"input_tokens": u.fresh_input_tokens,
"output_tokens": u.output_tokens,
}
if u.cache_read_tokens:
usage["cache_read_input_tokens"] = u.cache_read_tokens
if u.cache_write_5m_tokens or u.cache_write_1h_tokens:
usage["cache_creation_input_tokens"] = u.cache_write_5m_tokens + u.cache_write_1h_tokens
usage["cache_creation"] = {
"ephemeral_5m_input_tokens": u.cache_write_5m_tokens,
"ephemeral_1h_input_tokens": u.cache_write_1h_tokens,
}
if u.web_search_calls:
usage["server_tool_use"] = {"web_search_requests": u.web_search_calls}
return usage
def _gemini_usage(u: ScriptedUsage) -> dict[str, object]:
# promptTokenCount carries the cached count inside it; TEXT modality is the
# cached-inclusive text count so litellm's implicit-caching subtraction lands
# on the fresh figure. candidatesTokenCount includes reasoning + audio.
prompt_tokens = u.fresh_input_tokens + u.cache_read_tokens + u.audio_input_tokens
candidates = u.output_tokens + u.reasoning_tokens + u.audio_output_tokens
usage: dict[str, object] = {
"promptTokenCount": prompt_tokens,
"candidatesTokenCount": candidates,
"totalTokenCount": prompt_tokens + candidates,
}
if u.cache_read_tokens:
usage["cachedContentTokenCount"] = u.cache_read_tokens
if u.reasoning_tokens:
usage["thoughtsTokenCount"] = u.reasoning_tokens
prompt_details = [{"modality": "TEXT", "tokenCount": u.fresh_input_tokens + u.cache_read_tokens}]
if u.audio_input_tokens:
prompt_details.append({"modality": "AUDIO", "tokenCount": u.audio_input_tokens})
usage["promptTokensDetails"] = prompt_details
if u.audio_output_tokens:
usage["candidatesTokensDetails"] = [
{"modality": "TEXT", "tokenCount": u.output_tokens + u.reasoning_tokens},
{"modality": "AUDIO", "tokenCount": u.audio_output_tokens},
]
return usage
def _responses_usage(u: ScriptedUsage) -> dict[str, object]:
input_tokens = u.fresh_input_tokens + u.cache_read_tokens + u.audio_input_tokens
output_tokens = u.output_tokens + u.reasoning_tokens + u.audio_output_tokens
usage: dict[str, object] = {
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"total_tokens": input_tokens + output_tokens,
}
input_details: dict[str, object] = {}
if u.cache_read_tokens:
input_details["cached_tokens"] = u.cache_read_tokens
if input_details:
usage["input_tokens_details"] = input_details
if u.reasoning_tokens:
usage["output_tokens_details"] = {"reasoning_tokens": u.reasoning_tokens}
return usage
# ---------- per-wire responses ----------
def _openai_message(scenario: Scenario) -> dict[str, object]:
message: dict[str, object] = {"role": "assistant", "content": scenario.output.text}
if scenario.usage.web_search_calls:
message["annotations"] = [
{
"type": "url_citation",
"url_citation": {
"url": "https://scripted.example/source",
"title": "scripted source",
"start_index": 0,
"end_index": 1,
},
}
for _ in range(scenario.usage.web_search_calls)
]
return message
def _openai_chat_body(scenario: Scenario, requested_model: str) -> dict[str, object]:
body: dict[str, object] = {
"id": f"chatcmpl-{scenario.scenario_id}",
"object": "chat.completion",
"created": int(time.time()),
"model": scenario.output.response_model or requested_model,
"choices": [
{
"index": 0,
"message": _openai_message(scenario),
"finish_reason": scenario.output.finish_reason,
}
],
"usage": _openai_usage(scenario.usage),
}
if scenario.service_tier is not None:
body["service_tier"] = scenario.service_tier
if scenario.output.provider_cost is not None:
body["cost"] = scenario.output.provider_cost
return body
def _openai_chunk(scenario: Scenario, requested_model: str, **kw: object) -> dict[str, object]:
chunk: dict[str, object] = {
"id": f"chatcmpl-{scenario.scenario_id}",
"object": "chat.completion.chunk",
"created": int(time.time()),
"model": scenario.output.response_model or requested_model,
}
chunk.update(kw)
return chunk
def _openai_chat_sse(scenario: Scenario, requested_model: str) -> bytes:
_EMPTY_DELTA: Final[dict[str, object]] = {}
delta: dict[str, object] = {"role": "assistant", "content": scenario.output.text}
if scenario.usage.web_search_calls:
delta["annotations"] = _openai_message(scenario)["annotations"]
events: list[tuple[str | None, dict[str, object] | str]] = [
(
None,
_openai_chunk(
scenario,
requested_model,
choices=[{"index": 0, "delta": {"role": "assistant"}, "finish_reason": None}],
),
),
(
None,
_openai_chunk(
scenario,
requested_model,
choices=[{"index": 0, "delta": delta, "finish_reason": None}],
),
),
(
None,
_openai_chunk(
scenario,
requested_model,
choices=[
{
"index": 0,
"delta": _EMPTY_DELTA,
"finish_reason": scenario.output.finish_reason,
}
],
),
),
]
if scenario.stream_usage == "final_chunk":
events.append(
(None, _openai_chunk(scenario, requested_model, choices=(), usage=_openai_usage(scenario.usage)))
)
events.append((None, "[DONE]"))
return _sse(tuple(events))
def _anthropic_body(scenario: Scenario, requested_model: str) -> dict[str, object]:
return {
"id": f"msg_{scenario.scenario_id}",
"type": "message",
"role": "assistant",
"model": scenario.output.response_model or requested_model,
"content": [{"type": "text", "text": scenario.output.text}],
"stop_reason": "end_turn" if scenario.output.finish_reason == "stop" else scenario.output.finish_reason,
"usage": _anthropic_usage(scenario.usage),
}
def _anthropic_sse(scenario: Scenario, requested_model: str) -> bytes:
emit_usage = scenario.stream_usage == "final_chunk"
input_usage = {k: v for k, v in _anthropic_usage(scenario.usage).items() if k != "output_tokens"}
message_start: dict[str, object] = {
"type": "message_start",
"message": {
"id": f"msg_{scenario.scenario_id}",
"type": "message",
"role": "assistant",
"model": scenario.output.response_model or requested_model,
"content": [],
"stop_reason": None,
**({"usage": input_usage} if emit_usage else {}),
},
}
message_delta: dict[str, object] = {
"type": "message_delta",
"delta": {
"stop_reason": "end_turn" if scenario.output.finish_reason == "stop" else scenario.output.finish_reason
},
**({"usage": {"output_tokens": scenario.usage.output_tokens}} if emit_usage else {}),
}
return _sse(
(
("message_start", message_start),
(
"content_block_start",
{
"type": "content_block_start",
"index": 0,
"content_block": {"type": "text", "text": ""},
},
),
(
"content_block_delta",
{
"type": "content_block_delta",
"index": 0,
"delta": {"type": "text_delta", "text": scenario.output.text},
},
),
("content_block_stop", {"type": "content_block_stop", "index": 0}),
("message_delta", message_delta),
("message_stop", {"type": "message_stop"}),
)
)
def _gemini_body(scenario: Scenario, requested_model: str) -> dict[str, object]:
candidate: dict[str, object] = {
"content": {"parts": [{"text": scenario.output.text}], "role": "model"},
"finishReason": "STOP" if scenario.output.finish_reason == "stop" else scenario.output.finish_reason.upper(),
"index": 0,
}
if scenario.usage.web_search_calls:
candidate["groundingMetadata"] = {
"webSearchQueries": [f"query {i}" for i in range(scenario.usage.web_search_calls)]
}
return {
"candidates": [candidate],
"usageMetadata": _gemini_usage(scenario.usage),
"modelVersion": scenario.output.response_model or requested_model,
}
def _gemini_sse(scenario: Scenario, requested_model: str) -> bytes:
first = _gemini_body(scenario, requested_model)
if scenario.stream_usage == "absent":
first = {k: v for k, v in first.items() if k != "usageMetadata"}
events: list[tuple[str | None, dict[str, object] | str]] = [(None, first)]
if scenario.stream_usage == "final_chunk":
events.append(
(
None,
{
"candidates": [],
"usageMetadata": _gemini_usage(scenario.usage),
"modelVersion": scenario.output.response_model or requested_model,
},
)
)
return _sse(tuple(events))
def _responses_body(scenario: Scenario, requested_model: str) -> dict[str, object]:
output: list[dict[str, object]] = [
{"type": "web_search_call", "id": f"ws_{i}", "status": "completed"}
for i in range(scenario.usage.web_search_calls)
]
output.append(
{
"type": "message",
"id": f"msg_{scenario.scenario_id}",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": scenario.output.text,
"annotations": [],
}
],
}
)
return {
"id": f"resp_{scenario.scenario_id}",
"object": "response",
"created_at": int(time.time()),
"status": "completed",
"model": scenario.output.response_model or requested_model,
"output": output,
"usage": _responses_usage(scenario.usage),
}
def _responses_sse(scenario: Scenario, requested_model: str) -> bytes:
completed = _responses_body(scenario, requested_model)
if scenario.stream_usage == "absent":
completed = {k: v for k, v in completed.items() if k != "usage"}
created = {**completed, "status": "in_progress", "usage": None}
return _sse(
(
("response.created", {"type": "response.created", "response": created}),
(
"response.output_text.delta",
{
"type": "response.output_text.delta",
"item_id": f"msg_{scenario.scenario_id}",
"output_index": scenario.usage.web_search_calls,
"content_index": 0,
"delta": scenario.output.text,
},
),
("response.completed", {"type": "response.completed", "response": completed}),
)
)
def _render(scenario: Scenario, *, stream: bool, requested_model: str) -> RenderedResponse:
if scenario.wire == "anthropic_messages":
if stream:
return RenderedResponse(200, "text/event-stream", _anthropic_sse(scenario, requested_model))
return RenderedResponse(200, "application/json", _json_bytes(_anthropic_body(scenario, requested_model)))
if scenario.wire == "gemini_generate":
if stream:
return RenderedResponse(200, "text/event-stream", _gemini_sse(scenario, requested_model))
return RenderedResponse(200, "application/json", _json_bytes(_gemini_body(scenario, requested_model)))
if scenario.wire == "openai_responses":
if stream:
return RenderedResponse(200, "text/event-stream", _responses_sse(scenario, requested_model))
return RenderedResponse(200, "application/json", _json_bytes(_responses_body(scenario, requested_model)))
# openai_chat, together_chat, fireworks_chat share the OpenAI chat shape.
if stream:
return RenderedResponse(200, "text/event-stream", _openai_chat_sse(scenario, requested_model))
return RenderedResponse(200, "application/json", _json_bytes(_openai_chat_body(scenario, requested_model)))
# ---------- registry + request routing ----------
class _ScenarioStore:
def __init__(self) -> None:
self._lock: Final = threading.Lock()
self._scenarios: dict[str, Scenario] = {} # mutable-ok: server state, guarded by _lock
def put(self, scenario: Scenario) -> None:
with self._lock:
self._scenarios[scenario.scenario_id] = scenario
def drop(self, scenario_id: str) -> bool:
with self._lock:
return self._scenarios.pop(scenario_id, None) is not None
def get(self, scenario_id: str) -> Scenario | None:
with self._lock:
return self._scenarios.get(scenario_id)
_REQUEST_BODY: Final = TypeAdapter(dict[str, object])
def _request_body(body: bytes) -> dict[str, object]:
try:
return _REQUEST_BODY.validate_json(body)
except ValueError:
return {}
def _request_wants_stream(path_tail: str, body: bytes) -> bool:
if ":streamGenerateContent" in path_tail:
return True
if not body:
return False
return _request_body(body).get("stream") is True
def _request_model(body: bytes) -> str:
model = _request_body(body).get("model")
return model if isinstance(model, str) else "unknown"
def handle_request(store: _ScenarioStore, method: str, raw_path: str, body: bytes) -> RenderedResponse:
path = urlsplit(raw_path).path
segments = [segment for segment in path.split("/") if segment]
if method == "GET" and segments == ["health"]:
return RenderedResponse(200, "application/json", _json_bytes({"status": "ok"}))
if segments and segments[0] == "_scenarios":
if method == "POST" and len(segments) == 1:
try:
scenario = Scenario.model_validate_json(body)
except ValidationError as exc:
return RenderedResponse(400, "application/json", _json_bytes({"error": str(exc)}))
store.put(scenario)
return RenderedResponse(200, "application/json", _json_bytes({"scenario_id": scenario.scenario_id}))
if method == "DELETE" and len(segments) == 2:
deleted = store.drop(segments[1])
return RenderedResponse(
200 if deleted else 404, "application/json", _json_bytes({"deleted": deleted})
)
return RenderedResponse(404, "application/json", _json_bytes({"error": "unknown control route"}))
if len(segments) < 2 or method != "POST":
return RenderedResponse(404, "application/json", _json_bytes({"error": f"no route for {method} {path}"}))
scenario_id, mount = segments[0], segments[1]
scenario = store.get(scenario_id)
if scenario is None:
return RenderedResponse(404, "application/json", _json_bytes({"error": f"unknown scenario {scenario_id}"}))
if scenario.mount != mount:
return RenderedResponse(
400,
"application/json",
_json_bytes({"error": f"scenario {scenario_id} is wire {scenario.wire}, not mount {mount}"}),
)
tail = "/".join(segments[2:])
return _render(scenario, stream=_request_wants_stream(tail, body), requested_model=_request_model(body))
class _ScriptedHandler(BaseHTTPRequestHandler):
store: Final[_ScenarioStore] = _ScenarioStore()
def _dispatch(self, method: str) -> None:
length = int(self.headers.get("content-length") or 0)
body = self.rfile.read(length) if length else b""
rendered = handle_request(self.store, method, self.path, body)
self.send_response(rendered.status_code)
self.send_header("content-type", rendered.content_type)
self.send_header("content-length", str(len(rendered.body)))
self.end_headers()
self.wfile.write(rendered.body)
def do_GET(self) -> None:
self._dispatch("GET")
def do_POST(self) -> None:
self._dispatch("POST")
def do_DELETE(self) -> None:
self._dispatch("DELETE")
DEFAULT_PORT: Final = 9100
def serve(port: int = DEFAULT_PORT, bind_host: str = "127.0.0.1") -> None:
server = ThreadingHTTPServer((bind_host, port), _ScriptedHandler)
sys.stderr.write(f"scripted-provider listening on http://{bind_host}:{port}\n")
server.serve_forever()
if __name__ == "__main__":
port_arg = int(sys.argv[1]) if len(sys.argv) > 1 else DEFAULT_PORT
serve(port=port_arg)

View file

@ -0,0 +1,115 @@
"""Token-pricing e2e: every (frontier model, pricing-component case) cell runs a
scripted-usage call through a deployment registered on the cost-map proxy, and
the spend row plus response-cost header must equal literal arithmetic on the
test map's rates.
Nothing here touches a real provider or the bundled cost map: the proxy's
upstream is the scripted-provider sidecar and its entire cost map is
tests/e2e/cost_map.json.
"""
from __future__ import annotations
import pytest
from conftest import CostCalcClient, cost_rows, register_scenario_deployment
from cost_matrix import (
FRONTIER_MODELS,
Case,
FrontierModel,
cases_for,
expected_cost,
expected_token_columns,
)
from e2e_config import unique_marker
from lifecycle import ResourceManager
from models import ChatBody, ChatMessage, ChatStreamOptions
pytestmark = [pytest.mark.e2e, pytest.mark.cost_map_stack]
_MATRIX: list[tuple[FrontierModel, Case]] = [
(model, case) for model in FRONTIER_MODELS for case in cases_for(model)
]
def _case_id(param: tuple[FrontierModel, Case]) -> str:
model, case = param
return f"{model.map_key.replace('/', '-')}-{case.name}"
def _chat_body(model_name: str, marker: str, case: Case) -> ChatBody:
return ChatBody(
model=model_name,
messages=[ChatMessage(role="user", content=f"{marker} scripted pricing call")],
stream=case.stream,
stream_options=ChatStreamOptions(include_usage=True) if case.stream else None,
service_tier=case.service_tier,
)
class TestTokenPricing:
@pytest.mark.parametrize("model_case", _MATRIX, ids=_case_id)
@pytest.mark.covers("quota_management.spend_tracking.cost_matrix.logs_cost")
def test_scripted_usage_bills_at_map_rates(
self,
client: CostCalcClient,
resources: ResourceManager,
scoped_key: str,
model_case: tuple[FrontierModel, Case],
) -> None:
model, case = model_case
marker = unique_marker()
model_name, _handle = register_scenario_deployment(client, resources, model, case, marker)
response = client.proxy.transport.send(
"/chat/completions",
headers=client.proxy.transport.bearer(scoped_key),
json=_chat_body(model_name, marker, case),
stream=case.stream,
)
assert response.ok, (
f"{model.map_key}/{case.name}: proxy returned {response.status_code}: {response.body[:400]}"
)
assert response.stream_error is None, f"stream carried an error event: {response.stream_error}"
expected = expected_cost(model, case)
if case.exact_spend and not case.stream:
# Streamed responses commit headers before the bill is computed, so
# the x-litellm-response-cost header is asserted only on non-stream
# calls.
assert response.response_cost is not None and cost_rows.approx_equal(
response.response_cost, expected
), (
f"x-litellm-response-cost {response.response_cost} != expected {expected}"
)
row = cost_rows.poll_cost_row_where(
client.proxy,
scoped_key,
lambda r: r.metadata is not None and r.metadata.cost_breakdown is not None,
)
assert row is not None, f"no spend row with a cost breakdown landed for {model.map_key}/{case.name}"
if not case.exact_spend and case.expect_zero_bill:
# The provider reported no usage and this wire has no proxy-side
# recount, so the bill is exactly zero.
assert row.spend is not None and row.spend == 0, f"no-usage stream billed {row.spend}: {row}"
return
if not case.exact_spend:
# stream_usage=absent: the provider reported no usage, so the row's
# token counts are the proxy's own recount; only assert a bill landed.
assert row.spend is not None and row.spend > 0, f"no-usage stream billed nothing: {row}"
return
assert row.spend is not None and cost_rows.approx_equal(row.spend, expected), (
f"{model.map_key}/{case.name}: spend {row.spend} != expected {expected} "
f"(breakdown {row.breakdown.model_dump()})"
)
prompt_tokens, completion_tokens = expected_token_columns(model, case)
assert row.prompt_tokens == prompt_tokens, (
f"prompt_tokens {row.prompt_tokens} != {prompt_tokens}"
)
assert row.completion_tokens == completion_tokens, (
f"completion_tokens {row.completion_tokens} != {completion_tokens}"
)
cost_rows.assert_total_is_sum_of_components(row)

View file

@ -0,0 +1,186 @@
"""Wire-format e2e: one scripted upstream per provider wire, answering with a
usage payload where every token kind the wire can report is nonzero. The spend
row's gross input cost must equal fresh tokens at the input rate plus each cache
and audio component at its own rate -- proving the wire's usage shape landed the
cached tokens inside the total (OpenAI/Gemini) or as separate fields
(Anthropic), and that the biller subtracted them before billing fresh tokens.
Also covers the Responses API wire (an openai/gpt-5.5-pro deployment bridged by
the proxy to POST /responses) and a streamed Anthropic-messages case.
"""
from __future__ import annotations
import pytest
from conftest import CostCalcClient, cost_rows, register_scenario_deployment
from cost_matrix import (
FRONTIER_MODELS,
Case,
FrontierModel,
expected_breakdown,
expected_token_columns,
)
from e2e_config import unique_marker
from lifecycle import ResourceManager
from models import ChatBody, ChatMessage, ChatStreamOptions
from scripted_provider import ScriptedUsage
pytestmark = [pytest.mark.e2e, pytest.mark.cost_map_stack]
_MODELS: dict[str, FrontierModel] = {model.map_key: model for model in FRONTIER_MODELS}
# One scripted usage per wire, every reportable token kind nonzero.
_WIRE_USAGE: dict[str, tuple[str, ScriptedUsage]] = {
"openai_chat": (
"gpt-5.6",
ScriptedUsage(
fresh_input_tokens=80,
cache_read_tokens=40,
cache_write_5m_tokens=20,
cache_write_1h_tokens=10,
output_tokens=25,
reasoning_tokens=15,
audio_input_tokens=5,
audio_output_tokens=3,
),
),
"openai_responses": (
"gpt-5.5-pro",
ScriptedUsage(
fresh_input_tokens=80, cache_read_tokens=40, output_tokens=25, reasoning_tokens=15
),
),
"anthropic_messages": (
"claude-sonnet-5",
ScriptedUsage(
fresh_input_tokens=80,
cache_read_tokens=40,
cache_write_5m_tokens=20,
cache_write_1h_tokens=10,
output_tokens=25,
),
),
"gemini_generate": (
"gemini/gemini-3.8-flash",
ScriptedUsage(
fresh_input_tokens=80,
cache_read_tokens=40,
output_tokens=25,
reasoning_tokens=15,
audio_input_tokens=5,
audio_output_tokens=3,
),
),
"together_chat": (
"together_ai/moonshotai/Kimi-K3",
ScriptedUsage(
fresh_input_tokens=80,
cache_read_tokens=40,
cache_write_5m_tokens=20,
cache_write_1h_tokens=10,
output_tokens=25,
reasoning_tokens=15,
audio_input_tokens=5,
audio_output_tokens=3,
),
),
"fireworks_chat": (
"fireworks_ai/kimi-k3",
ScriptedUsage(fresh_input_tokens=80, cache_read_tokens=40, output_tokens=25),
),
}
class TestWireFormats:
@pytest.mark.parametrize("wire", tuple(_WIRE_USAGE))
@pytest.mark.covers("quota_management.spend_tracking.scripted_wire.logs_cost")
def test_wire_usage_shape_bills_each_component(
self,
client: CostCalcClient,
resources: ResourceManager,
scoped_key: str,
wire: str,
) -> None:
map_key, usage = _WIRE_USAGE[wire]
model = _MODELS[map_key]
case = Case(name="basic", usage=usage)
marker = unique_marker()
model_name, _handle = register_scenario_deployment(client, resources, model, case, marker)
response = client.proxy.transport.send(
"/chat/completions",
headers=client.proxy.transport.bearer(scoped_key),
json=ChatBody(
model=model_name,
messages=[ChatMessage(role="user", content=f"{marker} scripted wire call")],
),
)
assert response.ok, f"{wire}: proxy returned {response.status_code}: {response.body[:400]}"
expected = expected_breakdown(model, case)
row = cost_rows.poll_cost_row_where(
client.proxy,
scoped_key,
lambda r: r.metadata is not None and r.metadata.cost_breakdown is not None,
)
assert row is not None, f"{wire}: no spend row landed"
assert row.spend is not None and cost_rows.approx_equal(row.spend, expected.total), (
f"{wire}: spend {row.spend} != expected {expected.total} "
f"(breakdown {row.breakdown.model_dump()})"
)
breakdown = row.breakdown
assert breakdown.input_cost is not None and cost_rows.approx_equal(
breakdown.input_cost, expected.input_cost
), (
f"{wire}: gross input_cost {breakdown.input_cost} != expected {expected.input_cost}; "
"cached/written tokens billed at the input rate"
)
assert breakdown.output_cost is not None and cost_rows.approx_equal(
breakdown.output_cost, expected.output_cost
), f"{wire}: output_cost {breakdown.output_cost} != expected {expected.output_cost}"
prompt_tokens, completion_tokens = expected_token_columns(model, case)
assert row.prompt_tokens == prompt_tokens, (
f"{wire}: prompt_tokens {row.prompt_tokens} != {prompt_tokens}"
)
assert row.completion_tokens == completion_tokens, (
f"{wire}: completion_tokens {row.completion_tokens} != {completion_tokens}"
)
cost_rows.assert_total_is_sum_of_components(row)
@pytest.mark.covers("quota_management.spend_tracking.scripted_wire.logs_cost")
def test_anthropic_streamed_usage_bills_each_component(
self, client: CostCalcClient, resources: ResourceManager, scoped_key: str
) -> None:
map_key, usage = _WIRE_USAGE["anthropic_messages"]
model = _MODELS[map_key]
case = Case(name="stream", usage=usage, stream=True)
marker = unique_marker()
model_name, _handle = register_scenario_deployment(client, resources, model, case, marker)
response = client.proxy.transport.send(
"/chat/completions",
headers=client.proxy.transport.bearer(scoped_key),
json=ChatBody(
model=model_name,
messages=[ChatMessage(role="user", content=f"{marker} scripted anthropic stream")],
stream=True,
stream_options=ChatStreamOptions(include_usage=True),
),
stream=True,
)
assert response.ok, f"anthropic stream: proxy returned {response.status_code}: {response.body[:400]}"
assert response.stream_done, "anthropic stream did not reach its terminal event"
assert response.stream_error is None, f"stream carried an error event: {response.stream_error}"
expected = expected_breakdown(model, case)
row = cost_rows.poll_cost_row_where(
client.proxy,
scoped_key,
lambda r: r.metadata is not None and r.metadata.cost_breakdown is not None,
)
assert row is not None, "anthropic stream: no spend row landed"
assert row.spend is not None and cost_rows.approx_equal(row.spend, expected.total), (
f"anthropic stream: spend {row.spend} != expected {expected.total} "
f"(breakdown {row.breakdown.model_dump()})"
)
cost_rows.assert_total_is_sum_of_components(row)

352
tests/e2e/cost_map.json Normal file
View file

@ -0,0 +1,352 @@
{
"claude-haiku-4-5": {
"cache_creation_input_token_cost": 0.00021,
"cache_creation_input_token_cost_above_1hr": 0.00028000000000000003,
"cache_read_input_token_cost": 7e-06,
"input_cost_per_token": 7.000000000000001e-05,
"litellm_provider": "anthropic",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.00014000000000000001,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"claude-opus-5": {
"cache_creation_input_token_cost": 0.00015000000000000001,
"cache_creation_input_token_cost_above_1hr": 0.0002,
"cache_read_input_token_cost": 4.9999999999999996e-06,
"input_cost_per_token": 5e-05,
"litellm_provider": "anthropic",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.0001,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"claude-sonnet-5": {
"cache_creation_input_token_cost": 0.00018,
"cache_creation_input_token_cost_above_1hr": 0.00024000000000000003,
"cache_read_input_token_cost": 6e-06,
"input_cost_per_token": 6.000000000000001e-05,
"litellm_provider": "anthropic",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.00012000000000000002,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"fireworks_ai/deepseek-v4p1-flash": {
"cache_read_input_token_cost": 1.4e-05,
"input_cost_per_token": 0.00014000000000000001,
"litellm_provider": "fireworks_ai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.00028000000000000003,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"fireworks_ai/kimi-k3": {
"cache_read_input_token_cost": 1.2e-05,
"input_cost_per_token": 0.00012000000000000002,
"litellm_provider": "fireworks_ai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.00024000000000000003,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"fireworks_ai/qwen3p8-max": {
"cache_read_input_token_cost": 1.3e-05,
"input_cost_per_token": 0.00013000000000000002,
"litellm_provider": "fireworks_ai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_token": 0.00026000000000000003,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"gemini/gemini-3.1-pro-preview": {
"cache_read_input_token_cost": 9e-06,
"input_cost_per_audio_token": 0.00054,
"input_cost_per_token": 9e-05,
"input_cost_per_token_above_200k_tokens": 0.00072,
"input_cost_per_token_flex": 0.000135,
"input_cost_per_token_priority": 0.000153,
"litellm_provider": "gemini",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 0.0006299999999999999,
"output_cost_per_reasoning_token": 0.00045000000000000004,
"output_cost_per_token": 0.00018,
"output_cost_per_token_above_200k_tokens": 0.0008100000000000001,
"output_cost_per_token_flex": 0.00022500000000000002,
"output_cost_per_token_priority": 0.000243,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true,
"web_search_billing_unit": "per_query"
},
"gemini/gemini-3.8-flash": {
"cache_read_input_token_cost": 8e-06,
"input_cost_per_audio_token": 0.00048,
"input_cost_per_token": 8e-05,
"input_cost_per_token_above_200k_tokens": 0.00064,
"input_cost_per_token_flex": 0.00012,
"input_cost_per_token_priority": 0.000136,
"litellm_provider": "gemini",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 0.00056,
"output_cost_per_reasoning_token": 0.0004,
"output_cost_per_token": 0.00016,
"output_cost_per_token_above_200k_tokens": 0.00072,
"output_cost_per_token_flex": 0.0002,
"output_cost_per_token_priority": 0.000216,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true,
"web_search_billing_unit": "per_query"
},
"gpt-5.3-codex": {
"cache_read_input_token_cost": 3e-06,
"input_cost_per_token": 3.0000000000000004e-05,
"input_cost_per_token_above_200k_tokens": 0.00024000000000000003,
"input_cost_per_token_flex": 4.5e-05,
"input_cost_per_token_priority": 5.1e-05,
"litellm_provider": "openai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "responses",
"output_cost_per_reasoning_token": 0.00015000000000000001,
"output_cost_per_token": 6.000000000000001e-05,
"output_cost_per_token_above_200k_tokens": 0.00027,
"output_cost_per_token_flex": 7.500000000000001e-05,
"output_cost_per_token_priority": 8.099999999999999e-05,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"gpt-5.4-mini": {
"cache_creation_input_token_cost": 0.00012,
"cache_creation_input_token_cost_above_1hr": 0.00016,
"cache_read_input_token_cost": 4e-06,
"input_cost_per_audio_token": 0.00024,
"input_cost_per_token": 4e-05,
"input_cost_per_token_above_200k_tokens": 0.00032,
"input_cost_per_token_flex": 6e-05,
"input_cost_per_token_priority": 6.8e-05,
"litellm_provider": "openai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 0.00028,
"output_cost_per_reasoning_token": 0.0002,
"output_cost_per_token": 8e-05,
"output_cost_per_token_above_200k_tokens": 0.00036,
"output_cost_per_token_flex": 0.0001,
"output_cost_per_token_priority": 0.000108,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"gpt-5.5-pro": {
"cache_read_input_token_cost": 2e-06,
"input_cost_per_token": 2e-05,
"input_cost_per_token_above_200k_tokens": 0.00016,
"input_cost_per_token_flex": 3e-05,
"input_cost_per_token_priority": 3.4e-05,
"litellm_provider": "openai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "responses",
"output_cost_per_reasoning_token": 0.0001,
"output_cost_per_token": 4e-05,
"output_cost_per_token_above_200k_tokens": 0.00018,
"output_cost_per_token_flex": 5e-05,
"output_cost_per_token_priority": 5.4e-05,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"gpt-5.6": {
"cache_creation_input_token_cost": 3e-05,
"cache_creation_input_token_cost_above_1hr": 4e-05,
"cache_read_input_token_cost": 1e-06,
"input_cost_per_audio_token": 6e-05,
"input_cost_per_token": 1e-05,
"input_cost_per_token_above_200k_tokens": 8e-05,
"input_cost_per_token_flex": 1.5e-05,
"input_cost_per_token_priority": 1.7e-05,
"litellm_provider": "openai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 7e-05,
"output_cost_per_reasoning_token": 5e-05,
"output_cost_per_token": 2e-05,
"output_cost_per_token_above_200k_tokens": 9e-05,
"output_cost_per_token_flex": 2.5e-05,
"output_cost_per_token_priority": 2.7e-05,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"together_ai/moonshotai/Kimi-K3": {
"cache_creation_input_token_cost": 0.00030000000000000003,
"cache_creation_input_token_cost_above_1hr": 0.0004,
"cache_read_input_token_cost": 9.999999999999999e-06,
"input_cost_per_audio_token": 0.0006000000000000001,
"input_cost_per_token": 0.0001,
"input_cost_per_token_above_200k_tokens": 0.0008,
"input_cost_per_token_flex": 0.00015000000000000001,
"input_cost_per_token_priority": 0.00017,
"litellm_provider": "together_ai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 0.0006999999999999999,
"output_cost_per_reasoning_token": 0.0005,
"output_cost_per_token": 0.0002,
"output_cost_per_token_above_200k_tokens": 0.0009000000000000001,
"output_cost_per_token_flex": 0.00025,
"output_cost_per_token_priority": 0.00027,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
},
"together_ai/zai-org/GLM-5.3": {
"cache_creation_input_token_cost": 0.00033,
"cache_creation_input_token_cost_above_1hr": 0.00044,
"cache_read_input_token_cost": 1.1e-05,
"input_cost_per_audio_token": 0.00066,
"input_cost_per_token": 0.00011,
"input_cost_per_token_above_200k_tokens": 0.00088,
"input_cost_per_token_flex": 0.000165,
"input_cost_per_token_priority": 0.000187,
"litellm_provider": "together_ai",
"max_input_tokens": 2000000,
"max_output_tokens": 128000,
"max_tokens": 128000,
"mode": "chat",
"output_cost_per_audio_token": 0.00077,
"output_cost_per_reasoning_token": 0.00055,
"output_cost_per_token": 0.00022,
"output_cost_per_token_above_200k_tokens": 0.00099,
"output_cost_per_token_flex": 0.000275,
"output_cost_per_token_priority": 0.000297,
"search_context_cost_per_query": {
"search_context_size_high": 0.03,
"search_context_size_low": 0.01,
"search_context_size_medium": 0.02
},
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_web_search": true
}
}

View file

@ -63,3 +63,5 @@
- {id: quota_management.spend_tracking.key_attribution.health_rows_keep_service_account, module: quota_management, tier: P1, behavior: spend_tracking, variant: key_attribution, assertions: [health_rows_keep_service_account], exercised_on: [chat_completions], source: "proxy/health_check.py", rationale: "A /health probe's spend row stays keyed by the literal litellm-internal-health-check service account rather than a hash of it, so health spend never appears as an unattributed key"}
- {id: quota_management.spend_tracking.key_attribution.retrieve_batch_cost_joins_retrieving_key, module: quota_management, tier: P1, behavior: spend_tracking, variant: key_attribution, assertions: [retrieve_batch_cost_joins_retrieving_key], exercised_on: [batches], source: "proxy/batches_endpoints/endpoints.py", rationale: "The retrieve that first sees a batch in a terminal state prices it inline and writes its {provider_batch_id}_batch_cost row against the retrieving key, so the batch each run creates is one OpenAI fails at validation within seconds and the test retrieves it by its raw provider id with the same key until it is failed; a raw id is never owned by the CheckBatchCost poller, and the row must carry that key's token hash and alias"}
- {id: quota_management.spend_tracking.key_attribution.poller_batch_cost_joins_creating_key, module: quota_management, tier: P1, behavior: spend_tracking, variant: key_attribution, assertions: [poller_batch_cost_joins_creating_key], exercised_on: [batches], source: "enterprise/litellm_enterprise/proxy/common_utils/check_batch_cost.py", rationale: "The CheckBatchCost poller bills a completed, positive-cost batch created through a unified id against the key that created it, a different writer from the inline retrieve. No test claims this cell yet: OpenAI's completion window is 24h and both e2e stacks boot a fresh Postgres per build, so a completed batch is out of one run's reach and the managed list never shows an earlier run's batch; the cell stays visible as a gap until a run can hand a completed batch to the poller"}
- {id: quota_management.spend_tracking.cost_matrix.logs_cost, module: quota_management, tier: P1, behavior: spend_tracking, variant: cost_matrix, assertions: [logs_cost], exercised_on: [chat_completions], source: "litellm_core_utils/llm_cost_calc/utils.py", rationale: "A scripted-usage call through the cost-map proxy bills every reported token kind at the deployment's test-map rate (input, output, cache read, 5m/1h cache write, reasoning, audio, above-threshold tiers, flex/priority service tiers, web search, response-model override) and lands on the row's cost_breakdown, streamed or not"}
- {id: quota_management.spend_tracking.scripted_wire.logs_cost, module: quota_management, tier: P1, behavior: spend_tracking, variant: scripted_wire, assertions: [logs_cost], exercised_on: [chat_completions], source: "litellm_core_utils/llm_cost_calc/utils.py", rationale: "Each provider wire shape (openai chat, responses, anthropic messages, gemini generateContent, together, fireworks) parses usage into the same spend components: the gross input cost is fresh tokens at the input rate plus each cache/audio component at its own rate, streamed anthropic included"}

View file

@ -145,6 +145,22 @@ WEEKLY_ANOMALY_OPT_IN_ENV = "E2E_WEEKLY_ANOMALY"
MANAGED_FILES_OPT_IN_ENV = "E2E_MANAGED_FILES_STACK"
PROMPT_CACHING_OPT_IN_ENV = "E2E_PROMPT_CACHING_STACK"
REDIS_CHAOS_OPT_IN_ENV = "E2E_REDIS_CHAOS"
# The cost_calculation suite needs a proxy booted with LITELLM_MODEL_COST_MAP_URL
# pointing at tests/e2e/cost_map.json (its whole map is test-owned rates) plus a
# scripted-provider sidecar; deselected unless the opt-in env var is set.
COST_MAP_OPT_IN_ENV = "E2E_COST_MAP_STACK"
# Base URL of the proxy running the test cost map. Defaults to the shared proxy
# so a local run only has to set the opt-in and boot the proxy accordingly.
COST_MAP_PROXY_URL = os.environ.get("E2E_COST_MAP_PROXY_URL", PROXY_BASE_URL).rstrip("/")
# Where the test runner reaches the scripted-provider sidecar's control API.
SCRIPTED_PROVIDER_CONTROL_URL = os.environ.get(
"E2E_SCRIPTED_PROVIDER_CONTROL_URL", "http://127.0.0.1:9100"
).rstrip("/")
# The api_base root deployments register with: how the proxy (possibly in
# another container) reaches the sidecar's provider wire.
SCRIPTED_PROVIDER_PROXY_BASE = os.environ.get(
"E2E_SCRIPTED_PROVIDER_PROXY_BASE", SCRIPTED_PROVIDER_CONTROL_URL
).rstrip("/")
ANOMALY_SESSIONS = int(os.environ.get("E2E_ANOMALY_SESSIONS", "6"))
ANOMALY_TURNS_PER_SESSION = int(os.environ.get("E2E_ANOMALY_TURNS_PER_SESSION", "6"))
ANOMALY_TURN_ATTEMPTS = int(os.environ.get("E2E_ANOMALY_TURN_ATTEMPTS", "3"))

View file

@ -11,3 +11,4 @@ markers =
managed_files: needs a proxy running with require_managed_files enabled; deselected unless E2E_MANAGED_FILES_STACK is set
prompt_caching_stack: needs a proxy running with router_settings.optional_pre_call_checks including prompt_caching; deselected unless E2E_PROMPT_CACHING_STACK is set
redis_chaos: load test that pauses the proxy's Redis outright mid-run; needs a proxy booted from gateway/redis_chaos_ci_config.yml on the same host, and is deselected unless E2E_REDIS_CHAOS is set
cost_map_stack: needs a proxy whose whole cost map is tests/e2e/cost_map.json (LITELLM_MODEL_COST_MAP_URL) plus a scripted-provider sidecar; deselected unless E2E_COST_MAP_STACK is set