mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-29 01:42:19 +00:00
test(e2e): cover sail completion_window pricing and service_tier rewrite
Live proxy suites for the sail provider added in #42840: chat with service_tier flex, balanced and absent bills each window's cost-map rates read at run time and echoes no service_tier; /v1/responses with extra_body.metadata.completion_window balanced bills the _balanced keys; a flex-only sail model refuses the asap and balanced windows with sail's own 400 and leaves only a zero-spend failure row; openai keeps its own service_tier and bills the tier it reports back. MUTATIONS.md records the five production mutations, each killed, and the final green run. It is review evidence and not meant to merge.
This commit is contained in:
parent
d7bc17fa47
commit
bf60fa3960
8 changed files with 945 additions and 0 deletions
184
MUTATIONS.md
Normal file
184
MUTATIONS.md
Normal file
|
|
@ -0,0 +1,184 @@
|
|||
# Mutation evidence for the tests/e2e Sail service_tier suites (PR #42840)
|
||||
|
||||
Review evidence for branch `litellm_sail_tests_e2e`, created from tip d7bc17fa47ee9acacd441e6b3349c7aa9e1c6c62 (merge base 0c1c3e18d5250ec3a0e1e3f287e0b93e2906d900). This file is not meant to merge
|
||||
|
||||
## Tests on the branch
|
||||
|
||||
| Test | File |
|
||||
| --- | --- |
|
||||
| `TestSailServiceTier::test_chat_prices_each_service_tier_as_its_window_and_echoes_none` | tests/e2e/llm_translation/test_sail_service_tier_e2e.py |
|
||||
| `TestSailServiceTier::test_responses_extra_body_completion_window_prices_that_window` | tests/e2e/llm_translation/test_sail_service_tier_e2e.py |
|
||||
| `TestSailServiceTier::test_flex_only_model_serves_flex_and_refuses_the_windows_other_tiers_become` | tests/e2e/llm_translation/test_sail_service_tier_e2e.py |
|
||||
| `TestSailCompletionWindowPricing::test_each_service_tier_bills_its_windows_cost_map_rates` | tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py |
|
||||
| `TestSailCompletionWindowPricing::test_responses_completion_window_in_extra_body_bills_balanced_rates` | tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py |
|
||||
| `TestSailCompletionWindowPricing::test_flex_only_model_refuses_default_tier_and_bills_nothing` | tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py |
|
||||
| `TestOpenAIKeepsItsOwnServiceTier::test_openai_bills_the_tier_it_reports_and_refuses_balanced_unbilled` | tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py |
|
||||
|
||||
Shared helpers live in tests/e2e/completion_window_pricing.py (runtime cost-map lookup of the cheapest Sail model that carries base, `_balanced` and `_flex` rates, and of the cheapest flex-only Sail model, plus the 1e-9 relative tolerance helper). tests/e2e/models.py gained the six `*_balanced` and `*_flex` cost-map fields on `CostMapEntry`
|
||||
|
||||
## Harness
|
||||
|
||||
Proxy booted per tests/e2e/CONTRIBUTING.md from the repo checkout with `LITELLM_LOCAL_MODEL_COST_MAP=True`, real Postgres (`DATABASE_URL`), real Redis, `proxy_batch_write_at: 5`, and `SAIL_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY` from the 1Password Shared vault in the proxy environment. Every mutation run was: edit one production file, restart the proxy, run the tests named below with `--reruns 0`, restore the file, confirm `git status` is clean under `litellm/`. No test file was changed between the red and the green run of a mutation. The mutation runs happened before the final refactor that swapped `pytest.approx` for the typed `within_rel` / `all_within_rel` helpers; the assertion messages are identical, only the trailing `assert` explanation pytest prints differs. The final green run below is on the exact branch contents
|
||||
|
||||
At the merge base 0c1c3e18d5250ec3a0e1e3f287e0b93e2906d900 there is no `sail` provider, so every Sail test fails at `/model/new` before making a request. The mutations below are the meaningful bar: each one leaves the provider registered and breaks one behaviour the PR introduces
|
||||
|
||||
## Kill table: 5 killed / 5 applied
|
||||
|
||||
| Id | File | Mutation | Tests that went red |
|
||||
| --- | --- | --- | --- |
|
||||
| M1 | litellm/llms/openai_like/dynamic_config.py | `ServiceTier.BALANCED.value: "balanced"` -> `ServiceTier.BALANCED.value: "flex"` | `test_flex_only_model_serves_flex_and_refuses_the_windows_other_tiers_become` |
|
||||
| M2 | litellm/cost_calculator.py | `_provider_bills_by_completion_window` body -> `return False` | `test_responses_extra_body_completion_window_prices_that_window`, `test_responses_completion_window_in_extra_body_bills_balanced_rates` |
|
||||
| M3 | litellm/litellm_core_utils/llm_cost_calc/utils.py | removed `ServiceTier.BALANCED.value: ServiceTier.BALANCED.value,` from `_SERVICE_TIER_TO_COST_KEY_SUFFIX` | `test_chat_prices_each_service_tier_as_its_window_and_echoes_none`, `test_each_service_tier_bills_its_windows_cost_map_rates` |
|
||||
| M4 | litellm/llms/openai_like/dynamic_config.py | `"default": "asap"` -> `"default": "flex"` | `test_flex_only_model_serves_flex_and_refuses_the_windows_other_tiers_become`, `test_flex_only_model_refuses_default_tier_and_bills_nothing` |
|
||||
| M5 | litellm/cost_calculator.py | non-window-provider branch of `_service_tier_billed_by_completion_window`: `return service_tier` -> `return ServiceTier.FLEX.value` | `test_openai_bills_the_tier_it_reports_and_refuses_balanced_unbilled` |
|
||||
|
||||
Every test on the branch is red under at least one of M1 to M5
|
||||
|
||||
## M1 balanced window sent as flex
|
||||
|
||||
```diff
|
||||
@@ -21,7 +21,7 @@ from .json_loader import SimpleProviderConfig
|
||||
- ServiceTier.BALANCED.value: "balanced",
|
||||
+ ServiceTier.BALANCED.value: "flex",
|
||||
```
|
||||
|
||||
Run: `pytest tests/e2e/llm_translation/test_sail_service_tier_e2e.py tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py -k "flex_only or each_service_tier_bills"`
|
||||
|
||||
Result: `1 failed, 2 passed, 2 deselected in 171.37s`
|
||||
|
||||
Failing assertion (test_sail_service_tier_e2e.py:217):
|
||||
|
||||
```
|
||||
AssertionError: expected sail's 400 for service_tier='balanced', got 200: {"id":"resp_01a0d42b-225c-7fb0-8265-71b0f6285c28",...,"model":"sail-flex-6bb4ace67ce7","object":"chat.completion",...,"usage":{"completion_tokens":64,"prompt_tokens":28,"total_tokens":92,...}}
|
||||
assert 200 == 400
|
||||
```
|
||||
|
||||
The flex-only Sail model accepts `completion_window=flex` and refuses `balanced`; with balanced rewritten to flex the provider served the call instead of refusing it. The two billing tests kept passing under this mutation because balanced and flex both still price through the cost-map suffix path, which is why the flex-only refusal test exists
|
||||
|
||||
Restore: `git status --short -- litellm/llms/openai_like/dynamic_config.py` empty
|
||||
|
||||
## M2 completion_window billing gate off
|
||||
|
||||
```diff
|
||||
@@ -982,7 +982,7 @@ def _provider_bills_by_completion_window(custom_llm_provider: str | None) -> boo
|
||||
- return provider is not None and provider.special_handling.get("service_tier_as_completion_window") is True
|
||||
+ return False
|
||||
```
|
||||
|
||||
Run: `pytest ... -k responses`
|
||||
|
||||
Result: `2 failed, 5 deselected in 107.53s`
|
||||
|
||||
Failing assertions:
|
||||
|
||||
```
|
||||
test_sail_service_tier_e2e.py:177: AssertionError: resp_QeexFiEG... priced 3.96e-06, expected 3.08e-06 from input_tokens=16 output_tokens=14 input_tokens_details=_ResponsesInputDetails(cached_tokens=0) output_tokens_details=_ResponsesOutputDetails(reasoning_tokens=12) at the balanced rates TierRates(input=7e-08, output=1.4e-07, cache_read=2e-08)
|
||||
```
|
||||
|
||||
```
|
||||
test_sail_completion_window_pricing_e2e.py:157: AssertionError: balanced row resp_vHYEtxB9... billed (input, output, cache_read, reasoning, total) (1.8e-06, 8.1e-06, 0.0, 7.74e-06, 9.9e-06), expected (1.4000000000000001e-06, 6.300000000000001e-06, 0.0, 6.020000000000001e-06, 7.7e-06) from prompt_tokens=20 cached_tokens=0 completion_tokens=45 reasoning_tokens=43 at the balanced rates TierRates(input=7e-08, output=1.4e-07, cache_read=2e-08)
|
||||
```
|
||||
|
||||
Both `/v1/responses` calls were billed at the base (asap) rates once the gate stopped recognising Sail's `completion_window`
|
||||
|
||||
Restore: the mutate script's own restore step tripped on its text-match guard, so the line was restored by hand and `git status --short -- litellm/cost_calculator.py` confirmed empty before the next run
|
||||
|
||||
## M3 balanced cost suffix dropped
|
||||
|
||||
```diff
|
||||
@@ -69,7 +69,6 @@ _SERVICE_TIER_SUFFIXES: Final[tuple[str, ...]] = tuple(
|
||||
- ServiceTier.BALANCED.value: ServiceTier.BALANCED.value,
|
||||
```
|
||||
|
||||
Run: `pytest ... -k "each_service_tier or chat_prices_each"`
|
||||
|
||||
Result: `2 failed in 130.02s`
|
||||
|
||||
Failing assertions:
|
||||
|
||||
```
|
||||
test_sail_service_tier_e2e.py:146: AssertionError: x-litellm-response-cost per tier {'flex': 4.78e-06, 'balanced': 7.92e-06, None: 8.64e-06} != each window's rates times its usage {'flex': 4.78e-06, 'balanced': 6.16e-06, None: 8.64e-06} (rates WindowPricedModel(model='sail/deepseek-ai/DeepSeek-V4-Flash-0731', asap=TierRates(input=9e-08, output=1.8e-07, cache_read=2e-08), balanced=TierRates(input=7e-08, output=1.4e-07, cache_read=2e-08), flex=TierRates(input=5e-08, output=9e-08, cache_read=1e-08)))
|
||||
```
|
||||
|
||||
```
|
||||
test_sail_completion_window_pricing_e2e.py:157: AssertionError: balanced row resp_01a0d432-5695-7a4b-93d6-1aaf5f91188f billed (input, output, cache_read, reasoning, total) (1.62e-06, 9.36e-06, 0.0, 9e-06, 1.098e-05), expected (1.26e-06, 7.280000000000001e-06, 0.0, 7.000000000000001e-06, 8.540000000000001e-06) from prompt_tokens=18 cached_tokens=0 completion_tokens=52 reasoning_tokens=50 at the balanced rates TierRates(input=7e-08, output=1.4e-07, cache_read=2e-08)
|
||||
```
|
||||
|
||||
Balanced chat calls fell back to base rates; flex and absent stayed correct, so exactly the balanced cell moved
|
||||
|
||||
Restore: restored by hand after the script's guard tripped, `git status` clean under `litellm/`
|
||||
|
||||
## M4 default tier sent as flex
|
||||
|
||||
```diff
|
||||
@@ -23,7 +23,7 @@ _SERVICE_TIER_TO_COMPLETION_WINDOW: Final[Mapping[str, str]] = MappingProxyType(
|
||||
- "default": "asap",
|
||||
+ "default": "flex",
|
||||
```
|
||||
|
||||
Run: `pytest ... -k flex_only`
|
||||
|
||||
Result: `2 failed in 160.92s`
|
||||
|
||||
Failing assertions:
|
||||
|
||||
```
|
||||
test_sail_service_tier_e2e.py:217: AssertionError: expected sail's 400 for service_tier='default', got 200: {"id":"resp_01a0d435-10b7-75b2-9bdf-7512013c8494",...,"model":"sail-flex-e9a7dbaa3bc9",...}
|
||||
assert 200 == 400
|
||||
```
|
||||
|
||||
```
|
||||
test_sail_completion_window_pricing_e2e.py:292: AssertionError: expected the provider's 400, got 200: {"id":"resp_01a0d436-865b-774b-a4fe-b08fe6e2e592",...,"model":"sail-flexonly-656400d84243",...}
|
||||
assert 200 == 400
|
||||
```
|
||||
|
||||
Restore: `git status --short -- litellm/llms/openai_like/dynamic_config.py` empty
|
||||
|
||||
## M5 window billing hook leaks flex onto other providers
|
||||
|
||||
```diff
|
||||
@@ -999,7 +999,7 @@ def _service_tier_billed_by_completion_window(
|
||||
- return service_tier
|
||||
+ return ServiceTier.FLEX.value
|
||||
```
|
||||
|
||||
Run: `pytest tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py -k openai`
|
||||
|
||||
Result: `1 failed in 88.49s`
|
||||
|
||||
Failing assertion:
|
||||
|
||||
```
|
||||
test_sail_completion_window_pricing_e2e.py:157: AssertionError: openai default row chatcmpl-ERgTzeELuKMNg9PbdAIuPhXmJDUBO billed (input, output, cache_read, reasoning, total) (4.75e-05, 0.000255, 0.0, 0.000105, 0.00030250000000000003), expected (9.5e-05, 0.00051, 0.0, 0.00021, 0.0006050000000000001) from prompt_tokens=19 cached_tokens=0 completion_tokens=17 reasoning_tokens=7 at the openai default rates TierRates(input=5e-06, output=3e-05, cache_read=5e-07)
|
||||
```
|
||||
|
||||
The untiered OpenAI row was priced at OpenAI's `_flex` cost-map rates and its `cost_breakdown.service_tier` read `flex`, so the Sail-only hook had leaked into every provider's billing. The Sail tests all stay green under M5, which is the point of keeping the OpenAI test
|
||||
|
||||
Restore: `git status --short -- litellm/cost_calculator.py` empty
|
||||
|
||||
## Final green run at the tip, production code unmutated
|
||||
|
||||
`git status --short -- litellm/` empty, proxy restarted, then
|
||||
|
||||
```
|
||||
LITELLM_PROXY_URL=http://localhost:4000 LITELLM_MASTER_KEY=sk-1234 .venv/bin/python -m pytest tests/e2e/llm_translation/test_sail_service_tier_e2e.py tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py --reruns 0 -p no:cacheprovider -v -rA --tb=short
|
||||
```
|
||||
|
||||
```
|
||||
PASSED tests/e2e/llm_translation/test_sail_service_tier_e2e.py::TestSailServiceTier::test_chat_prices_each_service_tier_as_its_window_and_echoes_none
|
||||
PASSED tests/e2e/llm_translation/test_sail_service_tier_e2e.py::TestSailServiceTier::test_responses_extra_body_completion_window_prices_that_window
|
||||
PASSED tests/e2e/llm_translation/test_sail_service_tier_e2e.py::TestSailServiceTier::test_flex_only_model_serves_flex_and_refuses_the_windows_other_tiers_become
|
||||
PASSED tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py::TestSailCompletionWindowPricing::test_each_service_tier_bills_its_windows_cost_map_rates
|
||||
PASSED tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py::TestSailCompletionWindowPricing::test_responses_completion_window_in_extra_body_bills_balanced_rates
|
||||
PASSED tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py::TestSailCompletionWindowPricing::test_flex_only_model_refuses_default_tier_and_bills_nothing
|
||||
PASSED tests/e2e/quota_management/spend_tracking/test_sail_completion_window_pricing_e2e.py::TestOpenAIKeepsItsOwnServiceTier::test_openai_bills_the_tier_it_reports_and_refuses_balanced_unbilled
|
||||
======================== 7 passed in 264.95s (0:04:24) =========================
|
||||
```
|
||||
|
||||
## Cells the live providers would not let the tests prove
|
||||
|
||||
OpenAI `service_tier: balanced` with the latest `openai/gpt-5.5` row on the account behind the vault key is refused by OpenAI itself: `Invalid value: 'balanced'. Supported values are: 'auto', 'default', 'fast', 'flex', and 'priority'.` The test therefore proves the tier is forwarded untouched (OpenAI's own 400 comes back, with a zero-spend failure row and no cost_breakdown) and that the untiered call is billed at the tier OpenAI reports back. The requested cell "openai balanced bills the same as no tier" is UNVERIFIED: no OpenAI model on this account accepts `balanced`
|
||||
|
||||
Anthropic `service_tier: balanced` versus no tier is UNVERIFIED: every request through the vault's Anthropic key fails with `Your credit balance is too low to access the Anthropic API. Please go to Plans & Billing to upgrade or purchase credits.`, so no Anthropic test is on the branch rather than a test that skips or passes vacuously
|
||||
|
||||
Sail `/v1/responses` returns `metadata.completion_window: "standard"` alongside `supercache_*` fields even when the request carried `balanced`, so the tests do not pin Sail's echoed metadata (a vendor fact) and assert the balanced billing instead, which M2 shows is driven by the requested window
|
||||
140
tests/e2e/completion_window_pricing.py
Normal file
140
tests/e2e/completion_window_pricing.py
Normal file
|
|
@ -0,0 +1,140 @@
|
|||
"""Runtime-priced Sail deployments for the completion-window (service_tier) suites.
|
||||
|
||||
Sail bills a request by the `metadata.completion_window` it ran under, and the
|
||||
cost map carries one rate set per window: the base keys for `asap`, `*_balanced`
|
||||
and `*_flex`. Every rate here is read off the proxy's own
|
||||
/public/litellm_model_cost_map at run time, so the tests compare a bill against
|
||||
the same numbers the gateway priced it from and never pin a vendor's price.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Mapping, Sequence
|
||||
from dataclasses import dataclass
|
||||
from typing import Final
|
||||
|
||||
from models import CostMapEntry
|
||||
|
||||
SAIL_PROVIDER: Final = "sail"
|
||||
SAIL_API_KEY: Final = "os.environ/SAIL_API_KEY"
|
||||
COST_REL_TOLERANCE: Final = 1e-9
|
||||
|
||||
|
||||
def within_rel(actual: float | None, expected: float) -> bool:
|
||||
"""Spend math agrees to one part in 1e9, the tightest tolerance float sums hold; no value never agrees."""
|
||||
return actual is not None and abs(actual - expected) <= abs(expected) * COST_REL_TOLERANCE
|
||||
|
||||
|
||||
def all_within_rel(actual: Sequence[float], expected: Sequence[float]) -> bool:
|
||||
"""Every component of one bill agrees with the matching component of another."""
|
||||
return len(actual) == len(expected) and all(within_rel(a, e) for a, e in zip(actual, expected, strict=True))
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class TierRates:
|
||||
"""The three per-token rates one window bills at."""
|
||||
|
||||
input: float
|
||||
output: float
|
||||
cache_read: float
|
||||
|
||||
def bill(
|
||||
self, *, prompt_tokens: int, cached_tokens: int, completion_tokens: int, reasoning_tokens: int
|
||||
) -> tuple[float, float, float, float, float]:
|
||||
"""(input_cost, output_cost, cache_read_cost, reasoning_cost, total_cost) the
|
||||
gateway must file for this usage: cached prompt tokens at the cache-read rate,
|
||||
the rest of the prompt at the input rate, every completion token (reasoning
|
||||
included) at the output rate, reasoning also reported on its own line."""
|
||||
cache_read_cost: Final = cached_tokens * self.cache_read
|
||||
input_cost: Final = (prompt_tokens - cached_tokens) * self.input + cache_read_cost
|
||||
output_cost: Final = completion_tokens * self.output
|
||||
return (
|
||||
input_cost,
|
||||
output_cost,
|
||||
cache_read_cost,
|
||||
reasoning_tokens * self.output,
|
||||
input_cost + output_cost,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class WindowPricedModel:
|
||||
"""A Sail cost-map row priced for every completion window."""
|
||||
|
||||
model: str
|
||||
asap: TierRates
|
||||
balanced: TierRates
|
||||
flex: TierRates
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class FlexOnlyModel:
|
||||
"""A Sail cost-map row priced for the flex window and nothing else."""
|
||||
|
||||
model: str
|
||||
flex: TierRates
|
||||
|
||||
|
||||
def _rates(input_rate: float | None, output_rate: float | None, cache_rate: float | None) -> TierRates | None:
|
||||
if input_rate is None or output_rate is None or cache_rate is None:
|
||||
return None
|
||||
return TierRates(input=input_rate, output=output_rate, cache_read=cache_rate)
|
||||
|
||||
|
||||
def _base(entry: CostMapEntry) -> TierRates | None:
|
||||
return _rates(entry.input_cost_per_token, entry.output_cost_per_token, entry.cache_read_input_token_cost)
|
||||
|
||||
|
||||
def _balanced(entry: CostMapEntry) -> TierRates | None:
|
||||
return _rates(
|
||||
entry.input_cost_per_token_balanced,
|
||||
entry.output_cost_per_token_balanced,
|
||||
entry.cache_read_input_token_cost_balanced,
|
||||
)
|
||||
|
||||
|
||||
def _flex(entry: CostMapEntry) -> TierRates | None:
|
||||
return _rates(
|
||||
entry.input_cost_per_token_flex, entry.output_cost_per_token_flex, entry.cache_read_input_token_cost_flex
|
||||
)
|
||||
|
||||
|
||||
def _sail_chat_rows(cost_map: Mapping[str, CostMapEntry]) -> tuple[tuple[str, CostMapEntry], ...]:
|
||||
return tuple(
|
||||
(model, entry)
|
||||
for model, entry in sorted(cost_map.items())
|
||||
if entry.litellm_provider == SAIL_PROVIDER and entry.mode == "chat" and entry.deprecation_date is None
|
||||
)
|
||||
|
||||
|
||||
def _window_priced(model: str, entry: CostMapEntry) -> WindowPricedModel | None:
|
||||
asap: Final = _base(entry)
|
||||
balanced: Final = _balanced(entry)
|
||||
flex: Final = _flex(entry)
|
||||
if asap is None or balanced is None or flex is None:
|
||||
return None
|
||||
totals: Final = {asap.input + asap.output, balanced.input + balanced.output, flex.input + flex.output}
|
||||
return WindowPricedModel(model=model, asap=asap, balanced=balanced, flex=flex) if len(totals) == 3 else None
|
||||
|
||||
|
||||
def cheapest_window_priced_model(cost_map: Mapping[str, CostMapEntry]) -> WindowPricedModel:
|
||||
"""The Sail chat row with distinct asap, balanced and flex rate sets whose asap
|
||||
input rate is lowest, so the tier comparison runs on the cheapest eligible model."""
|
||||
priced: Final = tuple(
|
||||
priced_row
|
||||
for model, entry in _sail_chat_rows(cost_map)
|
||||
if (priced_row := _window_priced(model, entry)) is not None
|
||||
)
|
||||
assert priced, f"no sail chat row carries distinct asap, balanced and flex rates: {sorted(cost_map)}"
|
||||
return min(priced, key=lambda candidate: candidate.asap.input)
|
||||
|
||||
|
||||
def cheapest_flex_only_model(cost_map: Mapping[str, CostMapEntry]) -> FlexOnlyModel:
|
||||
"""The Sail chat row that carries flex rates and no balanced rates, cheapest first."""
|
||||
priced: Final = tuple(
|
||||
FlexOnlyModel(model=model, flex=flex)
|
||||
for model, entry in _sail_chat_rows(cost_map)
|
||||
if (flex := _flex(entry)) is not None and _balanced(entry) is None
|
||||
)
|
||||
assert priced, f"no sail chat row is priced for flex only: {sorted(cost_map)}"
|
||||
return min(priced, key=lambda candidate: candidate.flex.input)
|
||||
|
|
@ -13,6 +13,8 @@
|
|||
- {id: llm.chat_completions.openai.vision.nonstream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: openai, capability: vision, streaming: nonstream, assertions: [works], source: "model_prices json", rationale: "gpt-4o vision; high usage"}
|
||||
- {id: llm.chat_completions.openai.prompt_cache_5m.nonstream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: openai, capability: prompt_cache_5m, streaming: nonstream, assertions: [works], source: "model_prices json", rationale: "Prompt caching cost optimization"}
|
||||
- {id: llm.chat_completions.openai.service_tier.nonstream.works, module: llm, tier: P1, subject_endpoint: chat_completions, route: openai, capability: service_tier, streaming: nonstream, assertions: [works], source: "OpenAI service_tier param", rationale: "OpenAI scale-tier request option is forwarded and echoed"}
|
||||
- {id: llm.chat_completions.sail.service_tier.nonstream.works, module: llm, tier: P1, subject_endpoint: chat_completions, route: sail, capability: service_tier, streaming: nonstream, assertions: [works], source: "llms/openai_like/dynamic_config.py", rationale: "service_tier flex and balanced are rewritten to metadata.completion_window for sail, priced at that window's cost-map rates, and never echoed back as service_tier; a window the model does not offer surfaces as sail's own 400 naming metadata.completion_window (#42840)"}
|
||||
- {id: llm.responses.sail.service_tier.nonstream.works, module: llm, tier: P1, subject_endpoint: responses, route: sail, capability: service_tier, streaming: nonstream, assertions: [works], source: "llms/openai_like/dynamic_config.py", rationale: "A completion_window sent inside extra_body.metadata on /v1/responses reaches sail and prices the call at that window's cost-map rates (#42840)"}
|
||||
- {id: llm.chat_completions.openai.thinking.nonstream.works, module: llm, tier: P1, subject_endpoint: chat_completions, route: openai, capability: thinking, streaming: nonstream, assertions: [works], source: "model_prices json", rationale: "o-series reasoning; emerging"}
|
||||
- {id: llm.chat_completions.openai.structured_output.nonstream.works, module: llm, tier: P1, subject_endpoint: chat_completions, route: openai, capability: structured_output, streaming: nonstream, assertions: [works], source: "model_prices json", rationale: "response_schema extraction"}
|
||||
- {id: llm.chat_completions.anthropic.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: anthropic, capability: basic, streaming: nonstream, assertions: [works], source: "proxy_server.py:8455", rationale: "P0 route translated to Anthropic"}
|
||||
|
|
|
|||
|
|
@ -60,6 +60,9 @@
|
|||
- {id: quota_management.spend_tracking.stream_cache_read.bills_cache_read_rate, module: quota_management, tier: P1, behavior: spend_tracking, variant: stream_cache_read, assertions: [bills_cache_read_rate], exercised_on: [chat_completions], source: "litellm_core_utils/streaming_chunk_builder_utils.py", rationale: "A streamed call's reassembled usage keeps the cached-token detail so cache reads bill at the cache-read discount, not full input price (#34812)"}
|
||||
- {id: quota_management.spend_tracking.messages_bridge.keeps_cache_tokens, module: quota_management, tier: P1, behavior: spend_tracking, variant: messages_bridge, assertions: [keeps_cache_tokens], exercised_on: [messages], source: "llms/anthropic/experimental_pass_through/responses_adapters/handler.py", rationale: "A /v1/messages request served by a Responses-only OpenAI model keeps its cache-read tokens and their discounted billing across the bridge (#34957)"}
|
||||
- {id: quota_management.spend_tracking.service_tier.bills_tier_rates, module: quota_management, tier: P1, behavior: spend_tracking, variant: service_tier, assertions: [bills_tier_rates], exercised_on: [chat_completions], source: "cost_calculator.py", rationale: "A priority service_tier call bills input, output, and reasoning at the deployment's *_priority rates and records the tier on the row (#35923, #35925)"}
|
||||
- {id: quota_management.spend_tracking.service_tier.bills_served_tier, module: quota_management, tier: P1, behavior: spend_tracking, variant: service_tier, assertions: [bills_served_tier], exercised_on: [chat_completions], source: "cost_calculator.py", rationale: "An untiered OpenAI call is priced at the base rates on the tier OpenAI reports back, and a tier OpenAI refuses is forwarded to it untouched by the sail completion_window rewrite and leaves a zero-spend failure row (#42840)"}
|
||||
- {id: quota_management.spend_tracking.completion_window.bills_window_rates, module: quota_management, tier: P1, behavior: spend_tracking, variant: completion_window, assertions: [bills_window_rates], exercised_on: [chat_completions, responses], source: "cost_calculator.py", rationale: "A sail call bills input, output, cache read and reasoning at the cost-map rates of the completion_window it ran under: flex and balanced from service_tier on chat, balanced from extra_body.metadata.completion_window on responses, base rates when neither is set (#42840)"}
|
||||
- {id: quota_management.spend_tracking.completion_window.rejected_window_bills_nothing, module: quota_management, tier: P1, behavior: spend_tracking, variant: completion_window, assertions: [rejected_window_bills_nothing], exercised_on: [chat_completions], source: "cost_calculator.py", rationale: "service_tier=default reaches a flex-only sail model as completion_window asap, the provider's 400 is relayed with param metadata.completion_window, and the only row filed is a zero-spend failure row (#42840)"}
|
||||
- {id: quota_management.spend_tracking.cost_headers.additive_components, module: quota_management, tier: P1, behavior: spend_tracking, variant: cost_headers, assertions: [additive_components], exercised_on: [chat_completions], source: "proxy/common_request_processing.py", rationale: "The x-litellm-response-cost-* component headers sum to the total, input covers only fresh tokens, and reasoning stays a subset of output (#36965)"}
|
||||
- {id: quota_management.spend_tracking.passthrough_stream.injects_usage_cost, module: quota_management, tier: P1, behavior: spend_tracking, variant: passthrough_stream, assertions: [injects_usage_cost], exercised_on: [openai_passthrough], source: "proxy/pass_through_endpoints/streaming_handler.py", rationale: "With include_cost_in_streaming_usage on, the /openai passthrough's final streaming usage frame carries the proxy-computed cost (#36503). Uncovered: the flag is only settable in litellm_settings, and the shared e2e stack does not turn it on yet"}
|
||||
- {id: quota_management.spend_tracking.websearch_interception.bills_under_request_session, module: quota_management, tier: P1, behavior: spend_tracking, variant: websearch_interception, assertions: [bills_under_request_session], exercised_on: [messages], source: "integrations/websearch_interception/handler.py", fail_before_fix: proven, rationale: "A web_search server tool the proxy intercepts into litellm.asearch writes its own asearch spend row, and that row carries the parent request's session_id so the session view counts the search and its cost next to the turn that triggered it (LIT-8063)"}
|
||||
|
|
|
|||
|
|
@ -56,6 +56,7 @@ LlmRoute = Literal[
|
|||
"gemini",
|
||||
"hosted_vllm",
|
||||
"openai",
|
||||
"sail",
|
||||
"together_ai",
|
||||
"vertex",
|
||||
"xiaomi_mimo",
|
||||
|
|
|
|||
231
tests/e2e/llm_translation/test_sail_service_tier_e2e.py
Normal file
231
tests/e2e/llm_translation/test_sail_service_tier_e2e.py
Normal file
|
|
@ -0,0 +1,231 @@
|
|||
"""Live e2e: the gateway translates `service_tier` into Sail's `metadata.completion_window`.
|
||||
|
||||
Sail is a JSON-registered OpenAI-compatible provider that has no `service_tier`.
|
||||
A caller's `flex` or `balanced` becomes `metadata.completion_window` on the wire
|
||||
and `default` becomes `asap` (#42840). The proof a real provider gives: a window
|
||||
the model does not offer comes back as Sail's own 400 naming the rewritten
|
||||
field, a served call carries no `service_tier` at all, and its
|
||||
`x-litellm-response-cost` is that window's cost-map rates times the usage Sail
|
||||
reported. Every rate is read off the proxy's cost map at run time, never pinned.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from typing import Final
|
||||
|
||||
import pytest
|
||||
from pydantic import BaseModel
|
||||
|
||||
from completion_window_pricing import (
|
||||
SAIL_API_KEY,
|
||||
TierRates,
|
||||
cheapest_flex_only_model,
|
||||
cheapest_window_priced_model,
|
||||
within_rel,
|
||||
)
|
||||
from e2e_config import unique_marker
|
||||
from e2e_http import StreamingResponse
|
||||
from lifecycle import ResourceManager
|
||||
from models import ChatBody, ChatMessage, ChatResponse, LiteLLMParamsBody
|
||||
from passthrough_client import PassthroughClient
|
||||
|
||||
pytestmark = pytest.mark.e2e
|
||||
|
||||
MAX_COMPLETION_TOKENS: Final = 64
|
||||
|
||||
|
||||
class _WindowMetadata(BaseModel):
|
||||
completion_window: str
|
||||
|
||||
|
||||
class _WindowExtraBody(BaseModel):
|
||||
metadata: _WindowMetadata
|
||||
|
||||
|
||||
class _WindowedResponsesBody(BaseModel):
|
||||
model: str
|
||||
input: str
|
||||
max_output_tokens: int
|
||||
extra_body: _WindowExtraBody
|
||||
cache: dict[str, bool] = {"no-cache": True}
|
||||
|
||||
|
||||
class _ResponsesInputDetails(BaseModel):
|
||||
cached_tokens: int = 0
|
||||
|
||||
|
||||
class _ResponsesOutputDetails(BaseModel):
|
||||
reasoning_tokens: int = 0
|
||||
|
||||
|
||||
class _ResponsesUsage(BaseModel):
|
||||
input_tokens: int
|
||||
output_tokens: int
|
||||
input_tokens_details: _ResponsesInputDetails = _ResponsesInputDetails()
|
||||
output_tokens_details: _ResponsesOutputDetails = _ResponsesOutputDetails()
|
||||
|
||||
|
||||
class _ResponsesObject(BaseModel):
|
||||
id: str
|
||||
usage: _ResponsesUsage
|
||||
|
||||
|
||||
class _ErrorDetail(BaseModel):
|
||||
message: str
|
||||
type: str
|
||||
param: str | None = None
|
||||
code: str
|
||||
|
||||
|
||||
class _ErrorBody(BaseModel):
|
||||
error: _ErrorDetail
|
||||
|
||||
|
||||
def _sail_deployment(client: PassthroughClient, resources: ResourceManager, prefix: str, backend: str) -> str:
|
||||
model = f"{prefix}-{unique_marker()}"
|
||||
model_id = client.proxy.create_model(model, LiteLLMParamsBody(model=backend, api_key=SAIL_API_KEY))
|
||||
resources.defer(lambda: client.proxy.delete_model(model_id))
|
||||
return model
|
||||
|
||||
|
||||
def _chat_total(rates: TierRates, chat: ChatResponse) -> float:
|
||||
usage = chat.usage
|
||||
assert usage is not None and usage.prompt_tokens is not None and usage.completion_tokens is not None, (
|
||||
f"sail chat {chat.id} carried no usage: {chat}"
|
||||
)
|
||||
prompt_details = usage.prompt_tokens_details
|
||||
completion_details = usage.completion_tokens_details
|
||||
return rates.bill(
|
||||
prompt_tokens=usage.prompt_tokens,
|
||||
cached_tokens=(prompt_details.cached_tokens or 0) if prompt_details else 0,
|
||||
completion_tokens=usage.completion_tokens,
|
||||
reasoning_tokens=(completion_details.reasoning_tokens or 0) if completion_details else 0,
|
||||
)[4]
|
||||
|
||||
|
||||
def _served_chat(client: PassthroughClient, key: str, model: str, service_tier: str | None) -> StreamingResponse:
|
||||
outcome = client.proxy.transport.send(
|
||||
"/chat/completions",
|
||||
headers=client.proxy.transport.bearer(key),
|
||||
json=ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=f"{unique_marker()} Reply with the single word ok.")],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
service_tier=service_tier,
|
||||
),
|
||||
)
|
||||
assert outcome.status_code == 200, (
|
||||
f"sail chat with service_tier={service_tier!r} failed {outcome.status_code}: {outcome.body}"
|
||||
)
|
||||
return outcome
|
||||
|
||||
|
||||
class TestSailServiceTier:
|
||||
@pytest.mark.covers("llm.chat_completions.sail.service_tier.nonstream.works")
|
||||
def test_chat_prices_each_service_tier_as_its_window_and_echoes_none(
|
||||
self, client: PassthroughClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
priced = cheapest_window_priced_model(client.proxy.model_cost_map())
|
||||
model = _sail_deployment(client, resources, "sail-tier", priced.model)
|
||||
|
||||
outcomes = {tier: _served_chat(client, scoped_key, model, tier) for tier in ("flex", "balanced", None)}
|
||||
chats = {tier: ChatResponse.model_validate_json(outcome.body) for tier, outcome in outcomes.items()}
|
||||
assert {tier: chat.service_tier for tier, chat in chats.items()} == {
|
||||
"flex": None,
|
||||
"balanced": None,
|
||||
None: None,
|
||||
}, f"sail has no service_tier, yet the responses reported {[chat.service_tier for chat in chats.values()]}"
|
||||
|
||||
costs = {tier: outcome.response_cost for tier, outcome in outcomes.items()}
|
||||
expected = {
|
||||
"flex": _chat_total(priced.flex, chats["flex"]),
|
||||
"balanced": _chat_total(priced.balanced, chats["balanced"]),
|
||||
None: _chat_total(priced.asap, chats[None]),
|
||||
}
|
||||
assert costs.keys() == expected.keys() and all(within_rel(costs[t], expected[t]) for t in expected), (
|
||||
f"x-litellm-response-cost per tier {costs} != each window's rates times its usage {expected} "
|
||||
f"(rates {priced})"
|
||||
)
|
||||
assert len(set(costs.values())) == 3, f"the three windows did not price three different totals: {costs}"
|
||||
|
||||
@pytest.mark.covers("llm.responses.sail.service_tier.nonstream.works")
|
||||
def test_responses_extra_body_completion_window_prices_that_window(
|
||||
self, client: PassthroughClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
priced = cheapest_window_priced_model(client.proxy.model_cost_map())
|
||||
model = _sail_deployment(client, resources, "sail-resp", priced.model)
|
||||
|
||||
outcome = client.proxy.transport.send(
|
||||
"/v1/responses",
|
||||
headers=client.proxy.transport.bearer(scoped_key),
|
||||
json=_WindowedResponsesBody(
|
||||
model=model,
|
||||
input=f"{unique_marker()} Reply with the single word ok.",
|
||||
max_output_tokens=MAX_COMPLETION_TOKENS,
|
||||
extra_body=_WindowExtraBody(metadata=_WindowMetadata(completion_window="balanced")),
|
||||
),
|
||||
)
|
||||
assert outcome.status_code == 200, f"sail /v1/responses failed {outcome.status_code}: {outcome.body}"
|
||||
response = _ResponsesObject.model_validate_json(outcome.body)
|
||||
expected = priced.balanced.bill(
|
||||
prompt_tokens=response.usage.input_tokens,
|
||||
cached_tokens=response.usage.input_tokens_details.cached_tokens,
|
||||
completion_tokens=response.usage.output_tokens,
|
||||
reasoning_tokens=response.usage.output_tokens_details.reasoning_tokens,
|
||||
)[4]
|
||||
assert within_rel(outcome.response_cost, expected), (
|
||||
f"{response.id} priced {outcome.response_cost}, expected {expected} from {response.usage} at the "
|
||||
f"balanced rates {priced.balanced}"
|
||||
)
|
||||
|
||||
@pytest.mark.covers("llm.chat_completions.sail.service_tier.nonstream.works")
|
||||
def test_flex_only_model_serves_flex_and_refuses_the_windows_other_tiers_become(
|
||||
self, client: PassthroughClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
flex_only = cheapest_flex_only_model(client.proxy.model_cost_map())
|
||||
model = _sail_deployment(client, resources, "sail-flex", flex_only.model)
|
||||
|
||||
served = _served_chat(client, scoped_key, model, "flex")
|
||||
chat = ChatResponse.model_validate_json(served.body)
|
||||
expected = _chat_total(flex_only.flex, chat)
|
||||
assert within_rel(served.response_cost, expected), (
|
||||
f"{chat.id} priced {served.response_cost}, expected {expected} at the flex rates {flex_only.flex}"
|
||||
)
|
||||
|
||||
windows = {"default": "asap", "balanced": "balanced"}
|
||||
refusals = {tier: _refused_window(client, scoped_key, model, tier) for tier in windows}
|
||||
assert {tier: bool(re.search(rf'"{windows[tier]}"', seen)) for tier, seen in refusals.items()} == {
|
||||
"default": True,
|
||||
"balanced": True,
|
||||
}, f"service_tier must reach sail as completion_window {windows}; sail saw: {refusals}"
|
||||
|
||||
|
||||
def _refused_window(client: PassthroughClient, key: str, model: str, service_tier: str) -> str:
|
||||
"""Sail's own message for a completion_window the model does not offer, once the
|
||||
gateway has relayed it as a 400 naming the rewritten field."""
|
||||
outcome = client.proxy.transport.send(
|
||||
"/chat/completions",
|
||||
headers=client.proxy.transport.bearer(key),
|
||||
json=ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=f"{unique_marker()} Reply with the single word ok.")],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
service_tier=service_tier,
|
||||
),
|
||||
)
|
||||
assert outcome.status_code == 400, (
|
||||
f"expected sail's 400 for service_tier={service_tier!r}, got {outcome.status_code}: {outcome.body}"
|
||||
)
|
||||
error = _ErrorBody.model_validate_json(outcome.body).error
|
||||
relayed = re.fullmatch(
|
||||
r"litellm\.BadRequestError: SailException - (?P<upstream>.+)\n\n"
|
||||
rf"LiteLLM: model group '{re.escape(model)}' failed with the error above\. No fallback was attempted\.",
|
||||
error.message,
|
||||
re.DOTALL,
|
||||
)
|
||||
assert relayed is not None, f"the error is not sail's message relayed for {model}: {error.message!r}"
|
||||
assert (error.type, error.param, error.code) == ("invalid_request_error", "metadata.completion_window", "400"), (
|
||||
f"sail refused {error.param!r}, not the completion_window the gateway rewrote service_tier into: {error}"
|
||||
)
|
||||
return relayed["upstream"]
|
||||
|
|
@ -1155,6 +1155,12 @@ class CostMapEntry(BaseModel):
|
|||
input_cost_per_token: float | None = None
|
||||
output_cost_per_token: float | None = None
|
||||
cache_read_input_token_cost: float | None = None
|
||||
input_cost_per_token_balanced: float | None = None
|
||||
output_cost_per_token_balanced: float | None = None
|
||||
cache_read_input_token_cost_balanced: float | None = None
|
||||
input_cost_per_token_flex: float | None = None
|
||||
output_cost_per_token_flex: float | None = None
|
||||
cache_read_input_token_cost_flex: float | None = None
|
||||
supports_function_calling: bool | None = None
|
||||
supports_reasoning: bool | None = None
|
||||
supports_response_schema: bool | None = None
|
||||
|
|
|
|||
|
|
@ -0,0 +1,378 @@
|
|||
"""Live e2e: Sail bills each completion window at that window's cost-map rates.
|
||||
|
||||
Sail has no `service_tier`; it prices a request by `metadata.completion_window`
|
||||
(`asap`, `balanced`, `flex`). The gateway rewrites a caller's `service_tier` into
|
||||
that window for Sail only, and the cost calculator reads the window back off the
|
||||
request to pick the base, `*_balanced` or `*_flex` cost keys (#42840). Every rate
|
||||
these tests compare against is read off the proxy's own cost map at run time, and
|
||||
the three windows carry three distinct rate sets, so a bill computed from the
|
||||
wrong window cannot match. A non-Sail provider keeps its own `service_tier`
|
||||
semantics: OpenAI's row is priced on the tier OpenAI reports back, and a tier the
|
||||
provider refuses surfaces as the provider's 4xx with nothing billed.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from typing import Final
|
||||
|
||||
import pytest
|
||||
from pydantic import BaseModel, RootModel
|
||||
|
||||
from completion_window_pricing import (
|
||||
SAIL_API_KEY,
|
||||
TierRates,
|
||||
all_within_rel,
|
||||
cheapest_flex_only_model,
|
||||
cheapest_window_priced_model,
|
||||
within_rel,
|
||||
)
|
||||
from cost_rows import CostBreakdownRow, CostRow, poll_cost_row, register_priced_model
|
||||
from e2e_config import unique_marker
|
||||
from e2e_http import unwrap
|
||||
from lifecycle import ResourceManager
|
||||
from models import ChatBody, ChatMessage, ChatResponse, LiteLLMParamsBody, SpendLogsParams
|
||||
from proxy_client import ProxyClient
|
||||
from spend_e2e_client import SpendClient
|
||||
|
||||
pytestmark = pytest.mark.e2e
|
||||
|
||||
OPENAI_BACKEND: Final = "openai/gpt-5.5"
|
||||
OPENAI_API_KEY: Final = "os.environ/OPENAI_API_KEY"
|
||||
MAX_COMPLETION_TOKENS: Final = 64
|
||||
"""OpenAI validates `service_tier` against its own enum and rejects `balanced`
|
||||
(observed 2026-09-24: "Supported values are: 'auto', 'default', 'fast', 'flex',
|
||||
and 'priority'"), so the balanced leg on OpenAI can only prove the gateway forwards
|
||||
the tier untouched and bills nothing for the refusal."""
|
||||
|
||||
|
||||
class _WindowMetadata(BaseModel):
|
||||
completion_window: str
|
||||
|
||||
|
||||
class _WindowExtraBody(BaseModel):
|
||||
metadata: _WindowMetadata
|
||||
|
||||
|
||||
class _WindowedResponsesBody(BaseModel):
|
||||
model: str
|
||||
input: str
|
||||
max_output_tokens: int
|
||||
extra_body: _WindowExtraBody
|
||||
cache: dict[str, bool] = {"no-cache": True}
|
||||
|
||||
|
||||
class _ResponsesInputDetails(BaseModel):
|
||||
cached_tokens: int = 0
|
||||
|
||||
|
||||
class _ResponsesOutputDetails(BaseModel):
|
||||
reasoning_tokens: int = 0
|
||||
|
||||
|
||||
class _ResponsesUsage(BaseModel):
|
||||
input_tokens: int
|
||||
output_tokens: int
|
||||
input_tokens_details: _ResponsesInputDetails = _ResponsesInputDetails()
|
||||
output_tokens_details: _ResponsesOutputDetails = _ResponsesOutputDetails()
|
||||
|
||||
|
||||
class _ResponsesObject(BaseModel):
|
||||
id: str
|
||||
usage: _ResponsesUsage
|
||||
|
||||
|
||||
class _ErrorDetail(BaseModel):
|
||||
message: str
|
||||
type: str
|
||||
param: str | None = None
|
||||
code: str
|
||||
|
||||
|
||||
class _ErrorBody(BaseModel):
|
||||
error: _ErrorDetail
|
||||
|
||||
|
||||
class _StatusRowMetadata(BaseModel):
|
||||
status: str | None = None
|
||||
cost_breakdown: CostBreakdownRow | None = None
|
||||
|
||||
|
||||
class _StatusRow(BaseModel):
|
||||
request_id: str | None = None
|
||||
spend: float | None = None
|
||||
metadata: _StatusRowMetadata | None = None
|
||||
|
||||
|
||||
class _StatusRows(RootModel[list[_StatusRow]]):
|
||||
pass
|
||||
|
||||
|
||||
class _ChatUsage(BaseModel):
|
||||
"""The four token counts a chat bill is computed from."""
|
||||
|
||||
prompt_tokens: int
|
||||
cached_tokens: int
|
||||
completion_tokens: int
|
||||
reasoning_tokens: int
|
||||
|
||||
|
||||
def _chat_usage(chat: ChatResponse) -> _ChatUsage:
|
||||
usage = chat.usage
|
||||
assert usage is not None and usage.prompt_tokens is not None and usage.completion_tokens is not None, (
|
||||
f"chat response {chat.id} carried no usage: {chat}"
|
||||
)
|
||||
prompt_details = usage.prompt_tokens_details
|
||||
completion_details = usage.completion_tokens_details
|
||||
return _ChatUsage(
|
||||
prompt_tokens=usage.prompt_tokens,
|
||||
cached_tokens=(prompt_details.cached_tokens or 0) if prompt_details else 0,
|
||||
completion_tokens=usage.completion_tokens,
|
||||
reasoning_tokens=(completion_details.reasoning_tokens or 0) if completion_details else 0,
|
||||
)
|
||||
|
||||
|
||||
def _billed(row: CostRow) -> tuple[float, float, float, float, float]:
|
||||
breakdown = row.breakdown
|
||||
return (
|
||||
breakdown.input_cost or 0.0,
|
||||
breakdown.output_cost or 0.0,
|
||||
breakdown.cache_read_cost or 0.0,
|
||||
breakdown.reasoning_cost or 0.0,
|
||||
breakdown.total_cost or 0.0,
|
||||
)
|
||||
|
||||
|
||||
def _assert_billed_at(row: CostRow, rates: TierRates, usage: _ChatUsage, *, tier: str | None, window: str) -> None:
|
||||
expected = rates.bill(
|
||||
prompt_tokens=usage.prompt_tokens,
|
||||
cached_tokens=usage.cached_tokens,
|
||||
completion_tokens=usage.completion_tokens,
|
||||
reasoning_tokens=usage.reasoning_tokens,
|
||||
)
|
||||
assert (row.prompt_tokens, row.completion_tokens) == (usage.prompt_tokens, usage.completion_tokens), (
|
||||
f"{window} row {row.request_id} logged {row.prompt_tokens}/{row.completion_tokens} tokens, the response "
|
||||
f"reported {usage.prompt_tokens}/{usage.completion_tokens}"
|
||||
)
|
||||
assert all_within_rel(_billed(row), expected), (
|
||||
f"{window} row {row.request_id} billed (input, output, cache_read, reasoning, total) {_billed(row)}, "
|
||||
f"expected {expected} from {usage} at the {window} rates {rates}"
|
||||
)
|
||||
assert within_rel(row.spend, expected[4]), (
|
||||
f"{window} row {row.request_id} spend {row.spend} != its total_cost {expected[4]}"
|
||||
)
|
||||
assert row.breakdown.service_tier == tier, (
|
||||
f"{window} row {row.request_id} records pricing basis {row.breakdown.service_tier!r}, expected {tier!r}"
|
||||
)
|
||||
|
||||
|
||||
def _landed_row(proxy: ProxyClient, request_id: str) -> CostRow:
|
||||
row = poll_cost_row(proxy, request_id)
|
||||
assert row is not None, f"no spend row with a cost breakdown landed for {request_id}"
|
||||
return row
|
||||
|
||||
|
||||
def _row_shape(row: _StatusRow) -> tuple[float | None, str | None, bool]:
|
||||
metadata = row.metadata
|
||||
return (row.spend, metadata.status if metadata else None, bool(metadata and metadata.cost_breakdown))
|
||||
|
||||
|
||||
def _poll_key_rows(proxy: ProxyClient, key: str) -> tuple[_StatusRow, ...]:
|
||||
"""Every /spend/logs row filed under the key once at least one has landed."""
|
||||
landed = proxy.poll_logs_for_key(key)
|
||||
assert landed, f"no spend row landed for the key before the {proxy.poll_timeout}s deadline"
|
||||
return tuple(
|
||||
unwrap(
|
||||
proxy.transport.get(
|
||||
"/spend/logs",
|
||||
headers=proxy.transport.master,
|
||||
params=SpendLogsParams(api_key=key),
|
||||
response_type=_StatusRows,
|
||||
)
|
||||
).root
|
||||
)
|
||||
|
||||
|
||||
def _sail_chat(proxy: ProxyClient, key: str, model: str, service_tier: str | None) -> ChatResponse:
|
||||
chat = unwrap(
|
||||
proxy.chat(
|
||||
key,
|
||||
ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=f"{unique_marker()} Reply with the single word ok.")],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
service_tier=service_tier,
|
||||
),
|
||||
)
|
||||
)
|
||||
assert chat.id, f"sail chat with service_tier={service_tier!r} carried no id: {chat}"
|
||||
return chat
|
||||
|
||||
|
||||
class TestSailCompletionWindowPricing:
|
||||
@pytest.mark.covers("quota_management.spend_tracking.completion_window.bills_window_rates")
|
||||
def test_each_service_tier_bills_its_windows_cost_map_rates(
|
||||
self, client: SpendClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
priced = cheapest_window_priced_model(client.proxy.model_cost_map())
|
||||
model = register_priced_model(
|
||||
client.proxy, resources, "sail-windows", LiteLLMParamsBody(model=priced.model, api_key=SAIL_API_KEY)
|
||||
)
|
||||
|
||||
flex = _sail_chat(client.proxy, scoped_key, model, "flex")
|
||||
balanced = _sail_chat(client.proxy, scoped_key, model, "balanced")
|
||||
asap = _sail_chat(client.proxy, scoped_key, model, None)
|
||||
assert (flex.service_tier, balanced.service_tier, asap.service_tier) == (None, None, None), (
|
||||
"sail has no service_tier, yet the responses carried "
|
||||
f"{(flex.service_tier, balanced.service_tier, asap.service_tier)}"
|
||||
)
|
||||
|
||||
flex_row = _landed_row(client.proxy, flex.id or "")
|
||||
balanced_row = _landed_row(client.proxy, balanced.id or "")
|
||||
asap_row = _landed_row(client.proxy, asap.id or "")
|
||||
_assert_billed_at(flex_row, priced.flex, _chat_usage(flex), tier="flex", window="flex")
|
||||
_assert_billed_at(balanced_row, priced.balanced, _chat_usage(balanced), tier="balanced", window="balanced")
|
||||
_assert_billed_at(asap_row, priced.asap, _chat_usage(asap), tier=None, window="asap")
|
||||
totals = (flex_row.spend, balanced_row.spend, asap_row.spend)
|
||||
assert len(set(totals)) == 3, f"the three windows did not bill three different totals: {totals}"
|
||||
|
||||
@pytest.mark.covers("quota_management.spend_tracking.completion_window.bills_window_rates")
|
||||
def test_responses_completion_window_in_extra_body_bills_balanced_rates(
|
||||
self, client: SpendClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
priced = cheapest_window_priced_model(client.proxy.model_cost_map())
|
||||
model = register_priced_model(
|
||||
client.proxy, resources, "sail-responses", LiteLLMParamsBody(model=priced.model, api_key=SAIL_API_KEY)
|
||||
)
|
||||
window = "balanced"
|
||||
|
||||
response = unwrap(
|
||||
client.proxy.transport.post(
|
||||
"/v1/responses",
|
||||
headers=client.proxy.transport.bearer(scoped_key),
|
||||
json=_WindowedResponsesBody(
|
||||
model=model,
|
||||
input=f"{unique_marker()} Reply with the single word ok.",
|
||||
max_output_tokens=MAX_COMPLETION_TOKENS,
|
||||
extra_body=_WindowExtraBody(metadata=_WindowMetadata(completion_window=window)),
|
||||
),
|
||||
response_type=_ResponsesObject,
|
||||
)
|
||||
)
|
||||
row = _landed_row(client.proxy, response.id)
|
||||
usage = _ChatUsage(
|
||||
prompt_tokens=response.usage.input_tokens,
|
||||
cached_tokens=response.usage.input_tokens_details.cached_tokens,
|
||||
completion_tokens=response.usage.output_tokens,
|
||||
reasoning_tokens=response.usage.output_tokens_details.reasoning_tokens,
|
||||
)
|
||||
_assert_billed_at(row, priced.balanced, usage, tier=window, window=window)
|
||||
|
||||
@pytest.mark.covers("quota_management.spend_tracking.completion_window.rejected_window_bills_nothing")
|
||||
def test_flex_only_model_refuses_default_tier_and_bills_nothing(
|
||||
self, client: SpendClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
flex_only = cheapest_flex_only_model(client.proxy.model_cost_map())
|
||||
model = register_priced_model(
|
||||
client.proxy, resources, "sail-flexonly", LiteLLMParamsBody(model=flex_only.model, api_key=SAIL_API_KEY)
|
||||
)
|
||||
|
||||
outcome = client.proxy.transport.send(
|
||||
"/chat/completions",
|
||||
headers=client.proxy.transport.bearer(scoped_key),
|
||||
json=ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=f"{unique_marker()} Reply with the single word ok.")],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
service_tier="default",
|
||||
),
|
||||
)
|
||||
assert outcome.status_code == 400, f"expected the provider's 400, got {outcome.status_code}: {outcome.body}"
|
||||
assert outcome.response_cost == 0, f"a refused request was priced at {outcome.response_cost}"
|
||||
error = _ErrorBody.model_validate_json(outcome.body).error
|
||||
relayed = re.fullmatch(
|
||||
r"litellm\.BadRequestError: SailException - (?P<upstream>.+)\n\n"
|
||||
rf"LiteLLM: model group '{re.escape(model)}' failed with the error above\. No fallback was attempted\.",
|
||||
error.message,
|
||||
re.DOTALL,
|
||||
)
|
||||
assert relayed is not None, f"the error is not the provider's message relayed for {model}: {error.message!r}"
|
||||
assert (error.type, error.param, error.code) == (
|
||||
"invalid_request_error",
|
||||
"metadata.completion_window",
|
||||
"400",
|
||||
), (
|
||||
f"the provider refused {error.param!r}, not the completion_window the gateway rewrote service_tier "
|
||||
f"into: {error}"
|
||||
)
|
||||
assert re.search(r'"asap"', relayed["upstream"]), (
|
||||
f"service_tier=default must reach sail as completion_window asap; the provider saw: {relayed['upstream']}"
|
||||
)
|
||||
|
||||
rows = _poll_key_rows(client.proxy, scoped_key)
|
||||
assert [_row_shape(row) for row in rows] == [(0.0, "failure", False)], (
|
||||
f"the refused request left billable spend rows under the key: {rows}"
|
||||
)
|
||||
assert client.proxy.key_info(scoped_key).spend == 0.0, "the refused request added to the key's spend"
|
||||
|
||||
|
||||
class TestOpenAIKeepsItsOwnServiceTier:
|
||||
@pytest.mark.covers("quota_management.spend_tracking.service_tier.bills_served_tier")
|
||||
def test_openai_bills_the_tier_it_reports_and_refuses_balanced_unbilled(
|
||||
self, client: SpendClient, resources: ResourceManager, scoped_key: str
|
||||
) -> None:
|
||||
cost_map = client.proxy.model_cost_map()
|
||||
entry = cost_map[OPENAI_BACKEND] if OPENAI_BACKEND in cost_map else cost_map[OPENAI_BACKEND.split("/", 1)[1]]
|
||||
assert (
|
||||
entry.input_cost_per_token is not None
|
||||
and entry.output_cost_per_token is not None
|
||||
and entry.cache_read_input_token_cost is not None
|
||||
), f"{OPENAI_BACKEND} has no base rates in the cost map: {entry}"
|
||||
base = TierRates(
|
||||
input=entry.input_cost_per_token,
|
||||
output=entry.output_cost_per_token,
|
||||
cache_read=entry.cache_read_input_token_cost,
|
||||
)
|
||||
model = register_priced_model(
|
||||
client.proxy, resources, "openai-tiers", LiteLLMParamsBody(model=OPENAI_BACKEND, api_key=OPENAI_API_KEY)
|
||||
)
|
||||
prompt = f"{unique_marker()} Reply with the single word ok."
|
||||
|
||||
refused = client.proxy.transport.send(
|
||||
"/chat/completions",
|
||||
headers=client.proxy.transport.bearer(scoped_key),
|
||||
json=ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=prompt)],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
service_tier="balanced",
|
||||
),
|
||||
)
|
||||
assert refused.status_code == 400, f"expected OpenAI's 400, got {refused.status_code}: {refused.body}"
|
||||
assert refused.response_cost == 0, f"a refused request was priced at {refused.response_cost}"
|
||||
refused_error = _ErrorBody.model_validate_json(refused.body).error
|
||||
assert (refused_error.type, refused_error.param, refused_error.code) == (
|
||||
"invalid_request_error",
|
||||
"service_tier",
|
||||
"400",
|
||||
), f"OpenAI must receive service_tier itself, untouched by the sail rewrite: {refused_error}"
|
||||
|
||||
served = unwrap(
|
||||
client.proxy.chat(
|
||||
scoped_key,
|
||||
ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=prompt)],
|
||||
max_completion_tokens=MAX_COMPLETION_TOKENS,
|
||||
),
|
||||
)
|
||||
)
|
||||
assert served.id and served.service_tier, f"OpenAI reported no service_tier on the untiered call: {served}"
|
||||
row = _landed_row(client.proxy, served.id)
|
||||
_assert_billed_at(row, base, _chat_usage(served), tier=served.service_tier, window="openai default")
|
||||
|
||||
rows = _poll_key_rows(client.proxy, scoped_key)
|
||||
assert sorted((key_row.request_id == served.id, *_row_shape(key_row)) for key_row in rows) == [
|
||||
(False, 0.0, "failure", False),
|
||||
(True, row.spend, None, True),
|
||||
], f"expected one unbilled failure row and one billed row for {served.id} under the key: {rows}"
|
||||
Loading…
Add table
Reference in a new issue