litellm/tests/test_litellm/proxy/spend_tracking
tin-berri 22f68c0c6b
fix(spend): read what a request cost from the record instead of pricing it again (#35736)
The auto-router savings driver recomputes what the served request cost, but that
request is not a counterfactual: it ran, and the cost calculator already billed it and
wrote the number down. Recomputing means restating every pricing dimension the biller
applied, and the two this missed were enough to halve it. A request billed at a
priority tier is recomputed at standard rates, and a regional host's uplift is dropped
entirely, so the driver writes a savings figure into the same rollup row as the `spend`
it disagrees with. On `gpt-5.4-mini` at priority the row is billed 0.024 and the driver
prices the same usage at 0.012.

Neither omission cancels between the two arms, because both are per-model. The uplift
is a multiplier read off each model's own entry, so 1.1*A - 1.1*B is 1.1*(A-B) and a
model without one does not move at all. Tier coverage is sparser and asymmetric:
`gpt-5.6` has priority rates and `gpt-5.4-nano` has none.

`cost_breakdown` already carries the answer and already reaches the call site. The cost
calculator records it, it rides the standard logging payload into the spend log's
metadata, and OTEL, the log drawer and the response headers all read it rather than
re-deriving; this driver was the only downstream consumer in the tree still pricing a
completed request from its tokens. `input_cost` and `output_cost` sum to exactly what
the pricer returns, so the served arm reads them. Tool spend, discount and margin stay
out, since the counterfactual cannot be priced with them and charging them to one arm
alone would read as the router losing money on every tool call.

The baseline never ran, so it is still priced through the cost engine, now on the basis
the biller used. `CostBreakdown` carries that basis because it cannot be recovered
afterwards: the tier the biller used comes from `optional_params`, which no log record
keeps, and the served tier that does survive on the usage object is a different fact
with the opposite precedence. Rows written before this shipped carry no basis and price
at standard rates, exactly as they do today; there is no backfill.

Two smaller things in the same path. The router is passed as a provider rather than a
router, so a spend write that was never auto-routed no longer fetches and discards one,
and the complexity router resolves its messages once per hook instead of once per
consumer.
2026-08-03 20:46:09 -07:00
..
test_budget_reservation_redis_failure.py fix(register_model): preserve built-in cache pricing when registering custom overrides under unmapped keys (#30044) 2026-06-10 12:11:03 -07:00
test_cloudzero_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_compression_savings.py feat(spend): track prompt compression saved tokens in daily spend aggregates (#33810) 2026-07-18 17:47:54 -07:00
test_savings.py fix(spend): read what a request cost from the record instead of pricing it again (#35736) 2026-08-03 20:46:09 -07:00
test_spend_log_error_logger.py feat(spend-logs): opt-in suppression of stack traces in spend-tracking error logs 2026-05-02 00:44:34 +00:00
test_spend_management_endpoints.py feat(ui): show which log rows are the auto-router's own classifier calls (#35304) 2026-07-31 11:51:36 -07:00
test_spend_query_optimization.py fix(spend): bound the logs-tab pagination count to stop full-window scans (#31825) 2026-07-07 09:41:20 -07:00
test_spend_tracking_utils.py feat(spend-logs): record when a spend log row is the auto-router's own classifier call (#35300) 2026-07-30 19:26:42 -07:00