mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-15 23:31:29 +00:00
Adds auto-router as a third savings driver alongside compression and prompt caching: a summary card, a donut segment, and a series in the savings graph. Savings are the net dollars from routing a request to the selected model instead of a counterfactual baseline. The delta prices prompt tokens at each model's input rate and completion tokens at each model's output rate, then subtracts the cache-write penalty the selected deployment incurs on a cold cache. Cache-read discounts stay attributed to the prompt-caching driver to avoid double-counting. The result is floored at zero so an escalation to a pricier model never reads as negative savings. The baseline model defaults to claude-opus-5 and is operator-configurable per deployment via the auto_router_savings_baseline_model litellm_param. It flows AutoRouter to PreRoutingHookResponse to the metadata bucket to SpendLogsMetadata to the daily spend writer, mirroring the routing_decision path, and is stripped from untrusted caller metadata so it cannot be spoofed. Savings accrue into a new autorouter_savings_spend column on the six LiteLLM_Daily*Spend rollup tables; no LiteLLM_SpendLogs queries are added. |
||
|---|---|---|
| .. | ||
| litellm-dashboard | ||
| Dockerfile | ||
| nginx.conf | ||