Commit graph

4669 commits

Author SHA1 Message Date
Yassin Kortam
33e9f54dc8
fix(proxy): reserve the larger declared output budget for TPM limits (#37001)
The TPM pre-call reservation read `max_tokens or max_completion_tokens`, so a
request declaring both was charged for whichever field came first. A caller
sending `max_tokens=1` with `max_completion_tokens=10000` reserved 2 tokens and
was then free to consume ten thousand, since the provider honours the modern
field and litellm's own param mapping drops the legacy one for the gpt-5 and
o-series families.

Reserve against the larger of the declared budgets instead. Over-reserving is
the safe direction for a limiter: post-call reconciliation refunds the
difference between the reservation and actual usage, while under-reserving lets
the window be exceeded before anything notices.
2026-08-15 11:54:22 -07:00
devin-ai-integration[bot]
fe9451c6cd
fix(panw_prisma_airs): surface scan_id on allowed requests (#37037)
* fix(panw_prisma_airs): surface scan_id and scan metadata on allowed requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: ruff format panw guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(panw_prisma_airs): expose scan id header only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(panw_prisma_airs): inject http client instead of patching private api

Adds an http_client seam so the scan-id tests drive the real AIRS request/parse path through a mock transport, plus direct coverage for the scan-id header helper.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): expose guardrail scan id header to browser clients

Keeps the panw optional_fields block untouched to avoid a needless conflict with a sibling PR that deletes it.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-15 11:49:03 -07:00
Yassin Kortam
5a50fe0b46
feat(proxy): gate the Global Control Plane worker registry on an enterprise license (#36996)
The Global Control Plane (formerly documented as the HA Control Plane) is
documented as an Enterprise feature, but `worker_registry` carried no premium
check, so any OSS install could run one. Gate it at config load, matching the
`enforced_params` precedent, and fail startup rather than ignoring the registry
silently: a silently dropped registry degrades a control plane into an ordinary
proxy with no signal to the operator.

Also declare `worker_registry` and `general_settings.control_plane_url`, both
load bearing today and neither previously declared, so they appear in the
generated config schema.
2026-08-15 11:40:04 -07:00
Thijmen Stavenuiter
8c5fa0c0f9 feat: add show budget window usage 2026-08-15 20:06:21 +02:00
mateo
1e63134adb fix(slack_alerting): hold a pod lock so a fleet sends one deprecation alert per day
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-15 17:10:12 +00:00
mateo
816fa50394 refactor(proxy): make the deprecation loop entrypoint public and drop a dead None check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-15 08:35:12 +00:00
Mateo Wang
89f233a15d
Merge pull request #36720 from BerriAI/litellm_tiered_pricing_cache_creation
fix(cost): tiered pricing supports cache creation cost and is all-or-nothing
2026-08-14 21:51:54 -07:00
devin-ai-integration[bot]
6e7984e537
fix(proxy): requeue spend logs when the DB write fails with a transport error (#36716)
* fix(proxy): requeue spend logs when the DB write fails with a transport error

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): hardcode the spend log queue cap and drop the stale re-export

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): keep the spend log requeue within the type discipline budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply the spend log queue cap to producer appends too

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): lower the spend log queue cap to 1k and make it env configurable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): bound the spend log queue by bytes instead of row count

A row cap cannot bound memory: a row carries the whole prompt under store_prompts_in_spend_logs, so a cap that rides out an outage of counter-only rows is an OOM once prompts are stored. Every enqueue and dequeue now goes through one pair that tracks what the queue costs and drops the oldest rows past a 64 MB budget.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): make the spend log queue byte budget env configurable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): use a string default for the spend log queue byte budget env read

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): make the spend log queue byte total a public attribute

The queue it accounts for is already public, and a private name only bought reportPrivateUsage errors at every call site.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: shivam <shivam@berri.ai>
2026-08-14 20:49:45 -07:00
Mateo Wang
6c2dcb801b
Merge pull request #36907 from guptaishaan/fix/issue-36880-8933
fix: report real token usage on guardrail-blocked /v1/responses replies
2026-08-14 20:16:25 -07:00
Mateo Wang
cba2beaf42
Merge pull request #36965 from erensh27/feat/per-component-cost-headers
feat(proxy): per-component response cost headers
2026-08-14 18:15:23 -07:00
mateo
3f64cbe41b fix(ptu): empty a PTU deployment's tiered_pricing instead of dropping it
Dropping it falls back to the public cost map's tier table, whose rates outrank the
zeros written beside them, so a PTU deployment on a tiered model keeps billing its
traffic per token. Stored empty, the tiers no longer apply and the zeros win
2026-08-15 01:14:14 +00:00
mateo-berri
5970754a85 fix(ptu): clear a PTU deployment's tiered_pricing instead of zeroing it
tiered_pricing is a list, so the 0.0 the flat-rate zeroing stores does not
even validate. Supplying tiers alongside PTU config gets the same 400 as a
flat rate; tiers already stored are dropped from both blobs
2026-08-14 18:04:01 -07:00
mateo-berri
05beb7abb5 fix(proxy): emit uncached input cost so component headers sum to the total 2026-08-14 17:30:50 -07:00
mateo-berri
7562445273 test(proxy): assert production nesting semantics for component cost headers 2026-08-14 17:24:10 -07:00
Mateo Wang
d74cb6de1b
Merge pull request #36798 from BerriAI/litellm_azure_ai_docs_index_write_grant_rc
fix(azure_ai): recognize real Search doc endpoints so teams can read/write via passthrough
2026-08-14 17:24:06 -07:00
mateo-berri
3217b8edae fix(proxy): count worker heartbeats on the primary so replica lag cannot undercount 2026-08-14 17:23:06 -07:00
yucheng-berri
f9f5c03884
fix(mcp): drop caller host and configured upstream headers from logged metadata (#36901)
* fix(mcp): drop caller host and configured upstream headers from logged metadata

The synthetic request that carries MCP client headers into
add_litellm_data_to_request forwarded the caller's Host header, and
Request.url is built from it, so a caller chose the proxy_server_request
url and the metadata endpoint that every logging callback records.

_upstream_credential_headers also only knew the configured client side
auth header and the x-mcp- prefix family, so a header name declared in
mcp_servers.<name>.extra_headers reached logging metadata in cleartext.
Those names are admin chosen, so no prefix rule can recognize them; read
them off the server registry instead. The header is still forwarded
upstream, which is what extra_headers is for. authorization is left out
because clean_headers already strips it and claiming it here would move
authenticated_with_header on the oauth passthrough config.

The Responses bridge tests stub the server manager, so their fakes gain
the registry accessor the sanitizer now reads.

* fix(mcp): drop caller host from the sanitized header mapping too

The synthetic request stopped forwarding host, but the parallel sanitizer
did not, so a forged hostname still reached the guardrail payload and the
list_tools spend row. Drop it there as well.

Exempt the configured identity headers from the upstream credential set.
get_user_from_headers resolves end user attribution off the same request
this module reconstructs, and it only fills end_user_id when auth left it
unset, so claiming user_header_name or a user_header_mappings name would
lose attribution on the MCP paths that authenticate upstream.

Drop the isinstance guard on extra_headers entries: the field is typed
list[str], so the check is dead and basedpyright scores it.

* fix(mcp): accept a bare user_header_mappings entry when exempting identity headers

get_internal_user_header_from_mapping and get_customer_user_header_from_mapping
both normalize a single mapping to a one element list, and config_settings.md
documents the key as a dict. Iterating the bare form yields its keys instead,
so the exemption silently matched nothing and an identity header also named in
an MCP server's extra_headers was dropped after all.
2026-08-14 17:21:07 -07:00
mateo-berri
ed33687422 feat(proxy): auto-suppress the no-Redis banner for confirmed single-worker deployments 2026-08-14 17:11:05 -07:00
Shivam Rawat
4110812b52
Merge pull request #33881 from BerriAI/litellm_fix_redis_spend_buffer_requeue_33872
fix(proxy): requeue Redis spend buffer transactions when the DB commit fails
2026-08-14 17:08:05 -07:00
mateo-berri
e06e2e6c12 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_36907_head 2026-08-14 17:06:18 -07:00
tin-berri
2d3c3e3098
feat(shadow_eval): add reverse-direction shadow eval jobs (#36865)
Shadow eval only answered "should this key adopt this auto-router". Once a key
is on the router it is invisible to the feature, because the sampling gate skips
any request the shadowed router already served, so post-adoption quality
regressions go unmeasured.

Reverse mode inverts the arms: sample the traffic the router did serve and
duplicate it against a fixed baseline_model, judged by the same blind pairwise
judge. Same job table, same attempt rows, same aggregates.

real_* stays the arm the caller was served and shadow_* the duplicated one, so
in reverse real_model is the router's pick and shadow_model is the baseline. The
active-job slot becomes one per (key, direction) so both directions can run at
once, and tier attribution in reverse reads the control request's routing
decision rather than the shadow call's write-back.
2026-08-14 17:05:55 -07:00
tin-berri
d4d6bc2577
fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting (#34845)
The MCP sub-app is attached with app.mount("/mcp", ...) and a Starlette
mount never matches its bare prefix, so POST /mcp fell through to the
router's redirect_slashes 307. Behind a TLS-terminating ingress whose
peer address is not in uvicorn's forwarded-allow-ips (default: loopback
only) the redirect Location is built from the socket scheme as http://,
and MCP clients strip the Authorization header on the cross-origin
follow, so reconnects fail with ECONNRESET right after a successful
OAuth flow. The redirect also fires before auth, so the bare spelling
never returns the RFC 9728 WWW-Authenticate challenge that OAuth
clients need to start the flow.

Add an explicit /mcp route beside the existing /toolset/{name}/mcp and
/{name}/mcp spellings, forwarding to handle_streamable_http_mcp with
the same scope rewrite those routes already use (path=/mcp,
_original_path preserved for OAuth challenge URL selection). When the
mcp package is unavailable the route 404s, matching what the bare
sub-app serves on /mcp/ in that state. /mcp/, /mcp/{server},
/{server}/mcp and /toolset/{name}/mcp spellings are unchanged; the
exact-match route and the mount have disjoint match sets so
registration order cannot matter.
2026-08-14 17:04:32 -07:00
mateo-berri
e4f2ea12bc fix(responses_api): map bridged chat usage on guardrail-blocked replies
Move the blocked-usage mapping for /v1/responses next to
blocked_response_usage in guardrail_translation utils, map bridged chat
prompt/completion tokens to Responses API input/output tokens, and let
raise_passthrough_exception attach the blocked response so post-call
guardrail blocks report real usage
2026-08-14 17:04:27 -07:00
mateo-berri
9079e4c47b fix(proxy): return cost breakdown header values as a named tuple 2026-08-14 17:04:26 -07:00
mateo-berri
b14c4a8d45 fix(vector_stores): classify write endpoints before reads on substring collisions 2026-08-14 16:53:48 -07:00
Yassin Kortam
eb4b847268
fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown (#36961)
Anthropic's Models API declares max_input_tokens and max_tokens as nullable, not
optional, and the live vendor endpoint returns both keys on every entry. The
merged Anthropic-native listing dropped either key whenever LiteLLM could not
resolve a limit, so a client validating against a nullable-but-required schema
saw a malformed entry for any model the cost map does not know.
2026-08-14 16:52:39 -07:00
Mateo Wang
870a8cf764
Merge pull request #36974 from BerriAI/litellm_vllm_dropdown_labels
fix(ui): distinguish hosted and local vLLM in the provider dropdown
2026-08-14 16:28:57 -07:00
devin-ai-integration[bot]
40d999b693
fix(mcp): keep admin-entered oauth endpoints in management reads (#36888)
* fix(mcp): keep admin-entered oauth endpoints in management reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover configured oauth endpoints on the config load path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-14 16:27:37 -07:00
Yuneng Jiang
9592a5447f
Revert "fix(auth): stop the team fallback from widening model access (#36837)"
This reverts commit ab2333b6c4.

Every Admin UI login mints its session key against the sentinel team_id
`litellm-dashboard`, and no LiteLLM_TeamTable row is ever created for it.
That lookup is therefore a provably-absent row on every UI request, which
#36837 turned into a hard refusal with no override, so the whole dashboard
404s.

Reverting restores the token-derived fallback. The model-access widening
#36837 closed is reopened and needs a re-land that exempts the UI sentinel
team.
2026-08-14 16:03:51 -07:00
devin-ai-integration[bot]
865ed96765
fix(proxy): force prisma recreate on postgres cached-plan error (#36428)
`_query_first_with_cached_plan_fallback` recovers from Postgres's "cached
plan must not change result type" by recreating the Prisma client, which
drops both the server-side plans and the engine's client-side statement-name
cache. Since #30183 the shared reconnect path probes the writer with
`SELECT 1` first and skips the recreate when it answers, which is right for
the IAM token refresh it was added for and wrong here: the connection is
healthy, it is the session's prepared statements that are stale, so the probe
always passes and always vetoes the recreate. Callers now pass
`force_recreate` to skip that probe, and only the cached-plan fallback does.

Getting past the probe is not enough on its own. Both cooldown checks would
still skip the recreate for 15 seconds after any earlier reconnect, which
outlives the 10 second auth retry window, so a migration landing in that
window kept 503ing. `force=True` would fix that but would also let every
concurrent caller of the same burst kill the engine the first one just built.
The caller instead names the engine it observed before the query, and the
cooldown is waived only while that engine is still the live one, so the first
caller repairs the pool and the rest fall back to the normal cooldown.

That engine has to be the one the query actually ran on. `query_first` is a
top-level read, so with a read replica configured it is dispatched to the
reader and it is the reader's prepared statements that go stale, while
`writer_db` names a different engine with its own counter. The observation
and the cooldown comparison both go through `read_db`, added alongside
`writer_db` and backed by a `read_target` property on the routing wrapper
that `__getattr__` now dispatches through so the two cannot drift.

The observation carries the wrapper, not just its generation. `read_db`
resolves to the reader while it is available and to the writer once it is
not, and those counters are independent and both start at zero, so comparing
a bare number across that switch pits one engine's counter against another's.
Equal by coincidence waives the cooldown for an engine already replaced;
unequal gates a caller that needs the recreate. Identity settles it, and is
sound because the engine object is never re-pointed without the generation
also moving.

Three smaller holes on the way out. The waiver is withdrawn once a repair of
that same engine has been tried and failed, so a burst collapses onto one
attempt instead of each caller running its own recreate serially; the record
is keyed per engine rather than counted globally, so an unrelated reconnect
failure cannot suppress a stale reader's recovery and a writer failure cannot
evict the reader's record. And a forced recreate that the optimistic-lock
guard declines is no longer reported as a success on either the direct or the
heavy path, since the routing wrapper leaves the reader untouched in that
case; a decline is deliberately not counted as a failure, so the caller's own
backoff still gets its waiver on the next attempt.

A decline on the heavy path clears the dead-engine flag before raising. The
clear after the cycle is skipped by any raise, which is right for a failure
and wrong here, and the non-forced path already clears it on a decline, so
this restores that policy rather than inventing one. Stranding the flag would
route the next cycle back down the probe-free heavy branch, where the
refreshed generation matches and the recreate kills the healthy engine a
refresh just spawned, which is #29176.

Clearing that flag is necessary and not sufficient. The escalation check
re-arms it whenever the consecutive-failure count sits at the threshold, so a
decline that left the count alone sent the very next attempt back down the
same path. A decline is raised only at the generation guard, and the
generation moves only after a replacement connects, so a decline is proof
that a replacement succeeded and the count is reset on it.

Fixes #36418

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-14 15:04:01 -07:00
Scott Wilson
9858d021ee fix(guardrails): record MCP tool guardrail evaluations and blocks in usage monitor
MCP tool calls run their guardrails against a throwaway LLM-shaped dict
built by `ProxyLogging._convert_mcp_to_llm_format`, not against the dict
the tool call is logged from. `@log_guardrail_information` therefore
appended `standard_logging_guardrail_information` to that throwaway
dict's metadata bucket, where `get_standard_logging_object_payload`
never saw it, so the Guardrails Monitor reported zero evaluations and
zero blocks for all MCP traffic.

Thread the request's `litellm_logging_obj` into `pre_call_tool_check`
and `_create_during_hook_task` and bridge the guardrail records onto it:

- Seed `data["litellm_logging_obj"]`, which unified guardrails read and
  pass into `apply_guardrail`.
- Call `_sync_guardrail_info_to_logging_obj` in a `finally`, which is
  what native guardrails need and what makes the block path work: a
  blocked call raises straight out of `pre_call_tool_check`, so the
  record has to be attached before the exception leaves the frame.

Only the guardrail evaluation records are copied. The synthetic
request's messages and tool arguments are deliberately left behind --
they can carry end-user data and nothing in the monitor needs them.

In `call_mcp_tool`, flush the failure handlers before
`post_call_failure_hook` so the `status="failure"` standard logging
object exists when `_ProxyDBLogger.async_post_call_failure_hook` writes
the spend-log row the monitor's "Total Blocked" counts. Both handlers
gate on `should_run_logging("sync_failure")` / `("async_failure")` and
then mark it, so the `@client` wrapper's own post-raise logging is a
no-op and nothing is double-counted -- the same pattern
`_fire_mcp_tool_call_logging` already uses for `isError=True`.

Threaded through every MCP tool entry point: the managed-server path,
the local-OpenAPI registry path, the legacy registry fallback, and the
Responses API's `_execute_tool_calls`.
2026-08-14 17:52:36 -04:00
mateo-berri
08966c842b test(vector_stores): drop redundant route-map comment 2026-08-14 14:12:09 -07:00
mateo-berri
0a81e1b222 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_azure_ai_docs_index_write_grant_rc 2026-08-14 13:58:28 -07:00
mateo-berri
ae2a5e1472 fix(ui): distinguish hosted and local vLLM in the provider dropdown 2026-08-14 13:39:37 -07:00
shivam
3838969527 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_redis_spend_buffer_requeue_33872 2026-08-14 19:07:19 +00:00
abhinav
1241bd5ce1 feat(proxy): add per-component response cost headers
- Extract input_cost, output_cost, cache_read_cost, cache_creation_cost, reasoning_cost, and tool_usage_cost from logging object cost breakdown
- Populate x-litellm-response-cost-* component headers in ProxyBaseLLMRequestProcessing.get_custom_headers
- Ensure headers are omitted when cost breakdown is absent or values are None
- Add comprehensive test suite covering component headers, math invariants, caching, reasoning, and discounts/margins
2026-08-14 23:43:04 +05:30
Shivi Jain
2b23295f82 fix(proxy): reconcile project quota reservations 2026-08-14 23:31:18 +05:30
Armaan Sandhu
e1f3d6e158
feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery (#35455)
* feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery

* refactor(proxy): move Anthropic model-list formatter into llms/anthropic/common_utils

* fix(proxy): make model_list request param optional for direct callers

* style: apply ruff format to changed lines

* style: satisfy ruff strict-rule budget (UP006, I001)

* style: satisfy type-discipline budget (LIT002 mutable-ok, LIT009 pyright ignore)

* style: satisfy LIT001/LIT010 and drop explanatory comment per contributor rules

* fix(proxy): translate team model names in the Anthropic /v1/models response

* ci: trigger buildkite status report

* feat(proxy): carry token limits into the Anthropic-native /v1/models entries

* fix(proxy): cast the injected request so the anthropic-version guard is a real comparison

* fix(proxy): explain the model listing casts so the type-discipline gate passes

---------

Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-08-14 10:22:58 -07:00
Shivi Jain
312d12fe0c fix(proxy): enforce project ITPM/OTPM quota on every Responses WebSocket frame
The connection-level pre-call hook only ran once per WebSocket
connection, so a project caller could send unlimited high-token
response.create frames after a single minimal reservation. Adds
enforce_project_io_token_quota_for_frame to the v3 rate limiter and
wires it into both the native and managed WebSocket handlers via a
duck-typed litellm.callbacks lookup, so the SDK layer stays free of
proxy imports. A rejected frame gets an error event; the connection
stays open for the client to retry.

Also fixes the RET504 and BLE001 strict-lint-budget violations the
litellm_internal_staging merge introduced in
parallel_request_limiter_v3.py, which were failing the lint check.
2026-08-14 21:39:34 +05:30
Shivi Jain
a9227057a1 Resolving merge conflicts and verai comment for batch 2026-08-14 20:52:05 +05:30
Ishaan
81c27fc4a0 fix: report real token usage on guardrail-blocked /v1/responses replies
## TLDR

Signed-off-by: Ishaan <ishaangupta0408@gmail.com>
2026-08-14 09:15:33 +00:00
Marty Sullivan
c9e9c279fe fix(batches): decide batch cost ownership once per retrieve
The ownership question was asked twice for one retrieve: once before the provider
call to decide whether to suppress inline accounting, and again afterwards to
decide whether to mark the batch accounted. Between those two points the poller
can complete its first successful filtered query and become usable, so the two
answers disagree. The retrieve then accounts for the batch inline, having decided
the poller was unusable, while the later check sees a usable poller and leaves the
marker unset, so the poller accounts for the same batch again and its spend is
counted twice.

The retrieve now decides once and passes that decision to
update_batch_in_database, which prefers it over re-deriving one. Callers that
record no cost of their own leave it unset and keep deriving it as before, so the
cancel path is unchanged.
2026-08-14 02:18:40 -04:00
Marty Sullivan
ec52858865 fix(batches): only hand accounting to the poller once it can mark batches done
The handoff asked whether the poller was running, when what matters is whether it
will actually account for the batch. Those differ on a schema without the
batch_processed column: the poller cannot filter on it, so it falls back to a
query that excludes complete and completed rows, and it cannot set it either. A
caller retrieving a provider-completed batch before the poller saw it therefore
suppressed inline accounting, then marked the row complete, and the fallback query
could never find it again. Nobody accounted for that batch, so its cost escaped
the caller's budget entirely.

The poller now publishes batch_processed_support_confirmed, set only once a
filtered query has actually succeeded, and the handoff requires it. Defaulting to
unconfirmed keeps accounting on the retrieve path in exactly the cases the poller
would drop the batch, including the window before the poller's first cycle. All
four combinations account exactly once: unconfirmed leaves the retrieve
accounting and setting the marker, whether or not the column exists, and
confirmed is only reachable when the column is present, where the poller accounts
and sets it.

A scheduler that hands back something other than a bound method leaves no poller
to interrogate, which reads as unconfirmed rather than as working.
2026-08-14 01:46:56 -04:00
Marty Sullivan
5649098e1b fix(batches): account a managed batch's cost exactly once
Two components computed a managed batch's cost and each assumed it was the only
one. Retrieving a batch computed it through the @client decorator's success
callback, and CheckBatchCost computed it on its own schedule. Whichever observed
completion first decided the outcome, so cost was either counted once per
retrieve or not at all.

The lockout is the worse half. Retrieving a batch that had reached completion set
batch_processed=True, which is what takes a batch out of CheckBatchCost's queue,
since it selects batch_processed=False. That write claimed the cost had been
accounted for on behalf of a callback that had not run yet and was not awaited.
When the callback then failed the cost was gone permanently, with the poller
already retired and no retry left. Observed on a live proxy: two completed
batches whose callbacks raised inside the logging worker, one on a provider
output path that did not resolve and one on a batch whose output file id was
still None, both left marked processed with no spend row and no way to recover
them. Nothing logged at error level for the batches themselves.

The over-count is the other half. Nothing suppressed recomputation, so each
retrieve of an already-completed batch recorded that batch's full cost again. A
caller polling its own batch to see whether it had finished inflated spend by
however many times it looked.

The flag now means what its name says, and only the component that actually
recorded the cost sets it. When the poller is running it owns accounting, so
retrieving a managed batch records no cost and leaves the flag alone; the poller
computes once and sets it. When the poller cannot be relied on, either because
polling is disabled by config or because the enterprise job never registered,
the retrieve path is the only accountant and behaves exactly as before. Batches
with no managed object row are untouched either way, since neither the flag nor
the poller queue applies to them.
2026-08-14 01:36:52 -04:00
Marty Sullivan
d7afc1797c refactor(batches): share the trusted-credentials helper across both call paths
The helper that carries the credential snapshot into litellm_params lived private
in files/main.py, and the batch retrieve needed it too. It now sits beside
get_litellm_params, which is what it augments, so neither caller reaches into the
other's private surface. Typed as Mapping/MutableMapping of object rather than
Any, which the strict import rules ban.

The file-content route builds the snapshot through the same helper as the batch
route instead of assembling a conditional mapping inline, which drops two mutable
constructions and leaves one way to attach it. Its name loses the batch suffix now
that both routes use it.
2026-08-14 01:32:59 -04:00
Marty Sullivan
60fe4e464c fix(bedrock): resolve the managed-batch output bucket on the inline accounting path too
A third path reads a completed batch's output file, and it could not resolve the
bucket either. When cost is accounted from the retrieve itself rather than from
the poller, the batch success handler calls _handle_completed_batch, which fetches
the output file through _extract_file_access_credentials. That helper forwarded a
whitelist covering Azure and Vertex, gcs_bucket_name included, but nothing for
Bedrock, and retrieve_batch built its litellm_params through get_litellm_params,
whose fixed signature drops the trusted credential snapshot. So the snapshot never
reached the file read and it failed with "S3 bucket_name is required" for a bucket
the deployment had configured, leaving the batch's cost unrecorded.

Adding s3_bucket_name to that whitelist would not have worked. The Bedrock file
config deliberately resolves the bucket only from the immutable server-side
snapshot or the environment, never from a request param, because the bucket is
what managed file ids are validated against. The snapshot is therefore what has to
flow, exactly as it already does for the model-routed and cost-poller paths.

retrieve_batch now re-adds the snapshot after get_litellm_params, the same way the
file operations already do, the whitelist forwards it, and the proxy attaches it
for router-routed managed batches from the deployment behind the unified id.
Verified against a live proxy reading a real completed Bedrock batch: the cost row
appears within seconds of the retrieve carrying the batch's real spend and usage,
where before the read raised and no row was written.

Resolving those credentials is best effort. A batch whose deployment no longer
resolves, which happens when a model group is removed while batches are in
flight, still serves its status instead of failing the request on the lookup.
This matters for the OSS and polling-disabled configurations, where the retrieve
path is the only thing that accounts for a batch at all.
2026-08-14 01:17:56 -04:00
Marty Sullivan
460f0d29a9 test(files): capture routed retrieval calls immutably
The mock merged every call into one shared dict, so a second routed retrieval would
overwrite the first and the assertions would still pass. Keep one frozen snapshot per
call and assert exactly one call, which also makes an unintended second retrieval a
failure rather than something the merge hides
2026-08-14 01:17:56 -04:00
Marty Sullivan
c99a1ab0d7 fix(bedrock): resolve the managed-batch output bucket on the model-routed and cost-poller paths
get_configured_s3_bucket_name accepts the output bucket only from the immutable
_litellm_internal_model_credentials snapshot or AWS_S3_BUCKET_NAME. That refusal to read
litellm_params is deliberate: the bucket is what validate_managed_cloud_file_id checks a
file id against, so trusting a request-supplied value would let a caller redirect reads
to a bucket of their choosing

Two live entry points reach the Bedrock file-content transformation without ever building
that snapshot. The managed-files pre-call hook sets data["model"] for any id carrying
llm_output_file_id, which is every batch output, so get_file_content always takes the
model-routed branch; that branch called llm_router.afile_content directly, and
managed_files_obj.afile_content, the only caller that built the snapshot, is therefore
unreachable for batch output. CheckBatchCost spread the deployment credentials as plain
kwargs, and get_litellm_params does not carry s3_bucket_name across (gcs_bucket_name is
listed for exactly this reason, its S3 counterpart is not), so the poller lost the bucket
the same way

The result was that every completed Bedrock managed batch failed files.content with
"S3 bucket_name is required" and never had its cost tracked, leaving the row to be
re-polled every cycle. Both paths now resolve the deployment credentials and pass the
same MappingProxyType snapshot the managed-files hook already builds
2026-08-14 01:17:56 -04:00
Marty Sullivan
363e3f3f03 test(spend): annotate the batch cost row constants as Final 2026-08-14 01:16:48 -04:00
Marty Sullivan
9a9e7a58d3 fix(spend): give a batch's cost row a primary key of its own
request_id is the primary key of LiteLLM_SpendLogs and the flush inserts with
skip_duplicates, so a spend log whose id already exists is dropped with no error
raised and a "processed 1 spend log" line still logged. Batch cost accounting
produced exactly such an id twice over, and on a proxy with message redaction
enabled no batch cost row could be written at all.

get_spend_logs_id derived the id by md5-hashing the response for two call types,
aretrieve_batch and acreate_file. Redaction makes that hash a constant:
perform_redaction returns the fixed {"text": "redacted-by-litellm"} placeholder
for any shape it cannot redact, which is what a batch object and a file body both
become, so every such row hashed to md5('{"text": "redacted-by-litellm"}') =
00fcbef15a3b0097e14b0ca016ed30a0 regardless of provider, user, or amount. The
first row to claim that id owned it and every later row was discarded. Verified
against a live proxy: four payloads spanning two providers and three distinct
spend values all computed that id, and the table held one acreate_file row dating
to 2025-05-25, the row that had claimed it.

Keying off the batch's own identity instead is necessary but not sufficient,
because creating a batch already writes an acreate_batch row under exactly that
id, so the cost row becomes a duplicate of the batch's own creation row. Also
verified live: after the hash was removed the poller computed and flushed a
batch's cost, and the only row carrying that id was the acreate_batch row from
when the batch was submitted.

The id now comes from the response's own id, then the standard logging payload's
id, then litellm_call_id, and a batch cost row is namespaced with a _batch_cost
suffix so it cannot collide with the creation row. The middle term is what keeps
this correct under redaction: that payload is built from the unredacted response,
so it still carries the batch id after redaction has flattened the body. Keying
the cost row to the batch rather than to the call also keeps accounting the same
batch twice collapsing to one row instead of billing it twice. Every other call
type still derives its key exactly as before.

Cost and usage themselves are unaffected by redaction: the token columns fall back
to the standard logging payload and spend comes from its response_cost, neither of
which redaction touches. generate_hash_from_response had no other caller and is
removed with it.
2026-08-14 01:05:38 -04:00