The pure helpers express the same regression: the save-side drop must
leave the alias in place, and two registries must expand it to their
own ids. The full validate path is already covered by the
persists-verbatim test and the live e2e test.
A single-instance run cannot reproduce the two-region setup, but the
regression is fully visible in one: the alias must survive to /key/info
unrewritten, and the alias-granted key must still list the server's
tools. The broken write path stored the resolved server id instead.
Since PR #29128, key create/update/regenerate resolved every
object_permission.mcp_servers entry against the saving instance's
DB + config registry and persisted the resolved server ids. For
config-loaded servers the id is derived from a hash of the regional
URL, so in a shared-database multi-region deployment the rewrite baked
one region's ids into the row and every other region denied the key.
Grants written before v1.88.0 kept the raw alias and kept working,
which is why only newly provisioned keys broke.
Keep the validation and the stale-entry drop (the LIT-3278 fix), but
persist the caller's original identifiers for everything that resolves.
Read-time expand_permission_list already maps a name to each region's
local server id.
* fix(proxy): include litellm_model_table in GET /v2/team/list
GET /v2/team/list built its find_many queries without joining the
LiteLLM_ModelTable relation, so litellm_model_table (and the
model_aliases it carries) always read back as null there, same bug
class as GH #26312 which PR #33047 fixed on /team/info and /team/list
but never touched this endpoint.
* fix(test): assert observable output, not mock calls, in v2 team list test
The test-quality gate flagged the regression test for asserting on
find_many's call args instead of what the caller gets back. Rewritten
so the fake find_many only attaches litellm_model_table when its own
include kwarg asks for it, so the assertions are on the response.
* fix(proxy): drop invalid litellm_model_table include on deleted-team query
Greptile caught that LiteLLM_DeletedTeamTable has no litellm_model_table
relation in the Prisma schema, so passing that include on the deleted-team
find_many raised UnknownRelationalFieldError against a real database on
every GET /v2/team/list?status=deleted call. Confirmed live against
Postgres. Scope the fix to the active-team branch only, where the relation
exists; update the test to reflect that and assert the deleted branch no
longer requests it.
PR #38113 made a configured search_tool_name fail fast when the router
does not carry a matching search tool, which broke
test_pre_request_hook_modifies_request_body: it names test-search-tool
but never registers it. Stub the proxy router with that tool so the test
exercises the conversion path again.
litellm_call_id is populated from the client-settable x-litellm-call-id
header, so a request_id lookup can match more than one row across
tenants. Authorizing on a single arbitrary match let an attacker reuse
a victim's request_id as their own call id and read the victim's spend
log row. Widen the ownership check to require every matching row to
belong to the caller, failing closed on any foreign match.
A non-admin could widen a read-only (info_routes) key to llm_api or full
access through the preset carve-out. Read-only keys now stay read-only
unless a proxy admin widens them; the other preset transitions, including
the LIT-4891 llm_api to full access switch, still work. Also converts the
transition tests to assert on a returned outcome so the no-403 cases
carry real assertions.
Streamed responses through the proxy previously exposed no usable cost:
the x-litellm-response-cost header is unreadable mid-stream and the final
usage chunk carried only tokens, priced against an alias model name the
client cannot resolve. The include_cost_in_streaming_usage flag existed
but was off by default and only fixed the wire, not SDK clients.
Stamp usage.cost into the joined streaming response by default wherever a
final usage object is built: the chat-completions stream_chunk_builder,
the native /v1/responses RESPONSE_COMPLETED event, and synthetic response
events. Provider-reported cost always wins over the computed value, and
only positive computed costs are stamped so unpriceable alias responses
keep deferring to the logging object's own calculation. Per-chunk SSE
cost injection (/v1/messages, generateContent, passthrough) stays behind
the flag.
Also normalize non-litellm usage objects in stream_chunk_builder: openai
CompletionUsage lacks Usage.__contains__, so membership probes silently
returned False and client-side rebuilds dropped the wire cost and
recounted token usage locally. Wire token counts and cost now survive.
Resolves LIT-6427
Success spend rows are keyed by the upstream provider response id, so the
x-litellm-call-id response header value never found them. Add a nullable
indexed litellm_call_id column to LiteLLM_SpendLogs, populate it at write
time, and widen every request_id lookup surface (/spend/logs,
/spend/logs/ui, request details, ownership check) to match either id.
Replaces the metadata transfer approach: that block read top-level
litellm_metadata which guardrail info never populates, and the SLP
builder already reads the nested bucket it lands in, so it was dead
code and is reverted.
Real cause of missing evaluations: process_input_messages skips
apply_guardrail entirely when message scoping (skip_system_message,
skip_tool, scan_only_tool_results) leaves no scannable content, so
the guardrail shows up in applied_guardrails with no
guardrail_information entry. Now records a not_run entry unless the
guardrail records its own information.
max_budget and reset_at already live on the matching budget_limits entry, so
repeating them (as budget_limit and reset_at) only invited confusion about which
copy is authoritative.