Commit graph

43770 commits

Author SHA1 Message Date
ryan-crabbe-berri
6f390fc12e refactor(ui): read throttling off enforcement instead of a caveat code
`enforcement` now carries "throttled" as its own mode, so the table reads the
field rather than inferring it from a note that only explains the mechanism. A
row is throttling even when the note is absent, and a note without the mode no
longer makes the table claim one.

The enforcement badge is a lookup exhaustive over the union, so a fourth mode
fails this build instead of silently defaulting to "Blocks requests", which is
the claim that was wrong in the first place.
2026-08-19 18:34:35 -07:00
ryan-crabbe-berri
db6eec50ff feat(key budgets): model a throttling key as its own enforcement mode
A key with throttle_on_budget_exceeded shipped as enforcement "hard" with a note
explaining otherwise, so any client that did not parse the prose reported a denial
that never happened, and flagged that row as the one that stopped the request. That
is the same mistake the joined note string made, one field up.

`enforcement` gains "throttled". At the limit such a key is admitted by both budget
layers: the read-time check sets a throttle percentage and returns, and the
reservation releases the entry it built for that one counter. What follows is a
reduced tpm/rpm, not a rejection. status stays "exceeded" and comparison stays ">="
because both are still true, and "throttled" is what stops them reading as a block.

The scoping matters and is now testable: only the key's own max_budget throttles.
Key windows, team, team member, user, org, tag and end user all still raise on the
same key, and the flag does nothing at all without a rate limit to scale or a
configured percentage. The note drops to info, since `enforcement` now carries the
fact and the note only explains the mechanism, which is the same rule the other ten
codes are classified by.

Also replaces the lru_cache on is_info_route with precomputed sets. The cache keyed
on a route carrying resolved ids, so its working set was unbounded on exactly the
traffic that would need it: a proxy with end user budgets serving per-resource GETs
would have paid a miss plus cache churn every request. Matching an exact frozenset
and the one templated pattern is 51x faster than the generic matcher on all-distinct
routes, with no state. A test pins it against check_route_access over 372 routes.
2026-08-19 18:32:08 -07:00
ryan-crabbe-berri
06c622ec50 fix(ui): keep a cold per-model budget live instead of rendering it as dead
A cold counter is transient, not a property of the budget. A key created minutes
ago with a per-model cap is fully live and blocks on the next request over it,
but the row read "Cannot trip", greyed, sorted below unlimited scopes, telling
someone to ignore the cap that stops them a minute later. Only a permanent
property counts as dead now, which leaves the project budget whose spend is
never incremented and the personal budget a team key never applies.

The fixture hid it: no_counter was built with no notes and a computed remaining,
neither of which the resolver can emit, so the row under test was not the row
users see. Fixtures now run through the invariants _to_entry guarantees, which
immediately caught a second one.
2026-08-19 18:32:00 -07:00
ryan-crabbe-berri
6a931a9def test(ui): stop asserting server note prose, which was promised to change
`code` and `severity` are the contract and `text` is explicitly free to be
reworded, so pinning the resolver's sentences was asserting the one field
guaranteed to move. It went stale three times in a day, silently each time,
because a fixture copy keeps passing long after the server stops agreeing.

Note texts are synthetic and derived from the code. The wrapping guard uses a
string past 500 characters rather than a copy of the current worst case, so no
rewording upstream can move it and it still fails if truncation returns.
2026-08-19 18:28:06 -07:00
ryan-crabbe-berri
7fa030debc test(ui): re-measure the budget note width pins against the current resolver
The reworded caveats moved the ceiling. The longest single note is now
reservation_blocks_at_limit at 192 characters, not per_model_counters at 165,
and the widest row runs 382 characters across three notes rather than 276.
Re-point the width pin at the note that actually holds the record and refresh
the snapshotted texts, so the fixtures stop asserting prose the server retired.
2026-08-19 18:24:27 -07:00
ryan-crabbe-berri
ed7cee5d50 test(ui): pin per-model row sorting, and stop fixtures encoding old severities
A cap reachable under several request models emits a row each, so one can be
exceeded while its siblings are fine. Pin that the exceeded row floats to the top
and the siblings hold the server's order behind it rather than being reshuffled.

Severity now turns on whether the row already carries the fact in a field, which
is orthogonal to whether the row is dead, and all four combinations occur. Assert
the truth table so no fixture can quietly reassert that severity implies deadness.
2026-08-19 18:21:57 -07:00
ryan-crabbe-berri
6a6809a16a fix(key budgets): classify note severity on whether the row carries the fact
`end_user_route_only` was tagged info while its text is the only place the row says
it applies to requests naming that end user. Nothing in `scope`, `max_budget` or
`comparison` scopes it, so a client that does not know the code had no way to learn
the row might not be in play, on the row most likely to answer "what blocked me".

The definition was the weaker half of that. "The numbers are accurate" left a note
about applicability unclassifiable, because a row can be perfectly accurate and still
not apply. Severity now turns on one question with no exceptions: does a field on the
row already carry this fact. `enforcement` carries alert_only, `comparison` carries
the reservation note, `window_start` carries the rolling window, `max_budget` carries
the personal budget a team key ignores, and `spend_state` carries the missing counter,
so those stay info. The five that only the note carries stay warning, and the end user
scoping note joins them. No other code changes, which is the reason to trust the rule.

A test now pins all eleven against the generated union, so a twelfth code fails until
someone decides which it is, mirroring the exhaustive switch the dashboard compiles.
2026-08-19 18:18:36 -07:00
ryan-crabbe-berri
d16a51a969 fix(ui): tell per-model budget rows apart, and stop calling a throttle a block
Per-model caps now report one row per request model, so a cap reachable under two
names arrives twice with the same label. The scope cell shows the request model
whose counter each row measures, or the two rows read as duplicates.

A key that opted into throttle_on_budget_exceeded was rendered as blocking, with
a "Blocks at" threshold and a red exceeded badge, when going over actually slows
requests instead of rejecting them. Such a row is never the cause of a denial.

Deadness is now read from the caveat code alone. Severity no longer tracks it in
either direction, since dead codes ship under both values, so an unclassified
code is assumed live: calling a row dead when it is not invites dismissing the
budget that actually stopped the request.
2026-08-19 18:17:58 -07:00
ryan-crabbe-berri
83ef72a76a fix(key budgets): stop the report inventing denials and leaking keys through the URL
Round-2 review fixes on the budgets endpoint.

Making the route an info route also enrolled it in proxy-only error logging, which
hands the request path to every failure callback. That gate exists to stop management
endpoints leaking temporary keys, and /key/info is immune only because its key is a
query param. The path form now rejects a plaintext key and asks for the hash, since a
key in a URL also reaches access logs and span names regardless of this gate.

The reservation layer builds no counter for a non-positive cap, so reporting its
tightened operator there turned a team, tag or end user sitting at max_budget 0.0 into
"exceeded" when both layers admit the request. The tightening now needs a positive cap.

Per-model budgets are enforced per request model, not per cap, so reporting the highest
counter under a cap claimed a denial that no request would hit. Each request model that
maps onto a cap gets its own row, identified by the model whose counter it reports.
Deployment names join the candidates, because routing straight at a deployment keys the
counter on that name rather than on a model group.

The reservation note claimed more than the layer delivers: a request it cannot price up
front is gated by the read-time check alone. Reworded rather than re-implemented, since
introspection cannot know the request.

Severity is now the fallback for a code a client has not been taught yet, not a ranking:
warning means the numbers may be incomplete or misread, info means they are accurate.

Also dropped the token end-user cap plumbing, which could only ever attribute the
caller's request-scoped cap to the inspected key, and memoised is_info_route, which the
pattern change had put 25 uncached regex builds per request behind.
2026-08-19 18:11:51 -07:00
ryan-crabbe-berri
af0c016b4f feat(ui): branch the key budgets table on caveat codes and spend state
Each caveat now arrives as its own note with a stable code, so the table renders
one line per caveat at its own severity instead of one prose blob, and decides
whether a row is dead by code rather than by matching wording.

A row that structurally cannot trip, like a project budget whose spend is never
incremented, now says so and sorts below every row that can, since it is never
the answer to which budget stopped a request.

spend_state replaces inferring absence from a null spend. A no-counter-yet zero
keeps its $0.00 and its meter because it is genuinely zero, while a failed read
still shows Unknown with no meter. Any state this build predates is treated as
unreadable rather than drawn as a confident number.

CODE_KILLS_ROW is exhaustive over the code union, so a caveat added server-side
fails the build until it is classified, and severity is the runtime fallback for
a code that arrives from a newer server than this bundle.
2026-08-19 17:54:20 -07:00
ryan-crabbe-berri
9cbf871394 refactor(key budgets): give each budget caveat a code, a severity and its own row
`note` shipped as one prose string, so a client had to substring-match it to act
on anything, and two caveats arrived joined by "; " with no way to render them as
the list they were. Each caveat is now a `KeyBudgetNote` with a stable `code` to
branch on, an `info`/`warning` severity, and `text` that stays free to reword, and
`notes` is always a list rather than a nullable string.

Severity is the difference between a row that cannot trip, like a project budget
whose spend is never incremented, and one that needs reading, like a rolling window
or a scope the reservation layer already blocks at the limit.

The absence of a spend number was also being smuggled through the same field. It is
data quality, not a caveat, so it moved to `spend_state`: `live`, `no_counter` for a
per-model budget whose cache-only counter does not exist yet, and `unavailable` for
a read that failed. A failed read no longer renders as untouched headroom.

The endpoint has not shipped, so this costs nothing today and would be a breaking
change tomorrow.
2026-08-19 17:46:50 -07:00
ryan-crabbe-berri
1cbf830083 fix(ui): stop rendering an unreadable budget spend as a healthy $0.00
SpendBudgetCell coerces a null spend to 0 and draws an empty meter, so a
budget whose live counter could not be read rendered as untouched headroom.
The key budgets table now branches before that: an unreadable spend reads
"Unknown" against its limit with no meter, which is the difference between
"we could not tell you" and "there is nothing to tell".
2026-08-19 17:35:14 -07:00
ryan-crabbe-berri
33b519e014 test(ui): pin the worst-case budget note against re-truncation
The endpoint's longest reachable note is 276 characters over three clauses,
on an end_user row where budget reservation tightened the operator and a
custom auth callable is configured. Assert the whole string renders and that
the note carries no clipping utility, so re-adding truncate fails the test
rather than silently hiding a caveat.
2026-08-19 17:32:23 -07:00
ryan-crabbe-berri
7e6546e8cc fix(ui): let key budget notes wrap, and stop encoding a per-scope operator
Regenerates schema.d.ts for the end-user docstring change, which the types sync
gate needs.

Notes now carry two clauses, e.g. "alert only, never blocks; compared against
recorded spend rather than the live counter", and the scope cell truncated them
to one clipped line behind a native tooltip. They are the most load-bearing text
in the table, so they wrap now and the column is wider to suit.

The threshold helper always read `comparison` per row, but its doc comment and
the fixtures around it asserted that team enforces ">", which is no longer true:
reservation tightens team, tag and end_user to ">=", and disabling reservation
relaxes them back. Nothing about the rendering changes, but the tests now pin
that one scope can render either operator instead of teaching a wrong default.
2026-08-19 17:26:28 -07:00
ryan-crabbe-berri
2c608b05a6 fix(key budgets): restrict end user budget lookups to proxy admins
end_user_id was free-form and the lookup was a global find_unique with nothing tying the end
user to the key, its team or the caller. Any valid key could read any end user's alias, budget,
live spend and reset times, and a null entity_label told the caller whether an id existed at
all. /key/budgets made that reachable for every role, since a caller trivially satisfies the
key check by inspecting their own key.

End users are a proxy-global namespace, LiteLLM_EndUserTable carries no team or organization
column, so there is nothing to scope a non-admin against. Gate the parameter on proxy admin and
admin viewer, matching /customer/info and /customer/list. The gate is on the parameter, not the
route, so a non-admin reading their own key's budgets is unaffected.
2026-08-19 17:15:26 -07:00
ryan-crabbe-berri
1282a6d232 fix(key budgets): report the limit that actually stops a request
Four ways the report disagreed with enforcement.

The comparison came from auth_checks alone, but reservation runs first and refuses once spend
has reached the cap, so team, tag and end_user really block at >= while the row claimed >. A
team at 300 of 300 read "ok" next to a member at 50 of 50 reading "exceeded", with nothing on
the row to explain the difference. Every scope reservation covers now reports >=, status
follows that same operator, and a row whose operator was tightened says so. Scopes reservation
does not cover keep their read-time operator, and so does everything when
disable_budget_reservation is set.

Per-model spend is counted under the request model, but the report read the model_max_budget
key. A cap on "gpt-4o" with callers sending "openai/gpt-4o" enforced against a warm counter
while the row showed no spend at all and blamed a cold cache. Probe every model that routes to
the cap and report the highest, since each counter is compared against the cap on its own.
get_request_model_budget_key now owns that matching so enforcement and the report share it.

A single malformed budget_limits or model_max_budget entry dropped every window or model in
the scope, hiding budgets that were still being enforced. Bad entries are skipped one at a
time now.

The end user cap came from the row, but reservation reads the request token's first. Mirror
that precedence, and say plainly when a custom auth callable could be setting one out of view.
2026-08-19 17:10:23 -07:00
ryan-crabbe-berri
dd720b6385 fix(auth): match templated info routes by pattern
is_info_route compared the request path against info_routes with a bare `in`, so an entry
carrying a path parameter could never match: the incoming route holds a resolved id, not the
`{key_id}` template. /key/{key_id}/budgets was therefore refused for the view-only, team and
customer roles that /key/info serves, and its info_routes entry matched nothing at all.

Route it through check_route_access, the way is_management_route already does. It is the only
templated entry in the list today, so no other route changes reachability.
2026-08-19 17:07:40 -07:00
ryan-crabbe-berri
3b5ab06a95 fix(ui): state each key budget's real threshold in the Budgets tab
Scopes disagree on whether hitting the limit exactly is over it. team_member
enforces >= so 50 of 50 is already denied, while team enforces > so 300 of 300
still passes. The table rendered the numbers and the server's status but not the
operator, so two rows could show identical spend and limit with opposite
statuses and nothing on screen explaining why.

Each row with a limit now states its own threshold, "Blocks at >= $300.00"
against "Blocks at > $300.00", and a soft row says "Alerts at" so the wording
never promises a block it cannot make.
2026-08-19 15:58:16 -07:00
ryan-crabbe-berri
d77cb4a427 chore(ui): regenerate schema.d.ts for the tightened key budgets types
Picks up entity_type as a $ref to Litellm_EntityType and the now-required key
field on KeyBudgetsResponse.
2026-08-19 15:53:44 -07:00
ryan-crabbe-berri
4b8368a27d refactor(key management): tighten the key budgets response types
entity_type was declared str, so the generated client saw a bare string
even though the resolver already carried a Litellm_EntityType and only
called .value at the boundary. Declaring the enum keeps the wire format
identical and lets a caller match a BudgetExceededError's entity against
a row without comparing loose strings.

key is populated on every 200, since a request that resolves no key 404s
before the response is built, so declaring it optional understated the
contract.
2026-08-19 15:52:24 -07:00
ryan-crabbe-berri
83cbee007a fix(auth): stop an exhausted team member budget from blocking management routes
The team member budget check ran twice. common_checks calls
_check_team_member_budget behind skip_all_budget_checks, which correctly leaves
non-LLM routes alone, while a second inline copy in _user_api_key_auth_builder
was gated only on the zero-cost-model check and so fired on every route. A
member who had spent their in-team budget got a 429 from /key/info and could no
longer see which budget had stopped them.

The surviving check is a superset of the deleted one: it accepts a per-member
max_budget of 0 as an explicit disable where the inline copy required > 0, it
compares with >= where the inline copy used >, and it adds the team-level
team_member_budget_id default the inline copy never read. common_checks runs
unconditionally for every authenticated request, so nothing stops being
enforced on LLM routes.

Behaviour change: management and UI routes are no longer blocked by an
exhausted team member budget. LLM routes still are.
2026-08-19 15:38:55 -07:00
ryan-crabbe-berri
5833d99e40 feat(key management): add GET /key/{key_id}/budgets
A BudgetExceededError names one entity, so a caller who gets a 429 still has to
read auth source to work out which of the key, its windows, its per-model caps,
its team, their membership in that team, the owning user, org, project, the
key's tags, the end user or the proxy-wide limit produced it. This returns all
of them in one call, with the live spend and reset schedule of each, including
the scopes that are left unconfigured so they can be ruled out without opening
every object.

GET /key/budgets reports the calling key. Both routes reuse
_can_user_query_key_info, so reading another key's budgets needs the same rights
as reading its info, and 404 on an unknown key matches /key/info.

The report has to agree with enforcement or it is worse than nothing, so the
resolver consumes the same UserAPIKeyAuth get_key_object hands the auth path,
reads spend through get_current_spend, and shares the limit resolution with the
checks: counter key strings now come from one spend_counter_keys module, and the
team-member, personal-budget-on-team-key and budget-org-id rules were extracted
out of auth_checks for both callers. Each row carries the operator its check
actually uses, since they differ per scope, plus a note where a budget cannot
behave the way its numbers suggest.
2026-08-19 15:38:28 -07:00
ryan-crabbe-berri
5b573c552d feat(ui): add a Budgets tab to the virtual key detail page
Renders GET /key/{key_id}/budgets as a table of every budget that can gate the
key, so a 429 that names no entity can be traced to a row without reading auth
source. Rows sort blocking-first, and a soft budget that is over reads as
"Exceeded (alert only)" against "Blocks requests", so an alert can never be
mistaken for the thing that rejected the request. Scopes with nothing configured
still get a row, rendered "Unlimited" rather than $0, which is what lets someone
rule out the org and the team without clicking into them.

The panel is deliberately not keepMounted, so the request is lazy and its cells
cannot collide with the keepMounted Overview panel.

InheritedBudgetHint keeps its client-side team/org guess on the Overview card and
the keys list, where no tab is in reach, but its tooltip now says it is not the
full list and points at this tab, so the two cannot read as competing answers.
2026-08-19 15:00:12 -07:00
Mateo Wang
59c7e7a17d
Merge pull request #37360 from BerriAI/litellm_lit_5729_e2e_record_replay_seam
feat(e2e): add record/replay transport seam and fixture bundle format
2026-08-19 14:28:03 -07:00
tin-berri
0de60a2ff2
fix(mcp): stop reporting failed OpenAPI tool calls as successes (#37496)
An OpenAPI-backed MCP tool whose upstream answered 401 came back as a
successful tool result carrying the upstream's rejection as its content, so a
caller saw {"error":"invalid_token"} presented as data and the gateway recorded
the request in its own spend log as call_mcp_tool | success.

Three layers each erased the outcome. The request function returned
response.text whatever the status, _handle_local_mcp_tool caught every exception
and returned it as ordinary TextContent, and both dispatch sites then stamped
isError=False unconditionally. Fixing only the first, which is the obvious fix,
changes nothing, because the two above it still map failure onto the
success-shaped value.

The status is now classified where the response is held: a 401 becomes
MCPUpstreamAuthError so the caller is told to re-authenticate, and every other
non-2xx becomes MCPOpenApiUpstreamError, which carries the status and drops the
upstream body rather than serving it as tool content. _handle_local_mcp_tool no
longer swallows, and the call_tool arm keeps the auth error's type. Nothing new
renders these: call_mcp_tool and call_tool_rest_api already turn them into an
isError result naming the status and into a real 401 with WWW-Authenticate, and
the OpenAPI path simply never reached them.

The result is now byte-identical to the regular MCP path for the same failure.
2026-08-19 14:26:02 -07:00
tin-berri
afbfc3f8fa
fix(complexity-router): gate the reasoning override on a non-SIMPLE score (#37500)
Two or more reasoning keyword matches promoted a request straight to the
REASONING tier no matter what the weighted score said, so "hi, step by step,
pros and cons" scored 0.100 and still bought the most expensive tier.

Require the score to clear the simple_medium boundary before the override
applies. Promotion from MEDIUM or COMPLEX is unchanged; only prompts the
scorer already placed in the cheapest band stay there.
2026-08-19 14:19:24 -07:00
Mateo Wang
eec27a9cb3
Merge pull request #36593 from BerriAI/devin_ai_lit_5445_perplexity_stream_dict_cost
fix(streaming): accept provider cost objects when propagating usage cost
2026-08-19 14:18:36 -07:00
ryan-crabbe-berri
f1e143a87c
chore(ui): upgrade the dashboard to React 19 (#37411)
* chore(ui): upgrade the dashboard to React 19

Bumps react and react-dom from 18.3.1 to 19.2.8 with matching @types. Next 16 already required a React 19 peer, so this aligns the dashboard with what the framework expects and unblocks Base UI and shadcn work that assumes the React 19 ref model.

React 19 passes ref through as a regular prop, so the setup file's forwardRef tripwire and the ref-forwarding test's forwardRef case no longer describe real behavior; both now assert the React 19 contract instead. useRef<T>(null) now yields RefObject<T | null>, which is the one prop type MessageList had to widen.

* test(ui): wait for a Base UI select popup to open before clicking an option

The option lands in the DOM one render before the popup finishes entering, while its positioner still carries pointer-events: none, so clicking it throws. Waiting on the option's text alone was a race that React 19's flush timing loses, which is why four ToolPolicies cases went red on the bump.

chooseSelectOption in test-utils opens the trigger, finds the option by role, waits for it to stop being pointer-blocked, then clicks. It also replaces the last-match-by-text hack, which only worked because the popup happens to portal after the table.
2026-08-19 21:18:08 +00:00
yuneng-jiang
3d51eb378a
refactor(ui): migrate the antd Alert call sites onto the shared Alert (#37513)
Moves all 33 antd Alert usages across 20 dashboard files onto
src/components/shared/Alert, following the composition the rest of the
dashboard already uses: message becomes AlertTitle, description becomes
AlertDescription, showIcon becomes a lucide icon child, and closable
becomes an AlertAction ghost button.

antd type="success" has no counterpart on the shared Alert, so the two
success sites land on the default variant with a CircleCheck icon, which
is what cloudzero_export_modal and CloudZeroIntegrationSettings already
do for the same case.

LoginPage's dismissible SSO notice moves into its own SsoEnabledNotice
component in the same file: antd's closable carried its own dismiss
state, and inlining it pushed LoginPageContent past the complexity
budget.

Four files lose their last antd symbol, so their no-restricted-imports
suppressions are pruned by hand. antd import sites drop from 115 to 111
across 107 to 103 files, and the no-restricted-imports ratchet drops
from 119 to 115 over 110 to 106 files.

One test asserted antd's own ant-alert-info class; it is repointed to
the shared Alert's text-info variant class, which keeps the same
"info, not warning" check. Every other colocated test passes untouched.
2026-08-19 21:07:50 +00:00
tin-berri
a613773fca
feat(auto-router)!: scope shadow eval jobs to multiple keys (#37251)
* feat(auto-router): scope shadow eval jobs to multiple keys

A shadow eval job now covers a set of keys instead of exactly one, and each
key carries its own max_turns budget, so one key exhausting its budget leaves
its siblings sampling. The existing job row already is the per-key unit
(api_key_id, max_turns, stopped_at, and the one-active-per-key-and-direction
partial unique index all live on it), so multi-key is grouping rather than
schema surgery: a new group_id column ties N sibling rows written atomically
by one create_many, the API's job id becomes the group id, and pre-existing
jobs backfill group_id = id so their ids keep resolving. The sampler hot path
is untouched; its test file has a zero-line diff

Results come back pooled plus a per-key breakdown and responses list every key
with its own budget, stop state and read-time labels. The dashboard is adapted
minimally to the new shapes (the picker stays single-key and submits a one-key
list); the multi-select picker and per-key table land in the stacked UI PR

* fix(shadow_eval): derive completed from spent budgets and record operator stops

* fix(shadow_eval): stamp stops atomically and freeze counts at the stamp

The stop endpoint wrote stopped_by and stopped_at as two separate updates, so
a failure between them left a job reading stopped while its unstamped legs
kept sampling, and the retry got 400 already stopped. One UPDATE now stamps
stopped_by and every missing stopped_at together, preserving the stopped_at a
leg earned from its own budget via COALESCE

Attempt counts now exclude attempts that land after a leg's stopped_at, so an
in-flight attempt finishing just after an operator stop can never push a
legacy pre-stopped_by job over its budget and flip it from stopped to
completed at read time

* fix(shadow_eval): backfill stopped_by so legacy stops never read as completions

* chore(ui): regenerate api types for the shadow eval stop fields

* fix(shadow_eval): let the stop statement pick one winner under racing stops

Two operators can both pass the derived-status guard in the race window. The
stop UPDATE now claims only legs with stopped_by still null and the endpoint
judges by its row count, so exactly one caller ever gets the 200 and the loser
gets the same already-stopped 400 a late caller gets

* refactor(shadow_eval): make the stop statement the whole state machine

The status guard ran before the UPDATE, so a stop racing the last budgeted
attempt still claimed the job and it read stopped forever instead of
completed. The statement now claims the job only while a leg still samples
inside the window with no stop recorded, and the endpoint reads once after
writing: a racing operator, a same-instant budget spend, and a repeat stop all
get the 400 naming the status the job actually holds. The pre-write guard and
the hand-built response go away

* chore(ui): regenerate api types for the stop route description
2026-08-19 14:02:15 -07:00
yucheng-berri
a1afc2f433
refactor(ptu): give the rollup a source-agnostic deployment record (#37501)
The flat-cost rollup reads deployments only from LiteLLM_ProxyModelTable, so a PTU
deployment declared in config.yaml never accrues flat cost. Those deployments live in
llm_router.model_list as plain dicts whose id sits in model_info rather than on the entry,
so they do not satisfy the shape _parse_ptu_model reads.

Adds a frozen record in that shape and a factory that maps a router entry onto it, leaving
_parse_ptu_model byte-identical so the existing cases stand as evidence of no behaviour
change. Nothing calls the factory yet; the caller lands with the loader union.

_decode_model_info also stops handing back valid JSON that is not an object. It decoded
a list or a scalar and returned it as a mapping, so the caller read fields off it and
raised, losing the whole run rather than the one bad deployment.
2026-08-19 13:51:31 -07:00
yuneng-jiang
7675ba8717
test(ui): split the vitest suite into unit, component, integration and type projects (#37488)
* test(ui): split the vitest suite into unit, component and integration tiers

Every test file booted jsdom, including the ~1800 that assert pure functions
and never render. They now run as a separate vitest project in the node
environment, where the whole tier finishes in under four seconds.

The tiers are vitest projects rather than a naming convention, so CI can run
them as independent jobs. A .test.ts that renders React, a hook test being the
usual case, is listed explicitly and stays in the jsdom tier.

* test(ui): report per-test duration against a per-tier budget

A timeout only catches a hung test, and it has to stay generous enough to
survive a loaded runner, so it never reports the multi-second render tests that
make CI fail the moment the box is busy. Budgets are separate and far tighter:
50ms unit, 1s component, 3s integration.

The counts are laptop measurements, so the CI job is report-only for now.
Flipping it to blocking is one line once CI has published its own numbers.

* test(ui): run the tiers as separate CI jobs and stop clicking popups by text

The old job ran every file in one process, so the single slowest file set the
wall clock and a bigger box bought nothing. The tiers now run as separate jobs
with the component tier sharded four ways.

getByText and findByText match hidden nodes, so they resolve against a closed
Base UI popup whose positioner still carries pointer-events: none, and the
click lands or not depending on how far the open transition got. Two files
failed this way, one three runs in five and one every run. Querying the option
by role waits for it to be visible, and both are now stable. A lint rule keeps
the pattern from coming back.

* test(ui): give React Testing Library's async queries a CI-sized window

findBy* and waitFor run on asyncUtilTimeout, which defaults to 1000ms and is
independent of vitest's testTimeout. Raising the vitest timeout therefore did
nothing for them: a query still gave up after one second while the test had 59
seconds of budget left, which is why a loaded runner produced 'Unable to find
role=...' rather than a timeout.

UserSearchModal is the worked example. The role query it makes resolves in
249ms on a laptop and blew past 1000ms on CI, failing the run at 1494ms. Five
seconds keeps the same assertions and only widens the window a failing query
waits before reporting; a passing query still resolves the moment the element
appears.

* test(ui): calibrate the tier budgets from real CI numbers and report by default

The first CI run showed the laptop counts were badly off: component 176 local
against 326 on CI, integration 87 against 128. The maxima now come from that
run with headroom.

continue-on-error still painted the check red, which is the opposite of the
point, so the report-only decision moves into test-budgets.json as an explicit
enforce flag. The job passes and prints the counts; flipping enforce to true
makes it a gate.

* docs(ui): drop the CLAUDE.md edits from the tier split

Keeping this PR to the vitest, CI and test changes.

* ci(ui): run every tier in one job instead of eight check rows

Sharding bought nothing. Measured on the first run of this branch, the
component tier unsharded finishes in 198s while the integration tier is floored
at 384s by a single file, so integration was always the critical path and the
four component shards only added rows. One job running every project comes in
around 384s against the 426s the split jobs took.

Eight rows named things like 'component (2)' also told a reviewer nothing, on a
PR page that already carries forty checks.

The job keeps the id ui-unit-tests because guard-internal-staging requires that
exact context; renaming the jobs had silently stopped it reporting, which would
have blocked every merge on a check that no longer existed. The workflow's
display name becomes UI Tests since it runs more than unit tests.

The tier split itself is untouched: it lives in the vitest projects config, so
the unit tier still runs in node with no jsdom, and each tier keeps its own
timeout and budget.

* fix(ui): stop the type check from running the whole suite a second time

test:types was 'vitest --run --typecheck.only'. Under test.projects that flag
is ignored and the root-level typecheck block is not inherited, so the step
collected each project's normal include and ran all 8464 runtime tests instead
of type-checking. It took 542s on CI against 33s on the flat config it
replaced, and the job then ran the same suite again in the next step.

Typecheck now belongs to a project of its own, with an empty include so it
contributes no runtime tests, and the CI job runs one vitest invocation for all
four. The type tier adds about 3s to a full run and reports 'Type Errors: no
errors' rather than a suite of tests.

Verified it still catches things: breaking SortingState in DataTable.test-d.tsx
fails with 'Type number is not assignable to type string' and exit 1, and
restoring it passes.

* test(ui): scope the split down to the vitest tier projects

Removes everything from this branch that was not the tier split.

The three lint rules brought 1381 lines of grandfathered suppressions in
eslint-suppressions.json, which is 81% of the branch's added lines and
debt nobody is going to pay down. The per-test duration budget does not
scale as a CI step. Both are gone, along with the two query rewrites the
no-click-by-text rule forced: those files pass 10/10 at this base, quiet
and under load, so there was no failure behind them.

The workflow is byte-identical to the base again. It already runs
npm run test:types and then vitest related on pull requests, so PR cost
is unchanged; the split only repoints test:types at the new project.
That project is required, not optional: vitest silently ignores
--typecheck.only under test.projects, so without it the type script
collects the whole suite instead of the one typed file.

Restores the base 60s testTimeout on the unit tier. The 5s cap was not
part of the split and failed ChatShell.serverRootPath.test.ts, a 960ms
test, under load.
2026-08-19 20:40:14 +00:00
yuneng-jiang
e126975468
refactor(ui): migrate the antd Button call sites onto the shadcn Button (#37505)
Moves all 58 antd Button JSX sites across 24 files onto the shadcn
Button, leaving zero antd Button importers.

Prop mapping follows what already merged rather than a new convention:
type="primary" to the default variant, a bare button to outline (the
house default), type="text" to ghost, type="link" to link,
type="dashed" to outline plus border-dashed, danger to destructive,
size="small" to sm, htmlType to type, block to w-full, the icon prop to
a child, and loading to disabled plus aria-busy. Icon-only buttons take
the matching icon-* size. Inline style props that had a direct utility
equivalent moved to className, and the opacity toggle on the create key
submit is dropped since the base cva already carries disabled:opacity-50.

Base UI's Button defaults type to "button" for native buttons, the same
default antd used, so bare buttons inside a form do not start
submitting.

guardrail_info.tsx and tag_info.tsx lose their last antd import, so
their no-restricted-imports suppressions are pruned. The other 22 files
keep other antd symbols and keep their entries.

One behavior change: the MCP transports docs link now opens in a new
tab, matching every other external docs link in the dashboard, instead
of navigating the dashboard away.

The ModelSettingsModal loading assertion moves off antd's spinner
element onto aria-busy, which is what the rest of the suite already
asserts; that markup cannot survive removing antd.
2026-08-19 13:32:00 -07:00
Mateo Wang
0f19b5b9ab
Merge pull request #37361 from sytianhe/litellm_spend_log_timestamps
feat(spend-logs): add lifecycle timestamps
2026-08-19 13:29:49 -07:00
ryan-crabbe-berri
73e7105e60
fix(ui): restore tab strip styling and panel persistence lost in the shadcn migration (#37403)
Tremor's TabList defaulted to the underlined `line` variant and its TabPanel
rendered every panel, hiding the inactive ones with a class. The shadcn
TabsList defaults to a segmented pill and Base UI's TabsPanel unmounts a
hidden panel unless it carries `keepMounted`, so the conversion waves quietly
changed both on the pages that took tremor's defaults.

Restores the underline on the nine strips whose tremor markup carried no
`variant`, leaving the ones that were explicitly `variant="solid"` as pills,
and puts `keepMounted` back on the thirteen files whose panels used to stay
mounted, so filters, scroll position and in-progress input survive a tab
switch again.

Seeding the old usage page's activity state properly comes with it: it was
cast from `{}`, so the panel crashed on `data.length` the moment it mounted
before its fetch resolved, which only stayed hidden while the panel was
unmounted until first opened.
2026-08-19 13:28:20 -07:00
mateo-berri
6acf970009 chore(ui): regenerate dashboard API types for spend log timestamps 2026-08-19 12:41:52 -07:00
yuneng-jiang
4bb3152cc5
test(ui): drive fields with change events where the typing is not the behaviour (#37495)
user.type dispatches one event per character and re-renders the whole form on
each one, so a test that only needs a field to hold a value pays for every
keystroke. Where the value is all the test wants, fireEvent.change sets it in
one go.

Measured against a control on the same tree: the 100 converted files went from
654s to 619s of CPU, 5.4% cheaper, while the files nobody touched drifted 1.2%
the other way. Modest, and honest about it.

Rolled out one file at a time, running each before and after and keeping the
conversion only where the file stayed green. That rejected 35 files, all cases
where the keystrokes are the behaviour: Base UI comboboxes drive their filter
from real keyboard input and ignore a raw change event, and the same goes for
autocompletes, debounces and key handlers. Those keep user.type.
2026-08-19 12:11:37 -07:00
Mateo Wang
4d100bdd89
Merge pull request #37439 from BerriAI/litellm_decrease_anys_fable_round3
chore(typing): drop 1.3k basedpyright errors across 42 Any hotspot files
2026-08-19 12:07:26 -07:00
tin-berri
da7a10ebbd
fix(mcp): forward the per-server auth header on OpenAPI tool calls (#37410)
Both OpenAPI dispatch arms sourced the upstream credential only from the
deprecated global / BYOK mcp_auth_header and never from mcp_server_auth_headers,
so x-mcp-{alias}-authorization was silently dropped on spec_path servers and the
upstream API received no Authorization at all. The managed path already resolves
it through lookup_mcp_server_auth_in_headers, so the two had drifted.

_resolve_openapi_tool_auth now owns that resolution for both arms. A per-server
value is already a complete header value and is forwarded verbatim, while a BYOK
credential keeps its auth-type prefix, so the two are never conflated into
"Bearer Bearer <token>". The resolved credential is also handed to
resolve_openapi_upstream_auth, whose passthrough arm reads it through
_passthrough_token_from_mcp_auth_header and outranks the ContextVar.

server.py loses its inlined copy of the forwarded-header logic along with its
mcp_server is None guards, which are unreachable after the 503 raised above them.

Credit to the earlier analysis and approach in #33349, which this supersedes
against the current v2 credential resolver.
2026-08-19 11:15:19 -07:00
mateo-berri
19b5381e63 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_decrease_anys_fable_round3
# Conflicts:
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-19 10:52:48 -07:00
Mateo Wang
133e72c8fd
Merge pull request #36263 from BerriAI/litellm_replica_registry_read_through
fix(proxy): read through to the DB on registry misses so just-created models, guardrails, and agents resolve on sibling replicas
2026-08-19 10:46:35 -07:00
Mateo Wang
18d07e3a6b
Merge pull request #37432 from BerriAI/litellm_e2e_tag_denial_reason
test(e2e): pin the tag-routing denial to its actual cause
2026-08-19 10:43:27 -07:00
yuneng-jiang
481c08de4e
refactor(ui): port the MCP server forms off antd Form onto react-hook-form (#37483)
* refactor(ui): port the MCP server forms off antd Form onto react-hook-form

The MCP create and edit forms were the last antd `Form` graph in the
dashboard. antd's `onFinish` hands back only the fields mounted at submit
time, while react-hook-form with `shouldUnregister: false` hands back the
whole store, so a direct port would quietly widen every create and update
request.

All 14 files now bind through `MountedFormField`, whose mount registry
reproduces antd's mounted-only submit: both roots build their payload from
`projectMountedValues` instead of `getValues`. `mcpFormStore` carries the
rest of the FormInstance surface the two roots relied on, each piece
matched to what rc-field-form actually does rather than to what the API
name suggests: `setFieldsValue` deep-merges plain objects and writes an
explicit `undefined`, `resetFields` restores the seeded values rather than
clearing the key, and `onValuesChange` is rebuilt from a `watch`
subscription filtered to user input, carrying a single changed branch.

Two watches needed the mount gate moved rather than translated.
`MCPPermissionManagement` renders outside the transport gate that mounts
`auth_type`, so antd's watch read `undefined` there and mounted the
pass-through toggle; the effective auth type now arrives as a prop that
each root computes with exactly that gate. The edit root reads every watch
off the mounted projection for the same reason, which also stops the token
material in a saved server's credentials from reaching the tool preview.

`Form.List` becomes `useFieldArray` plus a `useMountedName` registration
for the list key itself, because antd registers a list as one field: a
per-user variable row keeps the `value` its scope hides, and an empty list
still submits `env_vars: []` instead of dropping the key.

* test(ui): cover the mount registry's unregister path on the real primitive

The existing MountedFormField suite drove a hand-written registry whose
register returned a no-op, so nothing exercised useMountRegistry's
ref-counting or the cleanup that React wires from useMountedName's effect
return value. A reviewer read that gap as a missing unregister.

These three cases drive the real hook through a gated tree: a key leaves
the submitted payload when its gate unmounts the field, a required field
that unmounts stops blocking submission, and a name held by two fields
survives one of them releasing it.

Verified by mutation: rewriting the effect body to discard the cleanup
turns the first two red, the second reporting the reviewer's exact
symptom, "expected [ 'server_name', 'token_url' ] to not include
'token_url'".

* test(ui): prove the permission panel's booleans reach the create payload

CreateMCPServer.integration.test.tsx mocks MCPPermissionManagement, so the
four booleans createServerPayload writes were invisible to every existing
create-side test. vi.mock is file-scoped, so rendering the real panel needs
its own file.

Four cases, each killed by a different mutation:

  unbind allow_all_keys                -> "sends allow_all_keys true"
  unbind available_on_public_internet  -> "sends the panel's defaults"
  invertedSwitchControl -> switchControl -> "sends ... false when the
                                            operator restricts"
  isOAuth2 gate forced open            -> "omits delegate_auth_to_upstream"

A payload assertion expecting false cannot detect an unbound field, since
Boolean(undefined) is false too, so the two cases carrying unbinding
detection are the ones asserting true. The other two are pinned by the
switch-inversion and mount-gate mutations instead.
2026-08-19 10:26:20 -07:00
Mateo Wang
b402fea745
Merge pull request #37451 from BerriAI/litellm_isolate_proxy_url_env_in_tests
fix(tests): keep a host PROXY_BASE_URL out of request-derived URL tests
2026-08-19 10:06:26 -07:00
mateo-berri
f8a23aab09 fix: gate guardrail read-through to active rows and serialize it with the reload reconcile 2026-08-19 01:46:33 -07:00
mateo-berri
1dbc7177bb chore(typing): close budget headroom left after merging staging 2026-08-19 01:03:54 -07:00
mateo-berri
882efc18ac Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_e2e_tag_denial_reason 2026-08-19 01:03:52 -07:00
mateo-berri
967970042a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_fable_round3 2026-08-19 00:59:46 -07:00
mateo-berri
9ef6a8826d test: trim the PROXY_BASE_URL fixture and regression docstrings 2026-08-19 00:56:37 -07:00
mateo-berri
e9c01da233 test: drop unused imports in the direct restore test 2026-08-19 00:54:46 -07:00