Commit graph

41837 commits

Author SHA1 Message Date
tin-berri
2798f39f5a
Merge pull request #34434 from BerriAI/litellm_lit4748_autorouter_logs
feat(ui): show in the log drawer and session sidebar when an auto-router served a request
2026-07-23 22:09:34 -07:00
tin-berri
c93c3f7582
Merge pull request #34439 from BerriAI/litellm_cache_leakage_header_layout
fix(ui): keep cache leakage time range picker inline at narrow widths
2026-07-23 21:53:42 -07:00
Tin Chi Lo
42aba4f32a feat(ui): show in the log drawer and session sidebar when an auto-router served a request
The dashboard already receives the requested model name as model_group on
every spend-log row, but LogEntry dropped the field, so nothing distinguished
an auto-routed request from a direct one.

Surface it precisely rather than by comparing requested against resolved:
model_group differs from model for plain aliases and wildcard deployments
too, so a bare mismatch tags almost every row and identifies nothing. The
indication is driven instead by which deployments are auto-routers, resolved
from every page of /v2/model/info and shared through context.

The request drawer header names the router in a badge next to the provider;
the session sidebar swaps the entry's leading icon. Rows that no auto-router
served render exactly as before.
2026-07-23 20:43:54 -07:00
yuneng-jiang
64aad5877a
fix(proxy): restore atomic user upsert when adding team members (#34457)
* fix(proxy): restore atomic user upsert when adding team members

Parallel /team/new calls naming the same not-yet-existing member were
returning 500 "Unique constraint failed on the fields: (`user_id`)".

The upsert in add_new_member passed an empty update branch. Prisma only
compiles an upsert down to a single INSERT ... ON CONFLICT when that branch
writes something; with an empty one it emits SELECT-then-INSERT instead, so
concurrent requests all read "no such user" and all insert. Postgres
statement logs confirm it: the empty form logs BEGIN/SELECT/INSERT/COMMIT,
the non-empty form logs INSERT ... ON CONFLICT ("user_id") DO UPDATE SET.

Re-state user_id in the update branch as a no-op so the native upsert path
comes back. The teams append stays in the filtered update below it, so an
already-existing member still cannot pick up a duplicate team id.

tests/test_team.py::test_team_new failed 9 of 15 runs against a live proxy
before this and 0 of 15 after. The existing unit test asserted only that
upsert had been called on a mock, so it passed either way; it now pins the
shape of both branches and fails when the update branch goes back to empty.

* test: point the live codex tests at gpt-5.3-codex

OpenAI deprecated gpt-5.2-codex, so test_openai_codex and
test_openai_codex_stream started failing against the live API with
model_not_found. gpt-5.3-codex is the current codex model; both tests pass
on it. The remaining gpt-5.2-codex references in the suite are mocked
transformation tests and are unaffected.

* test(e2e): update models page specs for the shared DataTable

The DataTable migration in #34363 changed three things the models page
specs were pinned to, and five tests went red.

Row click no longer opens the detail view; the Model ID cell owns that
now, so both specs click its `model-id-<id>` test id instead of the row.
The search box placeholder switched from an ASCII "..." to a real
ellipsis, so the specs use getByPlaceholder with a substring instead of
an exact attribute match that punctuation can break again. The results
count moved from `models-results-count` ("Showing 1 - 50 of 137 results")
to the shared pagination's `pagination-range` ("Showing 1-50 of 137").

The Team-BYOK test also filtered rows on the team alias, which the Team
ID column has never rendered in either the old or the new table; it
filters on the team id now, which is what the column actually shows and
what the assertion's own comment intends.

Verified against a local proxy serving a fresh build with the seeded
e2e postgres and mock upstream: all five failing tests pass, and the
full suite is 82 passed / 4 skipped at CI parity (workers=1).
2026-07-24 02:22:34 +00:00
Mateo Wang
15af874fcf
Merge pull request #34046 from BerriAI/litellm_openai_cache_write_tokens
fix(cost_tracking): map OpenAI cache_write_tokens for prompt cache creation billing
2026-07-23 19:19:46 -07:00
mateo-berri
d3f5c6dbf6 fix(cost_calculator): sum mirrored cache token fields once in combine_usage_objects
combine_usage_objects iterates prompt_tokens_details model_fields and sums each;
with cache_write_tokens and cache_creation_tokens now mirroring each other via
__setattr__, the pair was summed twice, doubling cache creation counts for
Anthropic batch cost calc, mid-stream fallback usage merges, and realtime usage.
Collapse the mirrored pair to one representative before summing.
2026-07-23 19:07:09 -07:00
Krrish Dholakia
eaf61eb34b test(cost_tracking): cover OpenAI Responses API cache cost breakdown itemization (#34309)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:09 -07:00
Krrish Dholakia
5e1b08b355 fix(spend_tracking): populate cache_creation_input_tokens for Responses API logs
On the /v1/responses path the response usage is not chat-Usage-shaped, so
additional_usage_values could not derive cache tokens from response_obj.usage
and the Admin UI Logs cache-creation token row stayed empty. Fall back to the
normalized standard_logging usage_object's prompt_tokens_details for both the
cache-read and cache-creation counts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:09 -07:00
Krrish Dholakia
2f502a1bfc fix(cost_tracking): map cache_write_tokens on Responses API usage path
The Responses API (/v1/responses) usage transform rebuilt prompt token
details and dropped OpenAI's input_tokens_details.cache_write_tokens, so
gpt-5.6 cache-creation tokens were never logged or billed via that route.
Map it in the transform, and make PromptTokensDetailsWrapper keep
cache_write_tokens and cache_creation_tokens in sync on assignment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:08 -07:00
Krrish Dholakia
e6ec153243 fix(cost_tracking): map OpenAI cache_write_tokens for prompt cache creation billing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:08 -07:00
Tin Chi Lo
dce1b0d1fd fix(ui): allow null for mcp_toolsets in the dashboard key response type
The generated schema declares object_permission.mcp_toolsets as
string[] | null; the handwritten KeyResponse shape omitted the null.
ObjectPermissionsView consumes the same value, so its prop type widens
with it
2026-07-23 18:33:08 -07:00
Tin Chi Lo
a469bb7924 fix(ui): keep a key's MCP toolsets when saving an edit
The key edit form seeded mcp_servers_and_groups from the key with only
servers and accessGroups, but handleKeyUpdate writes mcp_toolsets from
that same value, so every save posted an empty list and the backend
merge applied it literally. A key granted a toolset lost the grant on
any edit, including a budget change, and then got a 403 from
/toolset/<name>/mcp

Read toolsets in both places the form initializes from keyData, declare
mcp_toolsets on KeyResponse.object_permission so a write-without-read is
a type error, and carry toolsets through the create flow, which only
looked at servers and accessGroups
2026-07-23 18:06:57 -07:00
devin-ai-integration[bot]
c255f53bfb
ci: only run CodSpeed on backend changes (#34345)
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 16:58:55 -07:00
ryan-crabbe-berri
07f7fc224e
fix(proxy): reject failed atomic budget reservations under fail_closed_budget_enforcement (#34429)
* fix(proxy): reject request when budget reservation write fails under fail_closed_budget_enforcement

With general_settings.fail_closed_budget_enforcement set to true, the read-time
spend check already returns 503 when spend cannot be verified, but the atomic
pre-call reservation still failed open: reserve_budget_for_request swallowed
_CounterReservationUnavailable per counter and degraded to read-time-only
enforcement, so concurrent requests could all pass the same under-budget read
during a Redis outage and overspend past the configured budget.

Now the strict flag is threaded into reserve_budget_for_request and a failed
reservation write raises 503, releasing any counters that already reserved.
Default behavior with the flag absent or false is unchanged.

Fixes #33923

* fix(proxy): pass 503 budget-enforcement detail as plain string
2026-07-23 23:57:37 +00:00
Yuneng Jiang
caac7d8aa8
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/migrate-page-memory-9b2c09 2026-07-23 16:44:34 -07:00
Yuneng Jiang
58083b978c
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/migrate-page-memory-9b2c09 2026-07-23 16:43:15 -07:00
ryan-crabbe-berri
a507394841
fix(ui): find logs by request id across pages and dates (LIT-3981) (#31743)
* fix(spend): resolve spend logs by request_id across all dates (LIT-3981)

The /spend/logs/ui search only filtered the page already loaded, so a log id
copied from another page or from outside the active date window could not be
found. request_id is the primary key of LiteLLM_SpendLogs, so when it is
supplied on the internal UI route the mandatory date window is dropped and the
lookup resolves across all time. The date window stays required when no
request_id is given, and the public /spend/logs/v2 contract is unchanged.

A non-admin id lookup is gated by the same ownership check the detail endpoint
uses, so the relaxed window cannot be used to read another tenant's log by id

* fix(ui): send the logs request_id search to the server (LIT-3981)

The "Search by Request ID" box filtered only the rows already on the current
page, so an id from another page never matched. It now feeds the existing
server-side request_id filter via handleFilterChange, which debounces, resets
to page one, and rides the existing react-query key. The dead client-side
filter and its searchTerm state are removed; the session composition and dedup
logic is unchanged.

The box is now an exact request_id lookup, matching its label; the incidental
client-side model and user substring matching it used to do is dropped in
favor of the dedicated filters

* refactor(spend): model the request_id spend-log lookup as an explicit point lookup (LIT-3981)

The date-window relaxation for request_id lookups rode an apply_date_window flag threaded through the date validation and parsing. Model the two intents directly instead. A UI request_id query is a point lookup on the @id primary key that drops the time window and authorizes by row ownership; every other query, including the public /spend/logs/v2 route, takes the range-scan path that still requires a window

Because the ownership check fully authorizes the single row, the general user/team scoping is now skipped for id lookups rather than layered on top redundantly. The confusing `is_v2 or request_id is None` guard is gone, and moving the date requirement into the range-scan branch lets the type checker narrow the dates it parses

Behavior is preserved: the v2 contract still requires dates even when a request_id is supplied, and a non-owner is still rejected with 403. A regression test covers the non-admin owner id lookup, which resolves across all time and filters by the primary key alone
2026-07-23 16:40:19 -07:00
ryan-crabbe-berri
7aaaa055b7
feat(proxy): add overwrite_user_with_key_hash to stamp outgoing user param with key hash (#34417)
* feat(proxy): add overwrite_user_with_key_hash to stamp outgoing user param with key hash

Adds a litellm_settings flag that forces the outgoing user param to the
authenticated key's hashed token before the request is forwarded to the
provider. The value overrides any caller-supplied user, so providers see
a stable, tamper-proof identifier they can rate-limit or ban on, and the
hash matches user_api_key_hash in spend logs for easy mapping back to
the key owner. Off by default

* fix(proxy): hash non-sk credentials before stamping user param

UserAPIKeyAuth only hashes sk-prefixed keys and JWTs; custom-auth
credentials stay raw on api_key, so stamping them directly would forward
auth material to the provider. Pass through the two known hashed forms
(sha256 hex, hashed-jwt-*) and hash anything else

* refactor(proxy): stamp only standard virtual keys, skip jwt and custom auth

A hashed JWT rotates on every token re-issue so it is useless as a
stable ban id, and custom-auth credentials arrive raw on api_key.
Instead of hashing whatever we hold, the stamp now applies only when
api_key is the sha256 hex digest of a standard virtual key; other auth
methods are explicitly out of scope until the stamped identifier is
configurable

* fix(proxy): gate user stamping on server-set virtual key provenance

Shape alone cannot distinguish a key hash from a raw custom-auth
credential that happens to be 64 hex chars. Adds via_virtual_key, a
server-only marker on UserAPIKeyAuth following the
mcp_admitted_user_subject pattern: stripped from all validated input so
handlers and claims cannot forge it, set by post-construction assignment
only at the DB virtual-key auth return. Stamping now requires the marker
and the hash shape

* test(proxy): prove db auth path sets via_virtual_key marker

The stamping unit tests set the marker manually, so deleting the
assignment in _user_api_key_auth_builder would pass every existing test;
this exercises the real builder path with a mocked identity store and
fails if the marker is not set

* fix(proxy): stamp master-key requests with the master key alias

Master-key auth substitutes LITELLM_PROXY_MASTER_KEY_ALIAS for api_key
so the key and its hash never propagate; that made master-key traffic
bypass stamping and pass the caller-supplied user through. The master
path now sets via_virtual_key and the stamp gate accepts the alias
alongside the sha256 shape, so admin traffic gets the same tamper-proof
id that spend logs already record for it

* fix(proxy): restore via_virtual_key marker on key-cache hits

Cached PROXY_ADMIN auth objects early-return before the marked DB and
master-key returns, and cache serialization drops the exclude=True
marker, so cached admin traffic bypassed stamping. Key-cache entries are
written only after the proxy validated a virtual key or the master key,
so the cache-hit boundary restores the marker; the UI-login JWT fallback
constructs its token from a decrypted blob, not this cache, and stays
unmarked
2026-07-23 16:38:01 -07:00
shivam
ee9f0db1cf fix(bedrock-mantle): backfill usage on non-streaming /v1/messages responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 23:37:40 +00:00
ryan-crabbe-berri
e906a7e796
refactor(ui): extract shared tab-routing helpers and adopt them in Models + Endpoints (#34435)
* refactor(ui): extract shared tab-routing helpers

Every per-tab-routed page copy-pastes the same URL<->slug logic and the
same active-tab/redirect engine. Extract two reusable pieces:

- createTabRoutes(baseSegment, slugs) in utils/tabRoutes.ts returns
  { baseSegment, slugs, tabHref, slugFromPathname }, the trailing-slash
  href builder (via migratedHref) and the pathname->slug reader.
- useTabRouting({ routes, baseTabKey, visibleKeys, ready }) derives the
  active tab from the pathname, redirects an unknown/forbidden slug to
  base once ready, and returns an onTabChange navigator.

visibleKeys + ready exist so a role-gated page can pass its filtered tab
set and defer the redirect until permissions resolve, rather than
bouncing a user off a still-loading valid tab. Both are pure/unit-tested.
No page consumes them yet.

* refactor(ui): migrate Models + Endpoints onto the shared tab-routing helpers

Replace the page's hand-rolled tabRoutes.ts (base segment + slug tuple +
href builder + slugFromPathname) with createTabRoutes, keeping the
existing named exports as thin re-exports so callers and tests are
unchanged. The layout drops its local activeSlug/isKnownSlug/activeKey
derivation, its redirect useEffect and its router.push onChange in favor
of useTabRouting, passing the role-filtered visibleKeys and a ready flag
(!teamsLoading && !uiSettingsLoading) so the permission-gated redirect
behavior is preserved exactly. The antd tab bar, role-gated tab set, the
refresh button and the ?model=/?team= drill-in overlay are untouched; the
file's pre-existing antd import is now recorded in the suppressions
baseline since editing it makes it a linted-as-changed file.

The existing models-and-endpoints layout.test.tsx and tabRoutes.test.ts
pass unchanged, which is the regression guarantee.
2026-07-23 16:36:25 -07:00
yuneng-jiang
07726b4f60
refactor(ui): migrate agents to shadcn (#34365)
* test(ui): make the agents route's tests markup-agnostic before migration

Rewrites the two assertions that were coupled to antd's DOM and adds the
missing characterisation test for agent_cost_view, so the suite describes
behaviour rather than antd markup and can stay untouched across the shadcn
migration.

The skill selection test reached the checkbox with a querySelector on
input[type=checkbox]; antd renders an input while Base UI renders a
span[role=checkbox], so it now queries by role and accessible name, which
both libraries derive from the wrapping label.

The delete confirmation test queried role=dialog; antd Modal is a dialog
while Base UI AlertDialog is an alertdialog, so it now anchors on the
confirmation text and accepts either role.

agent_cost_view had no test at all; it gets one covering the null render,
the dollar-prefixed values, the omitted rows, and a zero cost that must not
be mistaken for unset.

All 55 tests pass against the current antd components.

* refactor(ui): migrate agents to shadcn

Replaces antd and Tremor with shadcn (base-vega) primitives across the five
files the agents route exclusively owns. Markup only; no behaviour, data
fetching or route structure changes.

Modal becomes AlertDialog, with a plain destructive Button in the footer
rather than AlertDialogAction, because that action is AlertDialog.Close and
would dismiss the dialog before the delete request settles, losing the
in-flight state. Alert, Tag, Spin, Space, Collapse, Descriptions, Typography
and the antd icons map onto alert, badge, ui-loading-spinner, flex/grid
utilities, collapsible, a definition list, semantic headings and lucide.

The shadcn CLI emits alert.tsx importing cva from class-variance-authority,
which this project does not depend on; it uses the cva object syntax from
lib/cva.config. The generated file fails to typecheck, so the adapted copy
lives in components/shared instead, per the convention that ui/ stays
CLI-managed.

Colour comes from tokens throughout, so the info callout is now the neutral
card style rather than antd's blue, and nothing hardcodes a colour in the way
of a later theme change.

The 55 tests in the route pass unchanged from the previous commit. The visual
gate re-baselined agents and all 34 other routes stayed pixel-identical.
2026-07-23 16:35:56 -07:00
Tin Chi Lo
2b77e8c4db fix(ui): keep cache leakage time range picker inline at narrow widths
The card header used flex-wrap, so the date picker was the element that
gave way when the row ran out of room; at higher browser zoom it dropped
onto its own line under the description. Pin the picker with shrink-0 and
let the title/description block shrink instead (min-w-0), so the copy
wraps to a second line and the picker stays on the right. Below md the
header stacks, since a 300px input plus its nowrap label leaves nothing
usable beside it.
2026-07-23 15:38:42 -07:00
tin-berri
43e7b96b83
Merge pull request #33978 from BerriAI/litellm_cost_optimization_tools
Some checks are pending
CodSpeed Benchmarks / benchmarks (push) Waiting to run
UI Unit Tests / ui-unit-tests (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
feat(cost-optimization): add spend-by-tool and cache leakage views
2026-07-23 13:43:23 -07:00
Tin Chi Lo
090fd491d4 feat(cost-optimization): sortable cache leakage columns and clearer token column name
Makes the three metric columns on the cache leakage table sortable, each with a
sensible first-click direction: most uncached tokens and biggest potential
savings first, worst cache hit rate first. Repeat clicks toggle the direction.
Renames Uncached input to Uncached input tokens, since the column is a token
count
2026-07-23 13:22:03 -07:00
tin-berri
df3050f538
Merge pull request #34407 from BerriAI/litellm_mcp_v1_obo_deadcode_cleanup
refactor(mcp): delete unreachable v1 OBO handler and gate REST OAuth on v2 resolver
2026-07-23 13:06:47 -07:00
Tin Chi Lo
d6d52d95e5 feat(cost-optimization): add by-model view to cache leakage table with plain-language columns
Adds a By virtual key / By model toggle to the cache leakage table. The model
view aggregates the daily activity model breakdown and is scoped to Anthropic
(Claude) models, which support prompt caching. Renames the columns to plain
language: Uncached input, Cache hit rate, and Potential savings (replacing
Realized caching savings and Est. savings left), with a tooltip on Potential
savings that spells out how it is calculated
2026-07-23 12:24:26 -07:00
mubashir1osmani
b5bc3631e1
refactor(e2e): drop require_env, read os.environ where a cred is used (#34413)
require_env hard-failed a test (and, for the shared litellm-ops secret, drove
piling every provider credential into one blob) whenever an optional cred was
absent. Most call sites either read a value the test actually uses or just
gated on the runner's env for a key the gateway consumes.

Read os.environ directly where the test uses the value; drop the presence-only
gates so those cases run against the proxy instead of pre-failing on the
runner's environment. Removes the require_env helper from e2e_config.
2026-07-23 19:14:22 +00:00
Tin Chi Lo
531854db4e fix(ui): hold the landing until the role hydrates before deciding the redirect
AuthContext sets token and clears authLoading in one effect, then a
second token-keyed effect populates userRole, so there is a render where
the user is signed in but userRole is still the initial empty string. The
positive internalUserRoles check reads that interim role as non-internal,
which let the api-keys dashboard paint for a frame before the role
arrived and the keyless redirect ran.

Treat "signed in on the post-login landing with an unhydrated role" as a
resolving state that holds the loading screen, so the dashboard never
flashes. Every login=success token carries a required user_role claim, so
the role always hydrates within a tick and this cannot hang; it is scoped
to the landing, so ordinary dashboard visits are unaffected.
2026-07-23 10:53:59 -07:00
Tin Chi Lo
bd73ca8c64 feat(cost-optimization): add spend-by-tool and cache leakage views
Adds GET /v1/tool/spend returning per-tool and daily tool spend with a
deduplicated request total, and a cache leakage breakdown on the Prompt
Caching tab of the Cost Optimization page. Tool-spend rows are validated at
the boundary with pydantic, the endpoint is scoped to proxy admins, date
params are cast to timestamptz for real-Postgres query_raw, and the leakage
math treats litellm-normalized prompt_tokens as cache-inclusive
(uncached = max(0, prompt - cache_read - cache_creation)).
2026-07-23 10:47:28 -07:00
ryan-crabbe-berri
1b2a7ce518
feat(ui): rebuild Organization Settings on react-hook-form + zod with a dirty-field PATCH (#34324)
* feat(ui): rebuild Organization Settings on react-hook-form + zod with a dirty-field PATCH

Replaces the antd Settings form in organization_view.tsx with OrgSettingsForm,
the first consumer of the shared RHF + zod form kit. The form derives a minimal
payload from RHF dirty tracking via pickDirty and sends it to the typed
PATCH /v2/organization/{organization_id}, so untouched fields are omitted,
emptied widgets clear with null ([] for lists), and the old full-send builder
with its length > 0 clear-dropping guards is deleted.

Adds src/lib/forms/useZodForm.ts so every form gets the z.input/z.output
generics and zodResolver wiring from one place, and forwardRefs ui/textarea
so RHF can register it under React 18

* fix(ui): forwardRef InputGroupTextarea to match the forwardRef'd Textarea

* test(ui): pin that an mcp server edit preserves existing org toolsets

* docs(ui): explain the useZodForm generics

* chore(ui): re-prune eslint suppressions after rebase onto staging
2026-07-23 10:45:04 -07:00
Tin Chi Lo
6ca40e7dcd refactor(mcp): delete unreachable v1 OBO handler and gate REST OAuth on v2 resolver
The v2 credential resolver owns oauth2_token_exchange end to end: any server
with a token-exchange config maps to a non-None TokenExchangeConfig spec, and
that config is in _create_mcp_client's override-exclusion set, so a caller
x-mcp-* override cannot force it back to v1 either. The v1 handler
resolve_mcp_auth reached at spec is None was therefore dead for OBO, including
its warn-then-proceed-unauthenticated fall-through. Delete auth/token_exchange.py
and the exchange branch, dropping the subject_token parameter that only fed it.

Separately, the REST listing and call paths still ran the v1 per-user OAuth
lookup for servers the v2 resolver owns. Unlike the two protocol-path call
sites they gated on auth_type == oauth2 only, with no to_server_spec check, so a
migrated authorization_code server did a DB round-trip whose Authorization
header _resolve_v2_auth then discards. Add the same guard via
_is_v1_resolved_oauth2_server, shared by the per-server lookup and the prefetch
preflight.

Also collapses MCPOAuth2TokenCache.async_get_token's now single-caller
require_client_credentials_flow kwarg and removes the dead
_get_bulk_user_oauth_headers helper (zero callers).
2026-07-23 10:41:15 -07:00
yuneng-jiang
0a4333580f
refactor(ui): migrate request logs table onto the shared DataTable (#34343)
* refactor(ui): migrate request logs table onto the shared DataTable

Moves the Request Logs tab off the local view_logs/table.tsx clone and onto the
shared DataTable in server sort, pagination, and filter mode. The container is
split into RequestLogsPanel (data owner: the spend-logs query, the session dedup
and composition pipeline, and the detail drawer), a thin RequestLogsTable, and
RequestLogsTableColumns. The clone itself stays for now because TopModelView and
TopKeyView still consume it

The advanced filter bar moves into the shared DataTableFilterDrawer, so filters
commit on Apply and render as removable chips. That makes the per-keystroke
debounce in the query hook redundant, and the hook now takes ColumnFiltersState,
PaginationState, and SortingState directly instead of carrying its own filter
shape. Reset still restores the default 24 hour window alongside the filters

Adds shared/PaginatedSearchSelect, a Base UI combobox with server-side search and
infinite scroll, and uses it for the Key Alias and Model filters. That retires the
three logs-only antd pickers (PaginatedKeyAliasSelect, PaginatedModelSelect,
FilterTeamDropdown) and the FilterComponent molecule they plugged into. The shared
TeamDropdown is deliberately untouched: six other surfaces still render it, five of
them as a bare child of an antd Form.Item that injects value/onChange implicitly

* test(ui): pin team-scoped key alias filtering in the logs filter drawer

The Key Alias filter narrows its options to the team selected in the same
drawer, a cross-filter dependency carried over from the antd picker it
replaced. Nothing covered it: the live QA pass explicitly did not exercise
it either, so it was the one behaviour in this migration that could regress
silently

Asserts the selected team id reaches useInfiniteKeyAliases, that the lookup
stays unscoped when no team is picked, and that the scope does not leak into
the Model lookup, which shares the same combobox but takes no team
2026-07-23 17:32:08 +00:00
tin-berri
3c2403a562
Merge pull request #33787 from BerriAI/litellm_mcp_chat_completion_linear_oauth
test(e2e): drive a real Linear OAuth MCP through chat completions under both ingress headers
2026-07-23 10:21:20 -07:00
yuneng-jiang
fd494d2fb2
refactor(ui): migrate models and endpoints table onto the shared DataTable (#34363)
* refactor(ui): migrate models and endpoints table onto the shared DataTable

Rebuilds the All Models table on the shared DataTable, following the 2a
treatment from the Models + Endpoints design: one card holding search, the
Team and View selectors, refresh, columns and filters, with the active
filters on a chip row and the pagination footer at the bottom.

Retires the last hand-rolled tremor renderer (all_models_table.tsx) and the
antd/tremor column defs in molecules/models/columns.tsx, replacing them with
a thin AllModelsTable consumer plus AllModelsTableColumns built from the
shared cell library.

Behavior is preserved end to end. The server sort field mapping now lives
next to the column ids so the two cannot drift. Status keeps its column and
its sort, hidden by default behind the Columns menu because the design shows
nine columns. Access groups collapse into a "+N more" tooltip instead of a
per-row expand toggle, and the full reset moves into the filter drawer
footer where the design puts it.

Adds the shadcn hover-card primitive (Base UI PreviewCard in the base-vega
style) for the model information hover, which needs an interactive surface a
tooltip cannot provide.

* fix(ui): stop the models tab re-querying on mount

The mount-time effect fires the debounced search with the initial empty
value, and its callback rebuilt the pagination object unconditionally. That
produced a second render (and a second query) roughly 300ms after mount with
no user input, which on a slow CI machine swapped the table's row nodes
mid-interaction and made a click land on a detached node.

resetToFirstPage now returns the existing state when already on the first
page, so React bails out instead of re-rendering. Pinned with a test that
asserts no additional query after the debounce settles; it fails without the
fix.
2026-07-23 10:20:16 -07:00
yuneng-jiang
bb388b2566
refactor(ui): migrate workflow runs to shadcn (#34370)
* test(ui): characterise the workflow runs detail drawer before migrating it

Pins the drawer's behaviour against the current antd implementation: the
metadata fields it surfaces, the timeline ordered by sequence number, the
empty-events copy, the messages section staying collapsed until opened, the
in-drawer refresh refetching, and the close control dismissing it.

Every assertion is role/text based so the same file can stay green once the
component moves off antd, without being edited.

* refactor(ui): migrate workflow runs to shadcn

Replaces the antd Drawer, Collapse, Button, Spin, Tooltip and Empty on the
Workflow Runs page with the installed Base UI primitives (Sheet, Collapsible,
Button, UiLoadingSpinner, Tooltip) and lucide icons, and moves the page's
hardcoded hex colours, fonts and geometry onto design tokens and utility
classes so the page can be themed. The only inline styles left are the gantt
bars' computed left/width, which are runtime values.

Behaviour is unchanged: the drawer's characterisation tests were written
against the antd version in the previous commit and pass here without being
edited.

Retires the file's now-unused antd no-restricted-imports suppression.
2026-07-23 10:19:04 -07:00
Tin Chi Lo
8929e09f49 fix(ui): gate the keyless connect redirect on internal-user roles
isAdminRole compares against a list that mixes raw and formatted role
strings: it holds raw org_admin but not the "Org Admin" that
formatUserRole produces, and AuthContext stores the formatted form. A
keyless org admin therefore read as a non-admin and was redirected to
the connect page.

Gate positively on internalUserRoles instead, which carries both
representations, so the redirect targets the persona it is meant for and
any role that is not unambiguously an internal user is left on the
dashboard. The shared admin list is left alone: completing it would
change org-admin access across every isAdminRole caller, which is a
roles-policy decision of its own.
2026-07-23 09:56:01 -07:00
tin-berri
712136d10d
Merge pull request #34225 from BerriAI/litellm_lit4658_oauth_discovery_logging
fix(mcp): log actionable OAuth discovery failures for misconfigured server urls
2026-07-23 09:51:31 -07:00
tin-berri
55b0046089
Merge branch 'litellm_internal_staging' into litellm_lit4658_oauth_discovery_logging 2026-07-23 09:39:11 -07:00
tin-berri
6bdb7918ea
Merge pull request #33190 from BerriAI/litellm_lit3637_session_admission
feat(mcp): admit gateway DCR session bearers at the aggregate /mcp scope
2026-07-23 09:33:59 -07:00
ryan-crabbe-berri
2bfd50ed37
test(e2e): cover key budget_duration resets on personal, team, and team-member keys (#33896) 2026-07-23 07:18:35 -07:00
Mateo Wang
3bba3633c7
Merge pull request #34338 from BerriAI/litellm_lit_4313_sagemaker_chat_streaming_ttft
fix(sagemaker): forward stream events as they arrive to cut TTFT
2026-07-23 00:37:39 -07:00
Tin Chi Lo
ffa0dffbc6 refactor(mcp): trim redundant comments and dedupe admission-arm tests
Compress the security rationale in the gateway-session admission path of
user_api_key_auth_mcp.py, keeping the load-bearing "why" and dropping the
restatement, and remove a garbled dead comment in get_allowed_tools_for_server

In the tests, hoist the duplicated _team / _admitted_subject fixtures to
module-level factories and parametrize the four fail-closed session-bearer
variants into one case. No behavior change; the 294 tests in the file still pass
2026-07-23 00:24:28 -07:00
Tin Chi Lo
a78130461f feat(mcp): gateway DCR session admission at the aggregate /mcp endpoint (LIT-3637)
Admits a keyless SSO user (no virtual key) at the aggregate /mcp endpoint from a gateway DCR
session bearer, resolving team/org/SCIM/budget authorization fresh on every call.

- Aggregate DCR front door: stateless /register (sealed llm_dcrc_ client ids), SSO-backed
  /authorize + /authorize/complete, and /token minting identity-only session tokens with PKCE,
  single-use codes/flows, and rotating refresh tokens.
- Admission: a session-shaped Authorization at the aggregate scope opens via _admit_gateway_session,
  reloads the live user, and runs the centralized policy gate; failures return the RFC 9728
  invalid_token challenge. Gated on the un-forgeable, server-only mcp_admitted_user_subject marker,
  so virtual-key and JWT auth are unchanged.
- Authorization model: an admitted subject is resolved as one plain UserAPIKeyAuth per grant source
  (its own grants, plus each team it is a live roster member of), each answered by the SAME resolver
  virtual keys use, then unioned. That branch is the FIRST statement of BOTH public resolvers, so no
  single-credential prelude runs for it and a fault in a lookup it never uses cannot deny its grants. A source team counts only while it is a live grantor: roster membership, not
  blocked, and neither the team nor its owning org over budget (enforced through the SAME
  _team_max_budget_check / _organization_max_budget_check owners common_checks uses for keys).
  Each team source carries that team's own org, so the existing org
  ceiling caps it; for a keyless source the org list only ever intersects (a ceiling must not become
  a grant) and an unresolvable ceiling denies rather than silently uncapping, on both the server and
  tool axes. _roster_team_object is the single owner of "which teams count": a team whose roster no
  longer lists the user neither grants servers nor throttles, in one place.
- Rate limits: the subject is bounded by its user rpm/tpm AND by the per-server mcp_rpm_limit of
  the team a call is ATTRIBUTED to — the same single source billing charges, from the same owner. A key charges its one pinned team's bucket; a keyless
  subject has no team_id, so admission stamps each granting team's limit map onto the auth
  (server-only field, stripped from validated input like the marker) and the limiter emits that
  team's mcp_per_team descriptor. Charging every granting team instead would let one cross-team user
  drain several teams' SHARED buckets on a single call and block their other members; and a server
  the user's OWN grant reaches charges no team bucket at all, because no team provided it. Per-KEY
  MCP limits do not apply because there is no key.
- Wrapper channels: the manager-level union treats the admitted subject by the same grant model.
  The admin-role short-circuit and the absolute no_mcp_servers early-return are key-credential
  rules and never apply to it (a session bearer is a third-party client credential, not the
  dashboard, and the subject's opt-out silences only its own source). Operator-open channels
  (allow_all_keys, the user's own BYOM submissions) are owned by one operator_open_server_ids
  helper that BOTH the server union and the admitted tool resolution consult (suppress-BYOM-when-
  explicitly-scoped is a key-credential rule and never applies to the subject, whose user row
  carries the DB-default empty mcp_servers), so an open-channel
  server is default-open for tools instead of listable but uninvokable.
- Redirect URIs: one owner, validate_redirect_uri_shape, decides redirect-URI hygiene (bad scheme,
  fragment, missing host, userinfo, backslash host) and resolves allowlisted native callbacks, shared
  by DCR registration and the OAuth endpoints. Registration keeps a deliberately wider trust policy
  than validate_trusted_redirect_uri: public dynamic registration accepts any https client, and its
  controls are mandatory S256 PKCE plus the consent screen.
- Egress leak-defense: a gateway admission credential (session bearer / bridge envelope) is scrubbed
  from EVERY egress header context, anchored to the credential shape, so it can never be forwarded
  upstream and replayed.
- Single-use guard: auth-code, refresh and connect-flow claims resolve the proxy's cross-worker redis
  cache themselves rather than trusting the cache passed in, and fail CLOSED on a Redis fault instead
  of falling back to a per-worker count that a captured id could replay through another worker.
- Sign-in return_to: one shared, never-raising helper persists a safe return_to for every sign-in
  branch (SSO/Okta/generic and username/password), and every branch RESUMES through the same
  _sso_return_to_redirect the SSO callback uses, so however a deployment signs in the stored value
  is honored identically (same-origin path directly; control_plane_url via the one-time login-code
  handoff). A stale cookie is ignored rather than failing a completed sign-in.

- Budgets, both halves: ENFORCEMENT (an already over-budget team or its owning org stops being a
  grantor, in the source gate) and ACCOUNTING (a team-derived tool call is billed to the granting
  team and ITS org, so that budget accumulates and the right organization is charged). A server the
  user's own grant reaches bills the user; when several teams grant one server the pick is the
  lowest team_id, stable and auditable. Billing rides a COPY, so authorization still sees the full
  union, and it is inert when the target server cannot be resolved from the tool name.

Deferred (tracked): client-selected server scoping of the session token (LIT-4680).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 00:24:28 -07:00
Yuneng Jiang
a928db1fab
fix(ui): keep the memory detail sheet inside the viewport on narrow screens
The migrated sheet asked for a flat 720px width while its only max-width
came from the primitive's sm:-scoped rule, so below the sm breakpoint no
cap applied at all: on a 375px viewport the sheet rendered 720px wide with
its left edge at -305px, and because it is position:fixed there was no
scroll to reach the hidden content. The primitive's own w-3/4 default did
not have this problem; the fixed pixel width is what removed the guard.

Caps the width to the viewport at every breakpoint and only asks for
720px from sm up. Verified in a browser at 375px, 700px and 1280px: the
sheet is now 375, 700 and 720 wide respectively, always at left 0.
2026-07-22 23:40:19 -07:00
Mateo Wang
86eee8bd05
Merge pull request #34319 from BerriAI/litellm_anthropic_output_format_remaining_keywords
fix(anthropic): strip all remaining output_format schema keywords rejected by Anthropic
2026-07-22 23:29:37 -07:00
Tin Chi Lo
3cdd6ab9a1 test(e2e): drive a real Linear OAuth MCP through chat completions under both ingress headers 2026-07-22 23:29:10 -07:00
yuneng-jiang
f7842cdeb7
fix(docker): bake non_root prisma engines at /opt/prisma so migrations run offline for any uid (#34325)
* fix(docker): bake non_root prisma engines at /opt/prisma so migrations run offline for any uid

The non_root image baked the prisma CLI and engines under /app/.cache and used
the CLI's default (library) engine mode. Prisma stopped baking the library
engine, so `prisma migrate deploy` fell back to downloading it at startup,
which needs network egress and a writable cache. Under an arbitrary non-root
uid (OpenShift restricted-v2), an air-gapped network, or a readOnlyRootFilesystem,
that download fails and the proxy starts on an empty schema while every DB
endpoint returns 500. The migration entrypoint exits 0 on that failure, so a
default-uid `docker run` with network never surfaced it

Bake to /opt/prisma, a fixed world-readable path no cache mount shadows, and
pin PRISMA_CLI_PATH plus PRISMA_CLI_QUERY_ENGINE_TYPE=binary so the baked binary
engine is used directly, matching Dockerfile and Dockerfile.database. A
build-time guard asserts the binary query engine is present, so a future prisma
change that stops baking it fails the image build instead of silently degrading
migrations

Adds docker/test_offline_migration.sh, run from image-scan, which migrates a
fresh Postgres with no egress as a non-root uid and asserts the schema was
created, the case a default-uid `docker run` with network cannot catch

* test(docker): move the offline migration check into a gated pytest and stop pinning XDG_CACHE_HOME at the read-only bake

The offline migration check lived in docker/ as a shell script. It now lives in
tests/proxy_migration_tests/ as a pytest gated on LITELLM_IMAGE, matching the
sibling schema-migration test gated on DATABASE_URL, and image-scan invokes it
with pytest instead of bash. It also asserts the migration entrypoint's exit
code alongside the table count, so a crash or a container-startup failure fails
loudly rather than only surfacing as a low table count

Runtime XDG_CACHE_HOME pointed at /opt/prisma/.cache, which is baked a+rX with
no write, so any XDG-aware library writing a cache at runtime would be denied
for every uid. Leave it unset so it falls back to $HOME/.cache (/app/.cache,
created here and owned by the runtime uid), matching Dockerfile and
Dockerfile.database which never pin XDG at runtime. A second test guards against
a future edit pointing a cache or home var back at the read-only bake
2026-07-22 23:03:39 -07:00
Yuneng Jiang
5a3ab862ee
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/migrate-page-memory-9b2c09
# Conflicts:
#	ui/litellm-dashboard/eslint-suppressions.json
2026-07-22 22:52:48 -07:00
Yuneng Jiang
d2f872d04f
refactor(ui): migrate memory page to shadcn
Replaces the antd Drawer with ui/sheet and the antd Button, Typography and
Space usage with ui/button plus token utilities, and swaps the @ant-design
PlusOutlined icon for lucide's Plus. Toasts now go through the shared
MessageManager so the route no longer imports antd directly.

The route's tests were written against the antd components in the previous
commit and are unchanged here, so they pass on both implementations.

MemoryEditModal is left alone because it is built on antd Form; the table
already sits on the shared DataTable.
2026-07-22 22:52:15 -07:00
Yuneng Jiang
9dda5d882c
test(ui): characterise the memory page's drawer and header before migrating
Adds role/text-based coverage for MemoryDetailDrawer (which had none) and
extends MemoryView's test past the mocked table to the header, the create
modal trigger and the detail drawer round trip. Both are green against the
current antd components, so they act as an unedited regression net for the
shadcn migration that follows.
2026-07-22 22:23:45 -07:00