Commit graph

41415 commits

Author SHA1 Message Date
Yucheng Zhu
198c7fe468 refactor(proxy/batches): make managed-file resolution purely additive, fall back like base
Restore the original PR's behavior on every path that did not already
resolve: no database, a lookup error, a missing managed-file row, or a
row without a storage_url all fall back to dispatching the original id,
which the managed-files deployment hook still maps. This drops the 404
and 503 fail-closed responses I had added, which were the only behaviors
that diverged from litellm_internal_staging.

The change is now strictly additive: when a managed-file row with a
storage_url exists, the unified batch branch substitutes it so providers
like Vertex receive a real gs:// path instead of the opaque token; every
other path behaves exactly as before. Verified live that non-managed,
managed-owner, multi-model load-balanced, and missing-row requests are
byte-identical to base
2026-07-24 17:53:24 -07:00
Yucheng Zhu
891c6e4123 fix(proxy/batches): fail closed with 503 when the managed-file lookup errors
A lookup exception previously fell back to dispatching the unresolved
unified token, which defeats the fail-closed guarantee: the token still
reaches the provider and can hit the same publishers-segment IndexError
the resolution prevents. Treat a lookup error like the missing-row case
and fail closed, but with a retryable 503 since the condition is
transient. No-database and no-storage_url rows still fall back unchanged
2026-07-24 17:06:46 -07:00
Yucheng Zhu
18dd984902 refactor(proxy/batches): scope managed-file handling to resolution, drop ownership check
Narrow this PR to its one problem: resolving a managed unified input_file_id
to its backend storage_url so provider batch handlers (Vertex parses a
publishers/ segment) receive a real location instead of the opaque token,
and failing closed with a 404 when the token has no backing row so it is
never dispatched into the provider crash.

Remove the cross-tenant ownership check (can_access_resource) added earlier.
Batch-create had no ownership enforcement before this PR, and the gap spans
every managed-file call type, so it belongs in the enterprise managed-files
pre-call hook (its acreate_batch branch) where files, batches and
fine-tuning are covered uniformly, not partially in this one endpoint. Filed
as a follow-up. This also removes the load-balanced-path ownership
inconsistency the bots flagged, since there is no ownership branch to skip.

Drop the inline comments flagged against the no-comments rule; behavior is
documented in the helper docstring and the test docstrings
2026-07-24 16:45:04 -07:00
Yucheng Zhu
df7fa74c93 fix(proxy/batches): do not divert unified files off the load-balanced branch
Excluding unified ids from the load-balanced branch (and not
unified_file_id) regressed a path that works on the base revision: a
multi-model managed file dispatched with an explicit router model under
load balancing was routed into the unified branch, which raises a 400
for anything other than exactly one target model. Verified live against
base (200, managed-files deployment hook remaps the unified id per
model) versus the guarded branch (400 Expected 1 model, got 2).

Restore the original three-condition load-balanced branch so that path
keeps working unchanged. Unified-file storage_url resolution and the
ownership 404 still apply on the non-load-balanced unified branch, which
is the common managed-batch flow; the load-balanced managed path retains
its existing behavior and its pre-existing enterprise-hook ownership gap,
unchanged from base
2026-07-24 15:44:05 -07:00
Yucheng Zhu
8e9e1254bf fix(proxy/batches): keep unified resolution in its own branch and fail closed on missing row
Cursor flagged that hoisting the storage_url substitution above the
load-balanced dispatch branch broke two things on that path: the
model_file_id_mapping deployment filter keys on the original unified id,
and the response returned the internal storage_url instead of the
unified id. Move the resolution back inside the unified branch and
exclude unified ids from the load-balanced branch so a managed file
always takes the resolving path (which restores input_file_id and the
unified_file_id hidden param on the response), and a load-balanced batch
keeps the original id for deployment filtering.

Also fail closed with a 404 when a unified id has no managed-file row
while a database is present: the id cannot be ownership-verified, and
dispatching it would both bypass the gate and hit the Vertex
publishers-segment IndexError. Owned rows without a storage_url (legacy)
still dispatch the original id
2026-07-24 01:31:51 -07:00
Yucheng Zhu
f2bb07dffc test(proxy/batches): default harness prisma_client to None
The batch routing harness left proxy_server.prisma_client at its module
global, which a sibling test in the same shard can leave as a MagicMock.
The unified-file rows that do not opt into managed-file resolution then
entered the resolver and awaited a non-awaitable mock, surfacing as a
503. Patch prisma_client to None by default so those rows stay a no-op;
resolution tests still override it explicitly
2026-07-24 01:15:02 -07:00
Yucheng Zhu
0a33ea40c6 fix(proxy/batches): fail closed when the managed file ownership lookup errors
A lookup exception previously fell back to dispatching the original
unified id with the ownership gate unexecuted; the managed-files
deployment hook maps unified ids from cache without re-checking
ownership, so a database outage let a caller dispatch another tenant's
file. Raise a clear 503 instead and lock the behavior with a regression
test. No-database and no-row cases still fall back unchanged
2026-07-23 23:59:36 -07:00
Yucheng Zhu
f8fc629286 fix(proxy/batches): enforce ownership and correct lookup key when resolving managed input_file_id
The adopted resolution queried LiteLLM_ManagedFileTable with the decoded
litellm_proxy string, but the unified_file_id column stores the raw base64
file id (see schema.prisma and the enterprise managed-files hook), so the
lookup never matched in production and silently fell back to the opaque id.
Query with the raw id instead and lock the key with a regression test.

Move the resolution above the dispatch branches so the load-balanced router
path receives the resolved storage_url too, enforce managed-file ownership
with the same can_access_resource semantics the files retrieve and download
endpoints use (404 on denial), and downgrade database failures to a logged
fallback instead of aborting batch creation. Unresolved ids still dispatch
unchanged because the managed-files deployment hook can map them via
model_file_id_mapping
2026-07-23 23:50:22 -07:00
htourinho-clgx
411b525a83 fix(proxy/batches): null-guard await on find_first for sync MagicMock test harnesses 2026-07-23 23:32:08 -07:00
htourinho-clgx
0e9f4ca09c fix: resolve unified_file_id to real storage_url before dispatching batch create
litellm.create_batch() against a Vertex AI-backed model crashes with an
opaque error when the input file was uploaded as a LiteLLM-managed
'unified file' (multi-model file upload). The base64-encoded
unified_file_id token is a LiteLLM-internal identifier, not a real
provider-side file reference, but the batches_endpoints create_batch
handler forwards it unchanged to llm_router.acreate_batch() /
litellm.acreate_batch() for the unified_file_id branch. Provider-specific
code that expects a real file location (e.g. Vertex AI's batch
transformation, which parses a 'publishers/' segment out of the GCS URI)
then fails on the opaque token.

Resolve the unified_file_id to its real backend location
(LiteLLM_ManagedFileTable.storage_url) before dispatch, mirroring the
same lookup already used by the files retrieve/download endpoints for
managed files. Falls back to the previous (unchanged) behavior if no
managed-file record exists.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-23 23:32:08 -07:00
tin-berri
2798f39f5a
Merge pull request #34434 from BerriAI/litellm_lit4748_autorouter_logs
feat(ui): show in the log drawer and session sidebar when an auto-router served a request
2026-07-23 22:09:34 -07:00
tin-berri
c93c3f7582
Merge pull request #34439 from BerriAI/litellm_cache_leakage_header_layout
fix(ui): keep cache leakage time range picker inline at narrow widths
2026-07-23 21:53:42 -07:00
Tin Chi Lo
42aba4f32a feat(ui): show in the log drawer and session sidebar when an auto-router served a request
The dashboard already receives the requested model name as model_group on
every spend-log row, but LogEntry dropped the field, so nothing distinguished
an auto-routed request from a direct one.

Surface it precisely rather than by comparing requested against resolved:
model_group differs from model for plain aliases and wildcard deployments
too, so a bare mismatch tags almost every row and identifies nothing. The
indication is driven instead by which deployments are auto-routers, resolved
from every page of /v2/model/info and shared through context.

The request drawer header names the router in a badge next to the provider;
the session sidebar swaps the entry's leading icon. Rows that no auto-router
served render exactly as before.
2026-07-23 20:43:54 -07:00
yuneng-jiang
64aad5877a
fix(proxy): restore atomic user upsert when adding team members (#34457)
* fix(proxy): restore atomic user upsert when adding team members

Parallel /team/new calls naming the same not-yet-existing member were
returning 500 "Unique constraint failed on the fields: (`user_id`)".

The upsert in add_new_member passed an empty update branch. Prisma only
compiles an upsert down to a single INSERT ... ON CONFLICT when that branch
writes something; with an empty one it emits SELECT-then-INSERT instead, so
concurrent requests all read "no such user" and all insert. Postgres
statement logs confirm it: the empty form logs BEGIN/SELECT/INSERT/COMMIT,
the non-empty form logs INSERT ... ON CONFLICT ("user_id") DO UPDATE SET.

Re-state user_id in the update branch as a no-op so the native upsert path
comes back. The teams append stays in the filtered update below it, so an
already-existing member still cannot pick up a duplicate team id.

tests/test_team.py::test_team_new failed 9 of 15 runs against a live proxy
before this and 0 of 15 after. The existing unit test asserted only that
upsert had been called on a mock, so it passed either way; it now pins the
shape of both branches and fails when the update branch goes back to empty.

* test: point the live codex tests at gpt-5.3-codex

OpenAI deprecated gpt-5.2-codex, so test_openai_codex and
test_openai_codex_stream started failing against the live API with
model_not_found. gpt-5.3-codex is the current codex model; both tests pass
on it. The remaining gpt-5.2-codex references in the suite are mocked
transformation tests and are unaffected.

* test(e2e): update models page specs for the shared DataTable

The DataTable migration in #34363 changed three things the models page
specs were pinned to, and five tests went red.

Row click no longer opens the detail view; the Model ID cell owns that
now, so both specs click its `model-id-<id>` test id instead of the row.
The search box placeholder switched from an ASCII "..." to a real
ellipsis, so the specs use getByPlaceholder with a substring instead of
an exact attribute match that punctuation can break again. The results
count moved from `models-results-count` ("Showing 1 - 50 of 137 results")
to the shared pagination's `pagination-range` ("Showing 1-50 of 137").

The Team-BYOK test also filtered rows on the team alias, which the Team
ID column has never rendered in either the old or the new table; it
filters on the team id now, which is what the column actually shows and
what the assertion's own comment intends.

Verified against a local proxy serving a fresh build with the seeded
e2e postgres and mock upstream: all five failing tests pass, and the
full suite is 82 passed / 4 skipped at CI parity (workers=1).
2026-07-24 02:22:34 +00:00
Mateo Wang
15af874fcf
Merge pull request #34046 from BerriAI/litellm_openai_cache_write_tokens
fix(cost_tracking): map OpenAI cache_write_tokens for prompt cache creation billing
2026-07-23 19:19:46 -07:00
mateo-berri
d3f5c6dbf6 fix(cost_calculator): sum mirrored cache token fields once in combine_usage_objects
combine_usage_objects iterates prompt_tokens_details model_fields and sums each;
with cache_write_tokens and cache_creation_tokens now mirroring each other via
__setattr__, the pair was summed twice, doubling cache creation counts for
Anthropic batch cost calc, mid-stream fallback usage merges, and realtime usage.
Collapse the mirrored pair to one representative before summing.
2026-07-23 19:07:09 -07:00
Krrish Dholakia
eaf61eb34b test(cost_tracking): cover OpenAI Responses API cache cost breakdown itemization (#34309)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:09 -07:00
Krrish Dholakia
5e1b08b355 fix(spend_tracking): populate cache_creation_input_tokens for Responses API logs
On the /v1/responses path the response usage is not chat-Usage-shaped, so
additional_usage_values could not derive cache tokens from response_obj.usage
and the Admin UI Logs cache-creation token row stayed empty. Fall back to the
normalized standard_logging usage_object's prompt_tokens_details for both the
cache-read and cache-creation counts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:09 -07:00
Krrish Dholakia
2f502a1bfc fix(cost_tracking): map cache_write_tokens on Responses API usage path
The Responses API (/v1/responses) usage transform rebuilt prompt token
details and dropped OpenAI's input_tokens_details.cache_write_tokens, so
gpt-5.6 cache-creation tokens were never logged or billed via that route.
Map it in the transform, and make PromptTokensDetailsWrapper keep
cache_write_tokens and cache_creation_tokens in sync on assignment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:08 -07:00
Krrish Dholakia
e6ec153243 fix(cost_tracking): map OpenAI cache_write_tokens for prompt cache creation billing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 19:07:08 -07:00
devin-ai-integration[bot]
c255f53bfb
ci: only run CodSpeed on backend changes (#34345)
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-23 16:58:55 -07:00
ryan-crabbe-berri
07f7fc224e
fix(proxy): reject failed atomic budget reservations under fail_closed_budget_enforcement (#34429)
* fix(proxy): reject request when budget reservation write fails under fail_closed_budget_enforcement

With general_settings.fail_closed_budget_enforcement set to true, the read-time
spend check already returns 503 when spend cannot be verified, but the atomic
pre-call reservation still failed open: reserve_budget_for_request swallowed
_CounterReservationUnavailable per counter and degraded to read-time-only
enforcement, so concurrent requests could all pass the same under-budget read
during a Redis outage and overspend past the configured budget.

Now the strict flag is threaded into reserve_budget_for_request and a failed
reservation write raises 503, releasing any counters that already reserved.
Default behavior with the flag absent or false is unchanged.

Fixes #33923

* fix(proxy): pass 503 budget-enforcement detail as plain string
2026-07-23 23:57:37 +00:00
ryan-crabbe-berri
a507394841
fix(ui): find logs by request id across pages and dates (LIT-3981) (#31743)
* fix(spend): resolve spend logs by request_id across all dates (LIT-3981)

The /spend/logs/ui search only filtered the page already loaded, so a log id
copied from another page or from outside the active date window could not be
found. request_id is the primary key of LiteLLM_SpendLogs, so when it is
supplied on the internal UI route the mandatory date window is dropped and the
lookup resolves across all time. The date window stays required when no
request_id is given, and the public /spend/logs/v2 contract is unchanged.

A non-admin id lookup is gated by the same ownership check the detail endpoint
uses, so the relaxed window cannot be used to read another tenant's log by id

* fix(ui): send the logs request_id search to the server (LIT-3981)

The "Search by Request ID" box filtered only the rows already on the current
page, so an id from another page never matched. It now feeds the existing
server-side request_id filter via handleFilterChange, which debounces, resets
to page one, and rides the existing react-query key. The dead client-side
filter and its searchTerm state are removed; the session composition and dedup
logic is unchanged.

The box is now an exact request_id lookup, matching its label; the incidental
client-side model and user substring matching it used to do is dropped in
favor of the dedicated filters

* refactor(spend): model the request_id spend-log lookup as an explicit point lookup (LIT-3981)

The date-window relaxation for request_id lookups rode an apply_date_window flag threaded through the date validation and parsing. Model the two intents directly instead. A UI request_id query is a point lookup on the @id primary key that drops the time window and authorizes by row ownership; every other query, including the public /spend/logs/v2 route, takes the range-scan path that still requires a window

Because the ownership check fully authorizes the single row, the general user/team scoping is now skipped for id lookups rather than layered on top redundantly. The confusing `is_v2 or request_id is None` guard is gone, and moving the date requirement into the range-scan branch lets the type checker narrow the dates it parses

Behavior is preserved: the v2 contract still requires dates even when a request_id is supplied, and a non-owner is still rejected with 403. A regression test covers the non-admin owner id lookup, which resolves across all time and filters by the primary key alone
2026-07-23 16:40:19 -07:00
ryan-crabbe-berri
7aaaa055b7
feat(proxy): add overwrite_user_with_key_hash to stamp outgoing user param with key hash (#34417)
* feat(proxy): add overwrite_user_with_key_hash to stamp outgoing user param with key hash

Adds a litellm_settings flag that forces the outgoing user param to the
authenticated key's hashed token before the request is forwarded to the
provider. The value overrides any caller-supplied user, so providers see
a stable, tamper-proof identifier they can rate-limit or ban on, and the
hash matches user_api_key_hash in spend logs for easy mapping back to
the key owner. Off by default

* fix(proxy): hash non-sk credentials before stamping user param

UserAPIKeyAuth only hashes sk-prefixed keys and JWTs; custom-auth
credentials stay raw on api_key, so stamping them directly would forward
auth material to the provider. Pass through the two known hashed forms
(sha256 hex, hashed-jwt-*) and hash anything else

* refactor(proxy): stamp only standard virtual keys, skip jwt and custom auth

A hashed JWT rotates on every token re-issue so it is useless as a
stable ban id, and custom-auth credentials arrive raw on api_key.
Instead of hashing whatever we hold, the stamp now applies only when
api_key is the sha256 hex digest of a standard virtual key; other auth
methods are explicitly out of scope until the stamped identifier is
configurable

* fix(proxy): gate user stamping on server-set virtual key provenance

Shape alone cannot distinguish a key hash from a raw custom-auth
credential that happens to be 64 hex chars. Adds via_virtual_key, a
server-only marker on UserAPIKeyAuth following the
mcp_admitted_user_subject pattern: stripped from all validated input so
handlers and claims cannot forge it, set by post-construction assignment
only at the DB virtual-key auth return. Stamping now requires the marker
and the hash shape

* test(proxy): prove db auth path sets via_virtual_key marker

The stamping unit tests set the marker manually, so deleting the
assignment in _user_api_key_auth_builder would pass every existing test;
this exercises the real builder path with a mocked identity store and
fails if the marker is not set

* fix(proxy): stamp master-key requests with the master key alias

Master-key auth substitutes LITELLM_PROXY_MASTER_KEY_ALIAS for api_key
so the key and its hash never propagate; that made master-key traffic
bypass stamping and pass the caller-supplied user through. The master
path now sets via_virtual_key and the stamp gate accepts the alias
alongside the sha256 shape, so admin traffic gets the same tamper-proof
id that spend logs already record for it

* fix(proxy): restore via_virtual_key marker on key-cache hits

Cached PROXY_ADMIN auth objects early-return before the marked DB and
master-key returns, and cache serialization drops the exclude=True
marker, so cached admin traffic bypassed stamping. Key-cache entries are
written only after the proxy validated a virtual key or the master key,
so the cache-hit boundary restores the marker; the UI-login JWT fallback
constructs its token from a decrypted blob, not this cache, and stays
unmarked
2026-07-23 16:38:01 -07:00
ryan-crabbe-berri
e906a7e796
refactor(ui): extract shared tab-routing helpers and adopt them in Models + Endpoints (#34435)
* refactor(ui): extract shared tab-routing helpers

Every per-tab-routed page copy-pastes the same URL<->slug logic and the
same active-tab/redirect engine. Extract two reusable pieces:

- createTabRoutes(baseSegment, slugs) in utils/tabRoutes.ts returns
  { baseSegment, slugs, tabHref, slugFromPathname }, the trailing-slash
  href builder (via migratedHref) and the pathname->slug reader.
- useTabRouting({ routes, baseTabKey, visibleKeys, ready }) derives the
  active tab from the pathname, redirects an unknown/forbidden slug to
  base once ready, and returns an onTabChange navigator.

visibleKeys + ready exist so a role-gated page can pass its filtered tab
set and defer the redirect until permissions resolve, rather than
bouncing a user off a still-loading valid tab. Both are pure/unit-tested.
No page consumes them yet.

* refactor(ui): migrate Models + Endpoints onto the shared tab-routing helpers

Replace the page's hand-rolled tabRoutes.ts (base segment + slug tuple +
href builder + slugFromPathname) with createTabRoutes, keeping the
existing named exports as thin re-exports so callers and tests are
unchanged. The layout drops its local activeSlug/isKnownSlug/activeKey
derivation, its redirect useEffect and its router.push onChange in favor
of useTabRouting, passing the role-filtered visibleKeys and a ready flag
(!teamsLoading && !uiSettingsLoading) so the permission-gated redirect
behavior is preserved exactly. The antd tab bar, role-gated tab set, the
refresh button and the ?model=/?team= drill-in overlay are untouched; the
file's pre-existing antd import is now recorded in the suppressions
baseline since editing it makes it a linted-as-changed file.

The existing models-and-endpoints layout.test.tsx and tabRoutes.test.ts
pass unchanged, which is the regression guarantee.
2026-07-23 16:36:25 -07:00
yuneng-jiang
07726b4f60
refactor(ui): migrate agents to shadcn (#34365)
* test(ui): make the agents route's tests markup-agnostic before migration

Rewrites the two assertions that were coupled to antd's DOM and adds the
missing characterisation test for agent_cost_view, so the suite describes
behaviour rather than antd markup and can stay untouched across the shadcn
migration.

The skill selection test reached the checkbox with a querySelector on
input[type=checkbox]; antd renders an input while Base UI renders a
span[role=checkbox], so it now queries by role and accessible name, which
both libraries derive from the wrapping label.

The delete confirmation test queried role=dialog; antd Modal is a dialog
while Base UI AlertDialog is an alertdialog, so it now anchors on the
confirmation text and accepts either role.

agent_cost_view had no test at all; it gets one covering the null render,
the dollar-prefixed values, the omitted rows, and a zero cost that must not
be mistaken for unset.

All 55 tests pass against the current antd components.

* refactor(ui): migrate agents to shadcn

Replaces antd and Tremor with shadcn (base-vega) primitives across the five
files the agents route exclusively owns. Markup only; no behaviour, data
fetching or route structure changes.

Modal becomes AlertDialog, with a plain destructive Button in the footer
rather than AlertDialogAction, because that action is AlertDialog.Close and
would dismiss the dialog before the delete request settles, losing the
in-flight state. Alert, Tag, Spin, Space, Collapse, Descriptions, Typography
and the antd icons map onto alert, badge, ui-loading-spinner, flex/grid
utilities, collapsible, a definition list, semantic headings and lucide.

The shadcn CLI emits alert.tsx importing cva from class-variance-authority,
which this project does not depend on; it uses the cva object syntax from
lib/cva.config. The generated file fails to typecheck, so the adapted copy
lives in components/shared instead, per the convention that ui/ stays
CLI-managed.

Colour comes from tokens throughout, so the info callout is now the neutral
card style rather than antd's blue, and nothing hardcodes a colour in the way
of a later theme change.

The 55 tests in the route pass unchanged from the previous commit. The visual
gate re-baselined agents and all 34 other routes stayed pixel-identical.
2026-07-23 16:35:56 -07:00
Tin Chi Lo
2b77e8c4db fix(ui): keep cache leakage time range picker inline at narrow widths
The card header used flex-wrap, so the date picker was the element that
gave way when the row ran out of room; at higher browser zoom it dropped
onto its own line under the description. Pin the picker with shrink-0 and
let the title/description block shrink instead (min-w-0), so the copy
wraps to a second line and the picker stays on the right. Below md the
header stacks, since a 300px input plus its nowrap label leaves nothing
usable beside it.
2026-07-23 15:38:42 -07:00
tin-berri
43e7b96b83
Merge pull request #33978 from BerriAI/litellm_cost_optimization_tools
Some checks are pending
CodSpeed Benchmarks / benchmarks (push) Waiting to run
UI Unit Tests / ui-unit-tests (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
feat(cost-optimization): add spend-by-tool and cache leakage views
2026-07-23 13:43:23 -07:00
Tin Chi Lo
090fd491d4 feat(cost-optimization): sortable cache leakage columns and clearer token column name
Makes the three metric columns on the cache leakage table sortable, each with a
sensible first-click direction: most uncached tokens and biggest potential
savings first, worst cache hit rate first. Repeat clicks toggle the direction.
Renames Uncached input to Uncached input tokens, since the column is a token
count
2026-07-23 13:22:03 -07:00
tin-berri
df3050f538
Merge pull request #34407 from BerriAI/litellm_mcp_v1_obo_deadcode_cleanup
refactor(mcp): delete unreachable v1 OBO handler and gate REST OAuth on v2 resolver
2026-07-23 13:06:47 -07:00
Tin Chi Lo
d6d52d95e5 feat(cost-optimization): add by-model view to cache leakage table with plain-language columns
Adds a By virtual key / By model toggle to the cache leakage table. The model
view aggregates the daily activity model breakdown and is scoped to Anthropic
(Claude) models, which support prompt caching. Renames the columns to plain
language: Uncached input, Cache hit rate, and Potential savings (replacing
Realized caching savings and Est. savings left), with a tooltip on Potential
savings that spells out how it is calculated
2026-07-23 12:24:26 -07:00
mubashir1osmani
b5bc3631e1
refactor(e2e): drop require_env, read os.environ where a cred is used (#34413)
require_env hard-failed a test (and, for the shared litellm-ops secret, drove
piling every provider credential into one blob) whenever an optional cred was
absent. Most call sites either read a value the test actually uses or just
gated on the runner's env for a key the gateway consumes.

Read os.environ directly where the test uses the value; drop the presence-only
gates so those cases run against the proxy instead of pre-failing on the
runner's environment. Removes the require_env helper from e2e_config.
2026-07-23 19:14:22 +00:00
Tin Chi Lo
bd73ca8c64 feat(cost-optimization): add spend-by-tool and cache leakage views
Adds GET /v1/tool/spend returning per-tool and daily tool spend with a
deduplicated request total, and a cache leakage breakdown on the Prompt
Caching tab of the Cost Optimization page. Tool-spend rows are validated at
the boundary with pydantic, the endpoint is scoped to proxy admins, date
params are cast to timestamptz for real-Postgres query_raw, and the leakage
math treats litellm-normalized prompt_tokens as cache-inclusive
(uncached = max(0, prompt - cache_read - cache_creation)).
2026-07-23 10:47:28 -07:00
ryan-crabbe-berri
1b2a7ce518
feat(ui): rebuild Organization Settings on react-hook-form + zod with a dirty-field PATCH (#34324)
* feat(ui): rebuild Organization Settings on react-hook-form + zod with a dirty-field PATCH

Replaces the antd Settings form in organization_view.tsx with OrgSettingsForm,
the first consumer of the shared RHF + zod form kit. The form derives a minimal
payload from RHF dirty tracking via pickDirty and sends it to the typed
PATCH /v2/organization/{organization_id}, so untouched fields are omitted,
emptied widgets clear with null ([] for lists), and the old full-send builder
with its length > 0 clear-dropping guards is deleted.

Adds src/lib/forms/useZodForm.ts so every form gets the z.input/z.output
generics and zodResolver wiring from one place, and forwardRefs ui/textarea
so RHF can register it under React 18

* fix(ui): forwardRef InputGroupTextarea to match the forwardRef'd Textarea

* test(ui): pin that an mcp server edit preserves existing org toolsets

* docs(ui): explain the useZodForm generics

* chore(ui): re-prune eslint suppressions after rebase onto staging
2026-07-23 10:45:04 -07:00
Tin Chi Lo
6ca40e7dcd refactor(mcp): delete unreachable v1 OBO handler and gate REST OAuth on v2 resolver
The v2 credential resolver owns oauth2_token_exchange end to end: any server
with a token-exchange config maps to a non-None TokenExchangeConfig spec, and
that config is in _create_mcp_client's override-exclusion set, so a caller
x-mcp-* override cannot force it back to v1 either. The v1 handler
resolve_mcp_auth reached at spec is None was therefore dead for OBO, including
its warn-then-proceed-unauthenticated fall-through. Delete auth/token_exchange.py
and the exchange branch, dropping the subject_token parameter that only fed it.

Separately, the REST listing and call paths still ran the v1 per-user OAuth
lookup for servers the v2 resolver owns. Unlike the two protocol-path call
sites they gated on auth_type == oauth2 only, with no to_server_spec check, so a
migrated authorization_code server did a DB round-trip whose Authorization
header _resolve_v2_auth then discards. Add the same guard via
_is_v1_resolved_oauth2_server, shared by the per-server lookup and the prefetch
preflight.

Also collapses MCPOAuth2TokenCache.async_get_token's now single-caller
require_client_credentials_flow kwarg and removes the dead
_get_bulk_user_oauth_headers helper (zero callers).
2026-07-23 10:41:15 -07:00
yuneng-jiang
0a4333580f
refactor(ui): migrate request logs table onto the shared DataTable (#34343)
* refactor(ui): migrate request logs table onto the shared DataTable

Moves the Request Logs tab off the local view_logs/table.tsx clone and onto the
shared DataTable in server sort, pagination, and filter mode. The container is
split into RequestLogsPanel (data owner: the spend-logs query, the session dedup
and composition pipeline, and the detail drawer), a thin RequestLogsTable, and
RequestLogsTableColumns. The clone itself stays for now because TopModelView and
TopKeyView still consume it

The advanced filter bar moves into the shared DataTableFilterDrawer, so filters
commit on Apply and render as removable chips. That makes the per-keystroke
debounce in the query hook redundant, and the hook now takes ColumnFiltersState,
PaginationState, and SortingState directly instead of carrying its own filter
shape. Reset still restores the default 24 hour window alongside the filters

Adds shared/PaginatedSearchSelect, a Base UI combobox with server-side search and
infinite scroll, and uses it for the Key Alias and Model filters. That retires the
three logs-only antd pickers (PaginatedKeyAliasSelect, PaginatedModelSelect,
FilterTeamDropdown) and the FilterComponent molecule they plugged into. The shared
TeamDropdown is deliberately untouched: six other surfaces still render it, five of
them as a bare child of an antd Form.Item that injects value/onChange implicitly

* test(ui): pin team-scoped key alias filtering in the logs filter drawer

The Key Alias filter narrows its options to the team selected in the same
drawer, a cross-filter dependency carried over from the antd picker it
replaced. Nothing covered it: the live QA pass explicitly did not exercise
it either, so it was the one behaviour in this migration that could regress
silently

Asserts the selected team id reaches useInfiniteKeyAliases, that the lookup
stays unscoped when no team is picked, and that the scope does not leak into
the Model lookup, which shares the same combobox but takes no team
2026-07-23 17:32:08 +00:00
tin-berri
3c2403a562
Merge pull request #33787 from BerriAI/litellm_mcp_chat_completion_linear_oauth
test(e2e): drive a real Linear OAuth MCP through chat completions under both ingress headers
2026-07-23 10:21:20 -07:00
yuneng-jiang
fd494d2fb2
refactor(ui): migrate models and endpoints table onto the shared DataTable (#34363)
* refactor(ui): migrate models and endpoints table onto the shared DataTable

Rebuilds the All Models table on the shared DataTable, following the 2a
treatment from the Models + Endpoints design: one card holding search, the
Team and View selectors, refresh, columns and filters, with the active
filters on a chip row and the pagination footer at the bottom.

Retires the last hand-rolled tremor renderer (all_models_table.tsx) and the
antd/tremor column defs in molecules/models/columns.tsx, replacing them with
a thin AllModelsTable consumer plus AllModelsTableColumns built from the
shared cell library.

Behavior is preserved end to end. The server sort field mapping now lives
next to the column ids so the two cannot drift. Status keeps its column and
its sort, hidden by default behind the Columns menu because the design shows
nine columns. Access groups collapse into a "+N more" tooltip instead of a
per-row expand toggle, and the full reset moves into the filter drawer
footer where the design puts it.

Adds the shadcn hover-card primitive (Base UI PreviewCard in the base-vega
style) for the model information hover, which needs an interactive surface a
tooltip cannot provide.

* fix(ui): stop the models tab re-querying on mount

The mount-time effect fires the debounced search with the initial empty
value, and its callback rebuilt the pagination object unconditionally. That
produced a second render (and a second query) roughly 300ms after mount with
no user input, which on a slow CI machine swapped the table's row nodes
mid-interaction and made a click land on a detached node.

resetToFirstPage now returns the existing state when already on the first
page, so React bails out instead of re-rendering. Pinned with a test that
asserts no additional query after the debounce settles; it fails without the
fix.
2026-07-23 10:20:16 -07:00
yuneng-jiang
bb388b2566
refactor(ui): migrate workflow runs to shadcn (#34370)
* test(ui): characterise the workflow runs detail drawer before migrating it

Pins the drawer's behaviour against the current antd implementation: the
metadata fields it surfaces, the timeline ordered by sequence number, the
empty-events copy, the messages section staying collapsed until opened, the
in-drawer refresh refetching, and the close control dismissing it.

Every assertion is role/text based so the same file can stay green once the
component moves off antd, without being edited.

* refactor(ui): migrate workflow runs to shadcn

Replaces the antd Drawer, Collapse, Button, Spin, Tooltip and Empty on the
Workflow Runs page with the installed Base UI primitives (Sheet, Collapsible,
Button, UiLoadingSpinner, Tooltip) and lucide icons, and moves the page's
hardcoded hex colours, fonts and geometry onto design tokens and utility
classes so the page can be themed. The only inline styles left are the gantt
bars' computed left/width, which are runtime values.

Behaviour is unchanged: the drawer's characterisation tests were written
against the antd version in the previous commit and pass here without being
edited.

Retires the file's now-unused antd no-restricted-imports suppression.
2026-07-23 10:19:04 -07:00
tin-berri
712136d10d
Merge pull request #34225 from BerriAI/litellm_lit4658_oauth_discovery_logging
fix(mcp): log actionable OAuth discovery failures for misconfigured server urls
2026-07-23 09:51:31 -07:00
tin-berri
55b0046089
Merge branch 'litellm_internal_staging' into litellm_lit4658_oauth_discovery_logging 2026-07-23 09:39:11 -07:00
tin-berri
6bdb7918ea
Merge pull request #33190 from BerriAI/litellm_lit3637_session_admission
feat(mcp): admit gateway DCR session bearers at the aggregate /mcp scope
2026-07-23 09:33:59 -07:00
ryan-crabbe-berri
2bfd50ed37
test(e2e): cover key budget_duration resets on personal, team, and team-member keys (#33896) 2026-07-23 07:18:35 -07:00
Mateo Wang
3bba3633c7
Merge pull request #34338 from BerriAI/litellm_lit_4313_sagemaker_chat_streaming_ttft
fix(sagemaker): forward stream events as they arrive to cut TTFT
2026-07-23 00:37:39 -07:00
Tin Chi Lo
ffa0dffbc6 refactor(mcp): trim redundant comments and dedupe admission-arm tests
Compress the security rationale in the gateway-session admission path of
user_api_key_auth_mcp.py, keeping the load-bearing "why" and dropping the
restatement, and remove a garbled dead comment in get_allowed_tools_for_server

In the tests, hoist the duplicated _team / _admitted_subject fixtures to
module-level factories and parametrize the four fail-closed session-bearer
variants into one case. No behavior change; the 294 tests in the file still pass
2026-07-23 00:24:28 -07:00
Tin Chi Lo
a78130461f feat(mcp): gateway DCR session admission at the aggregate /mcp endpoint (LIT-3637)
Admits a keyless SSO user (no virtual key) at the aggregate /mcp endpoint from a gateway DCR
session bearer, resolving team/org/SCIM/budget authorization fresh on every call.

- Aggregate DCR front door: stateless /register (sealed llm_dcrc_ client ids), SSO-backed
  /authorize + /authorize/complete, and /token minting identity-only session tokens with PKCE,
  single-use codes/flows, and rotating refresh tokens.
- Admission: a session-shaped Authorization at the aggregate scope opens via _admit_gateway_session,
  reloads the live user, and runs the centralized policy gate; failures return the RFC 9728
  invalid_token challenge. Gated on the un-forgeable, server-only mcp_admitted_user_subject marker,
  so virtual-key and JWT auth are unchanged.
- Authorization model: an admitted subject is resolved as one plain UserAPIKeyAuth per grant source
  (its own grants, plus each team it is a live roster member of), each answered by the SAME resolver
  virtual keys use, then unioned. That branch is the FIRST statement of BOTH public resolvers, so no
  single-credential prelude runs for it and a fault in a lookup it never uses cannot deny its grants. A source team counts only while it is a live grantor: roster membership, not
  blocked, and neither the team nor its owning org over budget (enforced through the SAME
  _team_max_budget_check / _organization_max_budget_check owners common_checks uses for keys).
  Each team source carries that team's own org, so the existing org
  ceiling caps it; for a keyless source the org list only ever intersects (a ceiling must not become
  a grant) and an unresolvable ceiling denies rather than silently uncapping, on both the server and
  tool axes. _roster_team_object is the single owner of "which teams count": a team whose roster no
  longer lists the user neither grants servers nor throttles, in one place.
- Rate limits: the subject is bounded by its user rpm/tpm AND by the per-server mcp_rpm_limit of
  the team a call is ATTRIBUTED to — the same single source billing charges, from the same owner. A key charges its one pinned team's bucket; a keyless
  subject has no team_id, so admission stamps each granting team's limit map onto the auth
  (server-only field, stripped from validated input like the marker) and the limiter emits that
  team's mcp_per_team descriptor. Charging every granting team instead would let one cross-team user
  drain several teams' SHARED buckets on a single call and block their other members; and a server
  the user's OWN grant reaches charges no team bucket at all, because no team provided it. Per-KEY
  MCP limits do not apply because there is no key.
- Wrapper channels: the manager-level union treats the admitted subject by the same grant model.
  The admin-role short-circuit and the absolute no_mcp_servers early-return are key-credential
  rules and never apply to it (a session bearer is a third-party client credential, not the
  dashboard, and the subject's opt-out silences only its own source). Operator-open channels
  (allow_all_keys, the user's own BYOM submissions) are owned by one operator_open_server_ids
  helper that BOTH the server union and the admitted tool resolution consult (suppress-BYOM-when-
  explicitly-scoped is a key-credential rule and never applies to the subject, whose user row
  carries the DB-default empty mcp_servers), so an open-channel
  server is default-open for tools instead of listable but uninvokable.
- Redirect URIs: one owner, validate_redirect_uri_shape, decides redirect-URI hygiene (bad scheme,
  fragment, missing host, userinfo, backslash host) and resolves allowlisted native callbacks, shared
  by DCR registration and the OAuth endpoints. Registration keeps a deliberately wider trust policy
  than validate_trusted_redirect_uri: public dynamic registration accepts any https client, and its
  controls are mandatory S256 PKCE plus the consent screen.
- Egress leak-defense: a gateway admission credential (session bearer / bridge envelope) is scrubbed
  from EVERY egress header context, anchored to the credential shape, so it can never be forwarded
  upstream and replayed.
- Single-use guard: auth-code, refresh and connect-flow claims resolve the proxy's cross-worker redis
  cache themselves rather than trusting the cache passed in, and fail CLOSED on a Redis fault instead
  of falling back to a per-worker count that a captured id could replay through another worker.
- Sign-in return_to: one shared, never-raising helper persists a safe return_to for every sign-in
  branch (SSO/Okta/generic and username/password), and every branch RESUMES through the same
  _sso_return_to_redirect the SSO callback uses, so however a deployment signs in the stored value
  is honored identically (same-origin path directly; control_plane_url via the one-time login-code
  handoff). A stale cookie is ignored rather than failing a completed sign-in.

- Budgets, both halves: ENFORCEMENT (an already over-budget team or its owning org stops being a
  grantor, in the source gate) and ACCOUNTING (a team-derived tool call is billed to the granting
  team and ITS org, so that budget accumulates and the right organization is charged). A server the
  user's own grant reaches bills the user; when several teams grant one server the pick is the
  lowest team_id, stable and auditable. Billing rides a COPY, so authorization still sees the full
  union, and it is inert when the target server cannot be resolved from the tool name.

Deferred (tracked): client-selected server scoping of the session token (LIT-4680).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 00:24:28 -07:00
Mateo Wang
86eee8bd05
Merge pull request #34319 from BerriAI/litellm_anthropic_output_format_remaining_keywords
fix(anthropic): strip all remaining output_format schema keywords rejected by Anthropic
2026-07-22 23:29:37 -07:00
Tin Chi Lo
3cdd6ab9a1 test(e2e): drive a real Linear OAuth MCP through chat completions under both ingress headers 2026-07-22 23:29:10 -07:00
yuneng-jiang
f7842cdeb7
fix(docker): bake non_root prisma engines at /opt/prisma so migrations run offline for any uid (#34325)
* fix(docker): bake non_root prisma engines at /opt/prisma so migrations run offline for any uid

The non_root image baked the prisma CLI and engines under /app/.cache and used
the CLI's default (library) engine mode. Prisma stopped baking the library
engine, so `prisma migrate deploy` fell back to downloading it at startup,
which needs network egress and a writable cache. Under an arbitrary non-root
uid (OpenShift restricted-v2), an air-gapped network, or a readOnlyRootFilesystem,
that download fails and the proxy starts on an empty schema while every DB
endpoint returns 500. The migration entrypoint exits 0 on that failure, so a
default-uid `docker run` with network never surfaced it

Bake to /opt/prisma, a fixed world-readable path no cache mount shadows, and
pin PRISMA_CLI_PATH plus PRISMA_CLI_QUERY_ENGINE_TYPE=binary so the baked binary
engine is used directly, matching Dockerfile and Dockerfile.database. A
build-time guard asserts the binary query engine is present, so a future prisma
change that stops baking it fails the image build instead of silently degrading
migrations

Adds docker/test_offline_migration.sh, run from image-scan, which migrates a
fresh Postgres with no egress as a non-root uid and asserts the schema was
created, the case a default-uid `docker run` with network cannot catch

* test(docker): move the offline migration check into a gated pytest and stop pinning XDG_CACHE_HOME at the read-only bake

The offline migration check lived in docker/ as a shell script. It now lives in
tests/proxy_migration_tests/ as a pytest gated on LITELLM_IMAGE, matching the
sibling schema-migration test gated on DATABASE_URL, and image-scan invokes it
with pytest instead of bash. It also asserts the migration entrypoint's exit
code alongside the table count, so a crash or a container-startup failure fails
loudly rather than only surfacing as a low table count

Runtime XDG_CACHE_HOME pointed at /opt/prisma/.cache, which is baked a+rX with
no write, so any XDG-aware library writing a cache at runtime would be denied
for every uid. Leave it unset so it falls back to $HOME/.cache (/app/.cache,
created here and owned by the runtime uid), matching Dockerfile and
Dockerfile.database which never pin XDG at runtime. A second test guards against
a future edit pointing a cache or home var back at the read-only bake
2026-07-22 23:03:39 -07:00
yuneng-jiang
baf85d7e38
refactor(ui): migrate search-tools info view to shadcn (#34323)
* test(ui): pin search-tools info view behavior before shadcn migration

Rewrite the markup-coupled copy-button assertions in SearchToolView to
role queries plus lucide icon-state, and add a role/text characterization
suite for SearchConnectionTest, which had none. Both are green against the
current antd/Tremor components so they can act as an unedited regression
net across the migration.

* refactor(ui): migrate search-tools info view to shadcn

Port the search-tools detail view and its two helpers off antd and Tremor
onto the installed shadcn primitives plus token utilities:

- SearchToolView (the info page reached by clicking a tool) now uses
  ui/button, ui/card and a plain CSS-grid header instead of Tremor
  Card/Grid/Title/Text and antd Button
- SearchToolTester swaps antd Input/Button/Spin and Tremor Card/Title for
  ui/input, ui/button and UiLoadingSpinner, with no inline styles
- SearchConnectionTest swaps antd Button/Divider/Typography and the inline
  keyframe spinner for ui/button, ui/separator and UiLoadingSpinner

Markup only; no behavior change. The list page (SearchTools) and its create
and edit forms stay on antd because they are Form-bearing and blocked until
the forms migration. Icons move from antd and heroicons to lucide. The
retired antd no-restricted-imports suppressions are pruned from the
baseline.

* fix(ui): drop redundant vertical padding in SearchToolTester card

The shadcn Card already applies py-6 and gap-6 to its flex children, so
the pt-6/pb-6/mb-6 added during the migration stacked on top of it and
roughly doubled the vertical whitespace. Keep only px-6 (Card has no
horizontal padding) and let the Card own the vertical rhythm, which
restores the original 24px spacing.

* style(ui): format SearchConnectionTest test file with prettier
2026-07-23 05:09:04 +00:00