Commit graph

42942 commits

Author SHA1 Message Date
Yassin Kortam
3615cccfef
fix(team): sweep dangling team references and cache on team delete (#36819)
* fix(team): sweep dangling team references and cache on team delete

delete_team drove all of its cleanup off the team's members_with_roles roster, so any
user row referencing the team by another route kept a dangling team id forever and the
deleted team stayed visible on /user/info. Nothing swept LiteLLM_UserTable.teams or
LiteLLM_TeamMembership by team id, schema.prisma declares no relation between the
membership table and the team table so there is no cascade to fall back on, and the
cached team object was never invalidated on delete.

Adds a sweep that runs before the team rows are dropped: it strips the deleted ids from
every user row that still lists them and removes every membership row for those teams.
Adds _delete_cache_team_object in auth_checks and calls it per deleted team so the
team_id:{team_id} entry cannot outlive the team.

The sweep is targeted, not indiscriminate: only the deleted ids are removed and the
other teams on a user record are left intact.

* fix(team): fail member_add when the team is deleted under the row lock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(team): correct the post-delete sweep note for the member_add lock path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-13 18:01:38 -07:00
Yassin Kortam
b344043eaf
fix(ui): stop a deselected MCP server keeping its grant on a virtual key (#36840)
* fix(ui): stop a deselected MCP server keeping its grant on a virtual key

The key editor sent `mcp_tool_permissions` unfiltered, and the MCP resolver
counts a server named only under `mcp_tool_permissions` as entitled, unioning
`tool_perm_servers` into `all_servers` at four sites in
`user_api_key_auth_mcp.py`. Deselecting a server, or removing the access group
that supplied it, therefore left a stale entry that kept the key reaching that
server with its old tool allowlist attached.

Reuse `extractMcpEntitlement`, which already landed for the internal-user
surface, so the key surface drops an entry only once the server is known and no
longer granted, and keeps it whenever a retained access group or toolset could
still supply it. The helper moves to a shared module so the key template does
not import a users page component.

Setting the map unconditionally is part of the same fix: the old
`Object.keys(...).length > 0` guard let the previous map ride through the
`object_permission` spread, which filtering to an empty map would otherwise hit
in exactly the case the fix is for.

* fix(ui): resolve retained MCP groups and toolsets per server when pruning tool permissions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): mock the MCP toolsets hook in the key update suite

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-13 18:00:42 -07:00
Yassin Kortam
e0f388cb5c
feat(ui): render request metrics on the /ui/chat surface (#36845)
/ui/chat never rendered a metrics bar. The Responses helper already parses
usage off the response.completed event, but the chat page passed positional
undefined where onTimingData, onUsageData and onTotalLatency sit, ChatMessage
had nowhere to hold them, and ChatMessages never rendered ResponseMetrics.

Thread the three callbacks through, persist the values on the assistant
message, and reuse the playground's ResponseMetrics to show latency, TTFT,
input/output/total tokens and cost. Also map the cost the proxy reports on the
streamed usage object, which only the chat-completions helper did before.
2026-08-13 17:31:13 -07:00
Emerson Gomes
4ea5749642
feat(azure-ai): add Grok 4.3 model metadata (#27932)
* Add Azure AI Grok 4.3 metadata

* Address Azure Grok 4.3 test feedback

* Drop empty tool choice in responses bridge

* style(azure-ai): update Grok metadata tests
2026-08-13 17:25:17 -07:00
Emerson Gomes
b85f557f30
fix: enable xhigh reasoning support for gpt-5.4-mini models (#26909)
* fix: sync gpt-5.4 reasoning capability flags

* fix(models): keep GPT-5.4 service tiers consistent
2026-08-13 17:24:58 -07:00
Emerson Gomes
603fe93758
feat(azure_ai): add Fireworks FW model pricing on Azure AI Foundry (#35613)
* feat(azure_ai): add Fireworks FW model pricing on Azure AI Foundry

* fix(azure_ai): drop incorrect FW-Kimi-K2.6-Code alias

* test(azure-ai): assert FW max token metadata

* feat(azure_ai): add Inkling and Nemotron 3 Ultra pricing
2026-08-13 17:24:38 -07:00
Yassin Kortam
04f5dedf69
feat(cli): make the hidden lite command list configurable (#36816)
* fix(cli): hide codex and opencode from the lite command listings

They stay registered and invokable, so existing `lite codex` users keep
working; they just no longer show up in `lite --help` or the interactive
shell's command list.

* feat(cli): make the hidden lite command list configurable

codex and opencode are supported, so hardcoding them as hidden was wrong. Let deployments curate their own listing with `lite config set hidden_commands codex,opencode` instead; nothing is hidden by default and hidden commands stay invokable.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-13 17:23:03 -07:00
tin-berri
83e890bdde
feat(ui): shadow evals tab beside auto-router usage (#36588) 2026-08-13 17:19:45 -07:00
Yassin Kortam
72ee0bb1c4
fix(cli): launch agents as a child process on Windows (#36822)
os.exec* has no process-replacement semantics on Windows, so `lite claude`
printed its routing line and returned to the prompt while Claude Code was left
detached without a usable console. Windows now spawns the agent, waits for it,
and exits with the child's status. Batch shims such as the npm-installed
claude.cmd go through cmd.exe because CreateProcess cannot run them directly,
and that command line is emitted verbatim with every token quoted so a spaced
path or an argument holding a shell metacharacter cannot be re-parsed by the
command processor. POSIX keeps using os.execvpe unchanged.
2026-08-13 17:06:25 -07:00
Yassin Kortam
56b08c19d6
fix(proxy/team): resolve member_delete cleanup by user id, not the addressed email (#36839)
/team/member_delete dropped the roster entry by matching user_email against
members_with_roles, then built its user-row lookup from that same raw email
instead of from the user_id the roster entry already carries. An email the user
row does not literally hold matched nothing, so the team id stayed in the user's
teams array and the team-membership row was left orphaned while the call still
returned 200.

/team/member_add resolves an email to a user case-insensitively but stores the
caller's casing on the roster, so inviting "Alice@Example.com" for a row holding
"alice@example.com" and removing by that same string is enough to reach it.

_cleanup_members_with_roles now returns the roster entries it removed, and both
the user-row update and the membership delete run against their user ids.
2026-08-13 17:00:52 -07:00
Yassin Kortam
ab2333b6c4
fix(auth): stop the team fallback from widening model access (#36837)
When get_team_object fails, the centralized auth gate rebuilds the team
from the token's own fields. A token whose team row was missing when the
key was read carries team_models=[] and team_blocked=False, and the
model-access check reads an empty model list as every model, so the
rebuilt team grants more than the real team ever did.

get_team_object reported a deleted team and a database that would not
answer as the same 404, so the fallback could not tell a definitive
answer from a degraded read. Raise a TeamNotFoundError subclass, still a
404 with the same detail so every other caller is unaffected, only when
the database answers and the row is absent.

A team that is provably gone now refuses, and no setting overrides that.
Otherwise the grant is merely unknown: a token carrying one may vouch,
since replaying a recorded grant cannot widen it, and a token carrying
none may not. allow_requests_on_db_unavailable still opts back out there,
and is only consulted once the failure is known to be a degraded read.
2026-08-13 16:59:58 -07:00
Yassin Kortam
b72dab8049
feat(ui): show provider prompt cache tokens in chat response metrics (#36827)
The chat metrics bar reported In/Out/Reasoning/Total/cost only, so a
playground user had no signal that provider prompt caching worked. The
cached-token counts were already visible in the Logs drawer, which meant
the answer to "does caching work here" lived on a different page.

Adds cacheReadTokens and cacheCreationTokens to TokenUsage and renders
them as two chips, reusing the prompt-cache tooltip wording already
introduced for the Logs drawer so both surfaces say the same thing.

A single helper, extractPromptCacheTokens, normalizes the three usage
shapes the playground consumes: Anthropic Messages
(cache_read_input_tokens / cache_creation_input_tokens), chat
completions (prompt_tokens_details) and the Responses API
(input_tokens_details). All three producers call it instead of parsing
per surface. Counts that are absent, zero or non-finite are dropped, so
providers without prompt caching render exactly what they render today.
2026-08-13 16:58:14 -07:00
Yassin Kortam
4bc27f1664
fix(auth): carry team grants in lite login session tokens (#36826)
CLI session tokens minted by /sso/cli/poll set team_id and team_alias but
never team_models or team_model_aliases, so the token carried a team with
none of that team's grants. /v1/models bails out to "unrestricted" when both
key_models and team_models are empty and listed the whole proxy, and team
model aliases never resolved because both can_team_access_model and the
pre-call rewrite read team_model_aliases off the token.

The team data was not close at hand: _fetch_cli_sso_team_details projected
full team rows down to team_id and team_alias before they reached the mint.
Widen that projection to include the team's models and its joined alias
table, and populate both fields at mint time.

Also stop writing the user's personal allowlist into the key models slot
when a team is bound, matching virtual-key semantics where a team-bound
credential is governed by the team grant.

Because an empty team grant is itself a real value meaning unrestricted, a
team whose grants cannot be resolved must not be minted as empty: that is
the same "unrestricted" bail-out this fix exists to close. The poll now
refuses to mint when the selected team has no complete cached detail.

That refusal is only safe because a login can no longer be pinned to a team
whose grants will never resolve. Deleting an organization drops its team
rows but leaves the memberships behind, so the login now offers only teams
whose rows still exist, and a lookup that fails outright fails the login
rather than caching a session that silently drops every team.
2026-08-13 16:56:47 -07:00
Mateo Wang
373a0fc506
Merge pull request #36717 from BerriAI/litellm_add_muse_spark_1_2
feat(model_prices): add meta/muse-spark-1.2 and its contributor tier
2026-08-13 16:19:06 -07:00
ryan-crabbe-berri
262ed530f8
fix(proxy): honor explicit null budget_duration on team and key create + clearable UI dropdowns (#36699)
* fix(proxy): honor explicit null budget_duration over default_team_params on /team/new

* fix(ui): clearable team budget reset with explicit Never resets option

* docs(proxy): align default_team_params docstrings with actual all-teams scope

* fix(proxy): honor explicit null budget_duration on /key/generate over configured defaults

* fix(proxy): keep upperbound_key_generate_params filling explicitly-null key params

* fix(proxy): restrict explicit-null default opt-out to budget_duration
2026-08-13 15:22:11 -07:00
yuneng-jiang
0c1355d54a
Merge pull request #36824 from BerriAI/litellm_/concurrent-view-creation
fix(proxy): tolerate a concurrent creator when creating spend views
2026-08-13 15:20:11 -07:00
yuneng-jiang
69792a9529
Merge pull request #36823 from BerriAI/litellm_/e2e-access-control-allowlist
test(e2e): assert the model allow-list permits, not only denies
2026-08-13 15:17:52 -07:00
Yuneng Jiang
e448049163
style(proxy): satisfy ruff format in create_views 2026-08-13 15:04:43 -07:00
devin-ai-integration[bot]
87736a767c
feat(ui): highlight Auto Router in the navbar announcement (#36315)
* chore(ui): remove Agent Platform announcement bell from navbar

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(ui): highlight Auto Router in the navbar announcement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): open Auto Router docs link in a new tab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "fix(ui): open Auto Router docs link in a new tab"

This reverts commit 3be861bafe.

* chore(ui): title the navbar announcement LiteLLM Auto Router

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-13 21:57:14 +00:00
mateo
cbcc3715c6 Merge branch 'litellm_internal_staging' into litellm_add_muse_spark_1_2 2026-08-13 21:54:41 +00:00
Yuneng Jiang
570b34988b
refactor(proxy): type the race helper's db against a Protocol
create_view_tolerating_race took the module's _db = Any. It now takes a
Protocol naming the single operation it calls, so the contract is checkable
at its call sites without retyping the rest of the module.

Kept free of Any deliberately: an earlier version typed the Protocol's
parameters as Any and pushed create_views.py from 26 basedpyright errors to
30 by adding reportExplicitAny. This version measures identical to the
baseline on both create_views.py (26) and utils.py (1355).
2026-08-13 14:53:59 -07:00
Mateo Wang
a4ab511d0b
Merge pull request #36805 from BerriAI/litellm_grok_4_6
feat(xai): day-0 pricing for grok-4.6
2026-08-13 14:49:41 -07:00
Yuneng Jiang
726292720c
fix(proxy): guard every view creation, not just the first and last
Against a real Postgres the previous commit still died on MonthlyGlobalSpend:
only 2 of the 8 creation sites went through the tolerant helper, so the losing
replica re-raised on the first unguarded one and skipped the rest.

The regression test now makes every CREATE lose the race and asserts all 8 are
still attempted, which fails on the partial fix.
2026-08-13 14:41:59 -07:00
yuneng-jiang
fa498391a4
Merge pull request #36129 from BerriAI/litellm_playground_shadcn
feat(ui): migrate playground chat controls to shadcn
2026-08-13 14:21:38 -07:00
mubashir1osmani
2a7f5f270c Merge remote-tracking branch 'berri/litellm_internal_staging' into litellm_playground_shadcn
# Conflicts:
#	ui/litellm-dashboard/eslint-suppressions.json
2026-08-13 13:54:11 -07:00
Yuneng Jiang
77e64c5d40
fix(proxy): tolerate a concurrent creator when creating spend views
Every replica booting against the same fresh database sees each view as
absent and issues the CREATE. Postgres fails all but one with a
duplicate-object error, and that exception propagated out of
create_missing_views, so every view after the first was never created and
/global/spend* 500'd for the life of the deployment.

Losing that race reaches the desired end state, so treat it as success.
Genuine DDL errors still propagate.
2026-08-13 13:42:57 -07:00
yuneng-jiang
69b0296ca3
Merge pull request #36793 from BerriAI/litellm_shadcn_logs_drawer_header_0813
refactor(ui): migrate SectionHeader and ToolsSection to shadcn
2026-08-13 13:41:36 -07:00
Yuneng Jiang
ed01f7316b
test(e2e): assert the model allow-list permits, not only denies
Every case in TestAccessControl asserted that something was refused. A gateway
that denied the allow-listed model too would have passed all of them, so the
suite could not tell "denied correctly" from "broken outright".

Adds the positive half: a key allow-listed for gemini-2.5-flash can call it and
gets back a real completion rather than a 200-wrapped error.

Also tightens the unknown-model case. It accepted any valid JSON, so a bare
"{}" or even "null" satisfied it. It now requires the OpenAI-shaped error
envelope with a message a client can actually surface, parsed through a typed
model instead of json.loads.
2026-08-13 13:37:31 -07:00
Yuneng Jiang
1dc0ea3d11
Merge branch 'litellm_internal_staging' into litellm_shadcn_logs_drawer_header_0813 2026-08-13 13:33:00 -07:00
yuneng-jiang
160548d40b
Merge pull request #36739 from BerriAI/litellm_/quirky-mcnulty-e432c4
refactor(ui): migrate TruncatedValue and OutputCard to shadcn
2026-08-13 13:31:55 -07:00
Mateo Wang
c1310de342
Merge pull request #36763 from BerriAI/litellm_decrease_anys_fable7
refactor: replace Any with precise types across responses, proxy, and llms modules
2026-08-13 13:24:47 -07:00
Yuneng Jiang
9f947a7406
Merge branch 'litellm_internal_staging' into litellm_/quirky-mcnulty-e432c4 2026-08-13 13:23:51 -07:00
yuneng-jiang
faea98ee7e
Merge pull request #36738 from BerriAI/litellm_/eloquent-wu-3ab1d5
refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn
2026-08-13 13:23:27 -07:00
Yuneng Jiang
2f7602b028
Merge branch 'litellm_internal_staging' into litellm_/eloquent-wu-3ab1d5 2026-08-13 13:16:48 -07:00
yuneng-jiang
c9b543dbfe
Merge pull request #36737 from BerriAI/litellm_/gifted-colden-355244
refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn
2026-08-13 13:16:22 -07:00
tin-berri
d8fda675cc
feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) (#36587) 2026-08-13 13:15:45 -07:00
Yuneng Jiang
b71e7a3213
Merge branch 'litellm_internal_staging' into litellm_/gifted-colden-355244 2026-08-13 13:07:51 -07:00
yuneng-jiang
0f5fa38ef2
Merge pull request #36735 from BerriAI/litellm_/interesting-meitner-e88370
refactor(ui): migrate TokenFlow and JsonViewer to shadcn
2026-08-13 13:04:50 -07:00
mubashir1osmani
f3f8c48b1d style(ui): format new playground tests and drop suppressions this stack fixed
Prettier flagged the two test files added while fixing review findings.
Migrating these components off Ant Design also retired the lint suppressions
they carried, so prune those 15 entries and leave the unrelated ones for the
PRs that made them stale.
2026-08-13 12:55:43 -07:00
yuneng-jiang
b8577516d4
Merge branch 'litellm_internal_staging' into litellm_/quirky-mcnulty-e432c4 2026-08-13 12:51:53 -07:00
yuneng-jiang
21dbd3381e
Merge branch 'litellm_internal_staging' into litellm_/eloquent-wu-3ab1d5 2026-08-13 12:51:51 -07:00
yuneng-jiang
229d258664
Merge branch 'litellm_internal_staging' into litellm_/gifted-colden-355244 2026-08-13 12:51:50 -07:00
yuneng-jiang
6f9ee6e28f
Merge branch 'litellm_internal_staging' into litellm_/interesting-meitner-e88370 2026-08-13 12:51:48 -07:00
Yuneng Jiang
81f324dc15
Merge branch 'litellm_internal_staging' of github.com:BerriAI/litellm into litellm_shadcn_logs_drawer_header_0813 2026-08-13 12:51:37 -07:00
Yuneng Jiang
dfdafbf89b
fix(ui): keep the tools panel mounted so a tool's expanded detail survives
antd's Collapse kept the panel mounted once opened, so a tool a user had
expanded stayed expanded after closing and reopening Tools. Base UI renders
only the open branch, so the migration silently reset every ToolItem.

The regression test passes against the antd original, fails against the
migration without keepMounted, and passes with it.
2026-08-13 12:50:56 -07:00
Anas Khan
7fcca523aa
fix(proxy/batches): stop forwarding custom_llm_provider twice in list and cancel (#32813)
* fix(proxy/batches): stop forwarding custom_llm_provider twice in list and cancel

The model-routing branches of list_batches and cancel_batch passed
custom_llm_provider as an explicit kwarg while also leaving it inside the dict
they splat, so every such call raised "got multiple values for keyword argument
'custom_llm_provider'" and returned a 500.

list_batches SCENARIO 2 called data.update(credentials) but never removed
custom_llm_provider before litellm.alist_batches(custom_llm_provider=..., **data);
it now uses prepare_data_with_credentials, the same helper the create and
retrieve branches already use, which pops it out.

cancel_batch SCENARIO 3 resolved the provider with
`provider or data.pop("custom_llm_provider", None) or ...`, so when the path
param provider was set the pop short-circuited and a body custom_llm_provider
stayed in data and collided with the explicit kwarg. The body value is now
popped unconditionally before the fallback chain, so the path param wins cleanly
and data no longer carries a duplicate.

Both paths already had strict-xfail regression tests documented "remove when
fixed"; those markers are dropped so the tests now guard the fix.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>

* fix(proxy/files): avoid duplicate custom_llm_provider in list

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>

---------

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-08-13 12:50:32 -07:00
Yuneng Jiang
b4092f476f
test(ui): find the section copy button by role instead of the antd icon
InputCard and OutputCard located SectionHeader's copy button by querying for a
descendant with aria-label="copy", which is the antd CopyOutlined icon. That
selector reaches into SectionHeader's internals, so migrating it off antd left
copyButton undefined and failed four tests.

getByRole("button", { name: /copy/i }) is green against both the antd and the
shadcn SectionHeader, verified by running these two files against each.
2026-08-13 12:47:45 -07:00
mubashir1osmani
8fd97803ce Merge remote-tracking branch 'berri/litellm_playground_shadcn' into litellm_playground_shadcn 2026-08-13 12:44:01 -07:00
mubashir1osmani
f4fb4e0bea fix(ui): label the virtual key source and unstick a cancelled model load
The key source trigger rendered the stored value, so the playground showed
session and custom instead of Current UI Session and Virtual Key. Name the
selected option on the trigger.

Clearing the key while models were loading left the selector disabled for
good: the in-flight load skips its reset once cancelled, and the branch that
handles an empty key returned without clearing the loading flag, so nothing
put it back. Clear it on that path too.
2026-08-13 12:43:25 -07:00
mubashir1osmani
3fdacfa6f9
fix(ui): restore playground model filtering by endpoint (#36130)
* fix(ui): restore playground model filtering by endpoint

Bring back the prior Chat model dropdown filter (including chat models
on responses/anthropic/interactions and image models on image_edits), and
map mode realtime so the realtime endpoint only lists compatible models

* fix(ui): exclude unknown model modes from playground endpoint filters

Modes outside ModelMode (batch, rerank, ocr, etc.) must not collapse to
chat-compatible, or conversational endpoints surface unusable models

* feat(ui): add shared vercel-style playground chat composer (#36131)

* feat(ui): adopt vercel-style chat composer for playground

Replace the compact single-line input with a PromptInput-style composer:
taller auto-growing textarea, rounded card shell, footer tools, and
stop button while a request is in flight

* style(ui): strengthen playground chat composer border and shadow

Make the shared chat input stand out with a fuller border, layered
shadow, and a slightly stronger focus ring

* fix(ui): size chat composer textarea with CSS field-sizing

Drop direct el.style.height mutation in favor of field-sizing:content

* fix(ui): keep the chat composer out of a nested form and focus its textarea

The composer wrapped everything in a native form, so MCP mode nested Ant
Design's tool-arguments form inside it, which is invalid HTML and let Enter
hit either form. The footer also relied on InputGroupAddon focusing the first
input in the group, which is the hidden file input from the attach controls
rather than the message textarea.

Drop the outer form and submit from the send button directly, and have the
addon focus the element marked as the group's control.

* refactor(ui): reuse the endpoint compatibility check when a model is picked

The endpoint guard added upstream duplicated the compatibility families this
PR introduces, so point it at isModelCompatibleWithEndpoint instead. Filtering
also means an incompatible model is no longer offered for an endpoint, so the
test that picked one now asserts it is absent.

* fix(ui): match the image-edit model mode the backend actually sends

model_prices_and_context_window.json labels these models image_edit, but the
mode enum spelled it image_edits, so once unknown modes started being filtered
out every image-edit model vanished from the playground, /v1/images/edits
included. The endpoint key keeps its own spelling.

The compatibility tests stubbed getEndpointType with a hand-written map that
repeated the same wrong spelling, which is how this stayed hidden, so they now
run against the real mapping.
2026-08-13 12:40:42 -07:00