* chore(ui): remove dead dashboard files and unused dependencies
knip flagged seven orphaned source/config files with no importers and
five declared dependencies that nothing in the tree uses. Removing them
shrinks the dashboard bundle's source surface and keeps the manifest
honest; vite stays installed transitively via vitest, so test tooling is
unaffected.
* fix(ci): restore serverRootPath.config.ts referenced by SERVER_ROOT_PATH workflow
The dead-code sweep removed e2e_tests/serverRootPath.config.ts, but its spec
(tests/login/serverRootPathRedirect.spec.ts) and the test_server_root_path.yml
workflow step still depend on it, so the redirect e2e job failed to load a
config that no longer existed.
* feat(ui): add admin flag to disable in-product UI nudges for everyone
Admins can now suppress the survey and Claude Code feedback popups for
all users via a single disable_ui_nudges UI setting, instead of relying
on each user dismissing them individually.
* fix(ui): suppress nudges while ui settings are loading
Gate nudgesDisabled on the ui-settings loading state so an admin with
disable_ui_nudges on doesn't see the survey prompt flash, and the
getInProductNudgesCall fetch doesn't fire, on a cold page load before
the flag resolves. Falls back to showing nudges if the fetch errors.
* test(ui): wrap CreateKeyPage test in QueryClientProvider
page.tsx now calls useUISettings (react-query), which needs a
QueryClient that layout.tsx supplies in production but the test did
not. Add the provider and mock getUiSettings so the query resolves.
With AI models capable of automated vulnerability discovery now publicly
available, we expect a large increase in report volume, much of it
unverified. Requiring a video of the exploit running against a live
instance raises the bar for submissions and keeps triage focused on
reproducible issues. Reports without a video will be closed and reopened
if one is added later.
Co-authored-by: stuxf <70670632+stuxf@users.noreply.github.com>
* fix(ui/mcp): reset OAuth hook state on modal close so a prior server's token no longer leaks into the next add-server session
* fix(ui/mcp): clear in-flight OAuth guard on reset and reset form/tools on modal close so nothing leaks on a parent-driven dismiss
models-and-endpoints, organizations, and virtual-keys each had a page.tsx
route under (dashboard)/ that is not in MIGRATED_PAGES, so the sidebar and
deep links never resolve to it and the route is unreachable. Each was a thin
wrapper that handed the shared view empty or no-op props (empty modelData with
a no-op setModelData, hardcoded empty organizations, no-op
setUserRole/setUserEmail), so reaching one would render a degraded page in any
case. The real wrapper belongs in the PR that flips each page into
MIGRATED_PAGES, written with eyes on it and a test
This continues the dead-scaffolding cleanup from #28891. The shared components
these wrappers rendered (ModelsAndEndpointsView, OrganizationFilters) stay,
since the legacy ?page= switch in app/page.tsx and src/components still import
them
* test(ui): add a data-driven App Router migration E2E smoke
Add a growing Playwright smoke for migrated pages: for each segment it deep-links
to the path route, asserts the URL and that the dashboard shell rendered, then
clicks off to a legacy page and asserts navigation still works. Driven by
e2e_tests/fixtures/migratedPages.ts, so adding a page is one line.
Runs in two situations against the same proxy: the default mount (npm run
e2e:migration) and a non-root SERVER_ROOT_PATH mount (npm run e2e:migration:root).
globalSetup now logs in at `${SERVER_ROOT_PATH}/ui/login` so the admin storage
state is valid under a prefix. Seeded with api-reference; append the rest as their
migrations merge.
* test(ui): support headed slow-motion + watch pauses in the migration smoke
Honor SLOWMO in the server-root-path config (the default config already did),
and add an env-gated E2E_WATCH_MS pause so a headed run lingers on each state.
Both are no-ops by default, so CI behavior is unchanged.
* test(ui): make the migration smoke a sidebar-click user journey
Rework the smoke from deep-linking to a real navigation journey: start at the
landing page, click the migrated page in the sidebar (expanding submenus for
nested items), assert the path route rendered, reload it (the check a wrong
server_root_path breaks), bounce to a legacy page and back, and — once two pages
are migrated — navigate directly between two migrated pages. Verifies via URL +
shell render, driven by the same fixture list.
* test(ui): address review on the migration smoke
Escape ROOT and segment before interpolating them into RegExp URL matchers so a
future segment containing regex metacharacters can't silently widen the match.
Make the server-root-path config fail fast when SERVER_ROOT_PATH is unset instead
of silently re-running the default mount and passing without exercising the prefix.
* test(ui): drop unused watch helper and fix stale smoke README
* test(ui): run the migration smoke under a server root path in CI
* test(ui): harden + instrument the server-root-path proxy reboot in CI
* test(ui): run the server-root-path migration smoke as its own CI job
Replace the in-place proxy reboot in e2e_ui_testing with a dedicated
e2e_ui_testing_server_root_path job that boots the proxy once with
SERVER_ROOT_PATH=/litellm, matching how every other proxy variant in the
config gets its own job rather than killing and relaunching the live proxy.
The reboot was failing deterministically: after pkill -9 and relaunch the
prefixed proxy never came back up on :4000 (connection refused), so the smoke
never ran. The readiness step that was supposed to surface the cause could
never reach its boot-log tail because CircleCI runs steps under bash -eo
pipefail and the preceding `curl -sv ... | tail` aborted the step with curl's
exit 7. Booting the proxy as the job's own background step lets any boot crash
land in that step's log instead of being swallowed.
The default e2e_ui_testing job is unchanged aside from dropping the reboot,
prefixed-readiness, and prefixed-smoke steps; the migration smoke still runs at
the root mount there via the default Playwright config.
The caller's PERSONAL max_budget was the wrong yardstick for /team/update: a
team's spend ceiling has nothing to do with the admin's own key budget. That
comparison was an unintended side effect of reusing _check_user_team_limits()
(which exists for the /team/new path) and broke the UI, which re-sends the
unchanged budget on every save.
New behavior on /team/update for standalone teams:
- A team admin (already authorized via _verify_team_access) may freely KEEP or
LOWER the team budget, and change models/tpm/rpm, without being gated by their
personal limits.
- GROWING a team's spend ceiling is a budget-authority action reserved for proxy
admins -> 403 for team admins. "Growing" covers both raising max_budget above
the team's current finite value and removing the cap entirely (max_budget=null,
detected via model_fields_set so an explicit null is distinguished from an
omitted field). For a team that currently has no cap, setting a finite value is
a restriction and is allowed.
- Org-scoped teams remain governed by _check_org_team_limits() (capped by the
org budget).
Also reverts the #29525 existing_team_max_budget workaround in
_check_user_team_limits() back to the create-only form; /team/new still enforces
the creator's personal caps.
docs(access_control): resolve the contradiction in the team-admin section —
team admins can keep/lower the budget and manage rate limits/models, but cannot
raise the team budget (proxy-admin only).
tests: unit + behavior coverage for raise-blocked, cap-removal-blocked (team
admin), raise/removal allowed (proxy admin), uncapped-team restriction allowed,
keep/lower/resend allowed, and unchanged create-path guards.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ui): load MCP tool configuration tools via the OBO/passthrough-aware GET path
* fix(mcp): admin-only include_disabled_tools so the settings UI shows toggled-off tools
* fix(ui): repopulate MCP server edit form when server data loads after mount (OAuth return)
* fix(ui): persist MCP OAuth token on save and return to the Settings tab after authorize
* fix(ui): scope MCP OAuth callback to the initiating form so create and edit flows don't cross-talk
* fix(ui): derive OAuth-return Settings tab via lazy state init instead of setState-in-effect
* Fix MCP OAuth edit token handling
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* feat(azure_ai): add MAI-Image-2.5 image generation support
Route azure_ai MAI models to /mai/v1/images/generations and map OpenAI size to width/height for the serverless API.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(azure_ai): address MAI image generation review feedback
Validate unsupported size values, default width/height independently, add MAI-Image-2.5 pricing, and expand test coverage.
@greptileai
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(azure_ai): add MAI image edit and expand model cost map
Add MAI image edit support with usage normalization for Azure response format,
and register MAI-Image-2.5-Flash and MAI-Image-2e pricing in the model map.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(azure_ai): validate MAI edit size by consuming map iterator
Greptile: lazy map() never evaluated int() so values like 1024xabc passed through.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(azure_ai): normalize MAI usage in generation response handler
Apply normalize_mai_image_usage before building ImageResponse so token-based
cost calculation works when Azure returns num_output_tokens fields.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(azure_ai): narrow MAI edit size param type for mypy
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix Azure MAI image response handling
* Fix MAI image generation base model routing
* fix(azure_ai): preserve zero num_output_tokens in MAI usage normalization
* fix(azure_ai): wrap MAI generation response JSON parsing in error handling
* fix(azure_ai): build MAI image edit URL correctly for /mai/ root bases
* fix(azure_ai): build MAI image generation URL correctly for /mai/ root bases
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Capture user_id and extra_info from metadata or litellm_metadata. The single-bag read dropped identity whenever a request carried a present litellm_metadata field (null or a user-supplied dict), since /chat/completions routes the authenticated identity into metadata while the guardrail read litellm_metadata first
* feat(vantage): include organization metadata in FOCUS Tags export
Join LiteLLM_OrganizationTable when building Vantage/FOCUS export rows so
organization_id and organization_alias appear in Tags for org-level filtering.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(focus): include api_requests in organization Tags tests
FocusTransformer now requires api_requests after staging merge; add the
column to test fixtures so integrations CI can run the Tags assertions.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
A team's BYOK models (rows in LiteLLM_ProxyModelTable with model_info.team_id set)
were left orphaned when the team was deleted; they lingered in the database and kept
showing on the Models + Endpoints page. delete_team now removes them via a new
delete_team_models helper that deletes the rows in one transaction and syncs the
in-memory router only after that transaction commits, run before the team rows are
deleted so a mid-flight failure never leaves the team gone with its models orphaned
Raise the PyJWT floor in pyproject (>=2.13.0,<3.0) and re-resolve uv.lock so
the proxy installs 2.13.0 instead of 2.12.0. Bump the ws transitive-version
override in the dashboard from 8.19.0 to 8.20.1 and regenerate package-lock;
jsdom and openai both dedupe onto the single 8.20.1 copy.
Both are routine dependency maintenance bumps to keep pinned versions current.
* fix(vertex): propagate Vertex AI metadata in streaming success callbacks
Streaming calls assembled via stream_chunk_builder were missing
vertex_ai_grounding_metadata and vertex_ai_url_context_metadata in
standard_logging_object.response. Merge metadata from chunks into the
assembled response and mirror non-streaming hidden_params on Gemini chunks.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(vertex): move streaming metadata merge into provider config hook
Address review feedback by delegating assembled-stream metadata propagation
to VertexGeminiConfig via BaseConfig.apply_assembled_streaming_response_metadata,
and only write chunk hidden_params when metadata is non-empty.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(redaction): scrub Vertex provider metadata when message logging is off
Clear vertex_ai_grounding_metadata and related fields from standard
logging responses and assembled streaming ModelResponse objects so
turn_off_message_logging cannot leak prompt-derived web search queries.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Use assembled model for streaming metadata hook
* Fix Vertex metadata redaction bypass in logging callbacks.
Scrub Vertex provider fields from litellm_params.metadata.hidden_params during perform_redaction so streaming success_handler merges do not leak prompt-derived metadata when message logging is disabled.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix Vertex streaming metadata from hidden params
* fix(vertex): mirror vertex_ai_safety_results on assembled streaming responses
The non-streaming transform_response stores safety data under
vertex_ai_safety_results, but the streaming path only wrote
vertex_ai_safety_ratings. Assembled streaming responses therefore never
carried vertex_ai_safety_results, so any consumer reading that field saw
a silent difference between streaming and non-streaming calls.
Set vertex_ai_safety_results alongside vertex_ai_safety_ratings in the
shared stream metadata setter and add it to the assembled metadata field
list so it propagates through stream_chunk_builder.
* fix(streaming): log provider streaming metadata hook failures instead of swallowing them
* refactor(vertex): share single Vertex metadata field tuple across redaction and streaming
* refactor(vertex): move Vertex metadata redaction helpers into llms/vertex_ai
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
glm-5p1 supports native tools on Fireworks; explicit false flags caused
drop_params to strip tools and tool_choice before the provider request.
Co-authored-by: Cursor <cursoragent@cursor.com>
On /team/update for a standalone (no-org) team, _check_user_team_limits()
compared the request max_budget against the caller's personal max_budget
whenever max_budget was present in the payload. A team admin whose personal
budget is lower than the team's budget could not edit any field (tpm_limit,
team name, etc.) because the UI re-sends the unchanged max_budget on every
update, tripping the personal-budget check.
Pass the team's current max_budget into _check_user_team_limits() and skip the
personal-budget comparison when the incoming value is unchanged or lower than
the team's current budget. Only genuine increases above the team's current
budget are still validated against the caller's personal limit, so no
over-relaxation. Proxy admins and the org-scoped path are unaffected.
Adds two regression tests for the standalone update path (unchanged budget +
tpm_limit change, and lowering the budget), both for a caller whose personal
budget is below the team budget.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(jwt-auth): defer to single-team DB fallback on claim mismatch
Extends the single-team DB fallback introduced in #26418 to two more
cases where it previously could not run:
* `find_and_validate_specific_team_id`: when `team_id_jwt_field` is
configured and a claim value is present in the token but the team
does not exist in the LiteLLM DB (HTTPException 404 from
`get_team_object`), return `(None, None)` instead of raising — the
auth_builder fallback then attributes the request to the user's
single DB team. Only HTTPException is caught; other errors (e.g.
"No DB Connected") still propagate.
* `find_team_with_model_access`: when none of the `team_ids_jwt_field`
groups resolve to a real LiteLLM team, return `(None, None)` instead
of raising 403 so the same fallback path runs. If at least one group
DID resolve to a team but none granted the requested model, the
original 403 is preserved (legitimate access denial — not a claim
mismatch). Tracked via the new `any_claim_team_resolved` flag.
The strict `is_required_team_id` raise and `enforce_team_based_model_access`
raise remain unchanged. Unit tests cover both new soft-fail paths and
guard each preserved path (strict required, enforce_team_based, the
preserved 403, and the non-HTTPException propagation).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(jwt-auth): narrow HTTPException catch to 404 (greptile review)
Address Greptile review comments on #28913:
* `find_and_validate_specific_team_id`: re-raise HTTPException when
`status_code != 404`, pinning the catch to the "team doesn't exist
in db" path documented for `get_team_object`. A future change that
introduces a different status code (e.g. 403 for a blocked team)
will now propagate instead of silently falling through to the
single-team DB fallback.
* Add `test_find_and_validate_specific_team_id_non_404_http_exception_propagates`
parametrised over 400 / 403 / 500 to lock in the contract.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(jwt-auth): gate claim-mismatch fallback behind opt-in flag
The unresolved-team-claim fallback added in the previous commit
weakened the strict claim-based authorization contract by default —
an authenticated user whose JWT carries a stale or invalid team
claim could still consume their single DB team's models/quota via
the fallback.
Gate both soft-fail paths in `find_and_validate_specific_team_id`
and `find_team_with_model_access` behind a new opt-in flag
`team_claim_fallback` on `LiteLLM_JWTAuth` (default False).
Default-off preserves the pre-existing strict behavior. Operators
who intentionally treat IdP team claims as advisory (e.g. machine
tokens whose group claims live in a separate namespace from
LiteLLM team_ids) opt in via config.
Adds two regression tests guarding the default-off behavior.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(model-management): allow deleting a BYOK model after its team is deleted
A team BYOK model (model_info.team_id set) became undeletable once its team
was deleted: POST /model/delete ran can_user_make_model_call, which looked the
team up and raised 400 "Team id=... does not exist in db" before the delete
could run, so the model lingered on the Models + Endpoints page with no way to
remove it.
Drop the team-existence prerequisite from the delete path. When the model's
team still exists the normal auth check runs unchanged; when it is gone a proxy
admin may delete the orphan and any other caller gets a 403. The check is
fail-closed, so a missing or errored team lookup can only block the delete or
require an admin, never grant a non-admin access. Add/update/health keep their
team-existence validation.
* refactor(model-management): drop redundant team lookup on model delete
Move the orphaned-team handling into can_user_make_model_call behind an
allow_missing_team flag instead of pre-checking team existence in delete_model.
The endpoint no longer issues its own litellm_teamtable lookup, so deleting a
model whose team still exists hits the team table once instead of twice. The
auth behavior is unchanged: a proxy admin can delete a model whose team was
deleted, any other caller gets a 403, and add/update/health keep the strict
"team must exist" validation.
* feat(galileo): add health check support for UI callback test
Register galileo in /health/services so the proxy UI callback connection test works.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(galileo): verify API key via /current_user health check
Call Galileo's current_user endpoint so the UI callback test validates credentials against the provider.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(ui): regenerate schema.d.ts for galileo health service
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(galileo): return IntegrationHealthCheckStatus from async_health_check
Fixes mypy assignment error in health_services_endpoint where response was
narrowed to IntegrationHealthCheckStatus from earlier branches.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix Galileo logging to match Langfuse across all endpoint types.
Stop skipping ingest when output is empty and log embeddings with a placeholder so embedding, speech, and other non-text responses are recorded like Langfuse.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(galileo): remove unreachable health-check guard and None output sentinel
The use_v2_api flag is derived from bool(api_key), so the inner
GALILEO_API_KEY check inside the v2 branch could never run; collapse the
credential validation into the username/password path with a combined
message. _serialize_galileo_output now returns an empty string for None,
so _get_galileo_input_output_content always yields a str and the
post-call None coalescing guard is no longer needed.
* test(galileo): cover async_health_check failure paths and empty model response
Add regression tests for the Galileo health check unhealthy branches
(missing project id, missing base url, missing credentials, auth
failure, and request exception) and for logging a model response with
no choices, which now queues an empty output instead of being skipped.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>