A failed fetchAvailableModels call disabled every preset with no way to recover short of closing and reopening the modal, even for a transient error. Surface the failure next to the template selector with a Retry action that re-runs the same useQuery.
The Template selector carried a manual required asterisk with nothing behind it: submitting with no template chosen fell through to the unrelated missing-tiers error instead of naming the actual problem. Add an explicit check ahead of the existing tier/classifier/semantic validation and an inline hint under the selector, reusing the showValidationErrors flag the rest of the form already uses.
sample_spec duplicated anthropic_family's exact model list under a different label, kept only as shape documentation; that duplication could drift silently if the real preset's models changed without a matching edit. Delete it and drop the now-unneeded filter in autorouter_presets.ts.
availableModelSet was rebuilt on every render while presets right above it was already memoized; wrap it in useMemo for consistency and to stop recomputing it on unrelated re-renders.
The model list drives preset availability and is scoped to accessToken. Loading it with a useState + useEffect + ignore-flag loader coupled the fetch lifecycle to manual resets, and each token-change edge (stale list, out-of-order resolution, config erased on reset) was handled by adding another line to that effect. This replaces the whole loader with useQuery keyed on accessToken, matching how EditFallbacks and the rest of the dashboard fetch model data.
react-query owns the race surface: a caller switch is a new query key, so the previous caller's list is never read for the new caller and out-of-order resolutions are discarded by key. There is no reset-on-token-change anymore, so a token change can no longer erase the user's in-progress configuration; only the model list is token-scoped, and preset availability recomputes from it. isLoading and isError replace the hand-rolled load-state union.
Net effect deletes the two model-related useState hooks, the loader effect, and the ignore guard. Tests updated to drive the query mock and assert re-gating on caller switch; mutation-checked (removing accessToken from the query key, skipping the loading gate, and swapping the nullish match_threshold prefill each turn a test red).
The prior fix for token-change races reset ALL configuration when the token changed, erasing user-edited tiers, keywords, semantic settings, and adaptive config alongside the preset selection and model-verification state. This was a regression: only the preset-tied state is token-scoped; user-entered config survives a token change.
Revised: on token change, reset only the preset selection (setSelectedPreset), model load state, and the cached model list. The user's manually-edited complexity-router config, keywords, and adaptive settings persist. This closes the availability race without erasing user input.
Updated test assertions to match the corrected design: "clears the preset selection" and "ignores a stale in-flight fetch".
* fix(proxy): redact credential headers from request logging copies
clean_headers preserves an Anthropic subscription OAuth token, and other
client-supplied provider credentials, so they can be forwarded upstream. The
same dict was also stored as proxy_server_request["headers"] and
metadata["headers"], so those credentials reached every logging callback and
the SpendLogs proxy_server_request column that the Admin UI logs page renders.
Build the observability facing copies through redact_credential_headers, and
drop the transport-only keys (provider_specific_header, headers, api_key) from
the request body snapshot since they have to keep the real values.
* fix(proxy): use the redacted header copy in the request debug log
The stdout secret filter matches Bearer and sk- shaped values, so an MCP auth
token printed by the request-header debug line survived it in cleartext.
* fix(proxy): resolve the configured MCP auth header name through the secret manager
get_secret_str also consults a configured secret manager, so a deployment that
stores the header name there now gets that header masked too. Drops the added
comments in favour of a named constant.
* perf(proxy): resolve the MCP auth header name once per process
get_secret_str issues a blocking secret-manager SDK call when one is configured,
and configured_credential_header_names runs on every proxied request.
* fix(proxy): read the MCP auth header name live, cache only the secret manager
The config reloader rewrites os.environ on an interval and after /config/update,
and MCPRequestHandler resolves the same setting per request, so caching the env
lookup left a renamed header logged in the clear until the process restarted.
Only the blocking secret-manager call stays cached.
* refactor(proxy): narrow header redaction to the reported credential set
Drops the MCP header-name resolution, its per-request config and secret-manager
lookups, and the x-mcp- prefix rule. Those cover a separate credential family
than the one this ticket reports and carried their own config-reload staleness
surface; they belong in their own change.
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds 61 unit tests on the modules #35694 extracted: 46 on the payload
builder, 15 on the OAuth redirect snapshot. They run in 9ms against 240s
for the 77 full-render tests they partly replace. Nine of nine mutants
were killed when the extracted logic was deliberately broken, so the
speed does not come at the cost of signal.
Deletes six cases across four blocks that rendered the whole modal to
assert one payload key belonging to a field they never touched. Every
test that proves a form field reaches the right payload key stays; those
cover field to form value to payload, which a unit test cannot reach.
Replaces "should not render when user is not an admin", which asserted
the admin title was absent and so passed for the wrong reason: the modal
does render for a non-admin, retitled. registerMCPServer was mocked but
never asserted anywhere, leaving the whole non-admin submission path
uncovered. It now drives a real submit and asserts the call lands there
and never on createMCPServer.
Renames the slow file to CreateMCPServer.integration.test.tsx and
documents the three tiers in the dashboard CLAUDE.md. No production code
changes.
Team-scoped DD credentials (dd_api_key, dd_site) set via POST /team/{id}/callback were silently dropped because _request_blocked_callback_params blocks them from standard_callback_dynamic_params. The security block is correct for request-level injection, but team callback_vars are admin-configured and trusted.
Store the raw init kwargs on the Logging instance and read dd_* params from there in _process_dynamic_callback_list instead of from standard_callback_dynamic_params.
Adds an integration test that exercises the full Logging.__init__ flow with team callback_vars to prevent regression.
Co-authored-by: Aanchal Khandelwal <aan2210khandelwal@gmail.com>
* feat(ui): expose an Auto-Router session affinity toggle
session_affinity on ComplexityRouterConfig defaults to True, and neither the
create form nor the edit modal ever emitted the key, so every auto-router built
in the UI silently pinned each session to its first turn's model for an hour
with no way to see or change that.
Adds an "Advanced: Session Affinity" switch to both surfaces, defaulted on to
match the backend field. Both paths now write the key explicitly instead of
falling through to the backend default, so a stored config states what the
router actually does. A stored config with the key absent hydrates as on, since
those routers are running with affinity enabled today; showing them as off would
report the opposite of reality and persist it on the next save.
* feat(complexity_router): default session affinity off and expose it in the UI
session_affinity defaulted to True and the Auto-Router UI never emitted the
key, so every router built there silently pinned each session to whatever model
its first turn classified into for an hour, refreshed on every hit. There was
no way to see that from the UI and no way to change it without hand-editing
config.yaml.
The default flips to False, so every turn is classified on its own merits and
lands on the cheapest adequate tier. Pinning is now opt-in.
The toggle added in the previous commit follows the field: it renders off, and
both the create tab and the edit modal keep writing the key explicitly, so a
stored config states what the router does instead of inheriting a default that
can move under it.
Behavior change for existing routers: those created before this have no
session_affinity key stored, so they pick up the new default and start
reclassifying every turn. That gives up the provider prompt cache the pin was
preserving, and a multi-turn session can now change model between turns. Set
session_affinity: true to keep the old behavior.
Key and team `router_settings.model_group_alias` was accepted, persisted and
echoed back by `/key/info`, but never applied at request time, so the request
ran on the group the caller asked for. `route_request` forwards only the
settings the Router accepts as per-request kwargs, and `model_group_alias` is
not one of them: the Router resolves aliases from its own instance attribute,
which holds the global config map and is shared across requests.
Resolve the alias in the proxy instead, alongside the existing model-alias
rewrites and ahead of the pre-call hooks, so per-model limits and guardrails
key off the group that actually serves the request. Authorize the alias target
before the rewrite; model access was checked against the requested group, so a
key whose alias points at a group it cannot call gets the usual 403 rather than
being quietly served it.
Resolves LIT-4879
Google Cloud has renamed Vertex AI RAG Engine to "RAG Engine" and
Vertex AI Search to "Agent Search" in its console. Users following our
setup instructions hit a naming mismatch when they cross-reference the
GCP console. Keep "Vertex AI" as the primary term (the generic new
names would make our provider UI ambiguous) and surface the new names
as secondary asides only where users leave the UI for the console.
Resolves LIT-3081
PR #35492 was authored before the ruff sweep removed Union from the typing
imports in litellm/llms/openai/common_utils.py, so the merge landed an
annotation referencing Union without an import. The annotation is evaluated at
class-definition time, so importing litellm raises NameError and every test
shard on litellm_internal_staging fails at collection. Rewrites the annotation
(and the same latent one in openai.py) as httpx.Client | httpx.AsyncClient |
None, matching the file's PEP 604 style, so no typing import is needed at all.
Pulls four modules out of the 1398-line create component, which drops to
896 lines. No behavior changes: CreateMCPServer.test.tsx is untouched and
all 77 of its tests pass against the refactored component, which is the
review contract for this PR.
createServerPayload.ts is a pure form-values-to-payload function whose
failures are a tagged union instead of inline notification calls, so the
transformation is reachable without a DOM. createOAuthUiState.ts owns the
snapshot that survives the OAuth authorize redirect, keeping every
presence guard the inline version had. AwsSigV4Fields and
OpenApiByokFields are the two largest JSX blocks, moved verbatim so they
can be diffed as moves.
The create/edit setToken divergence, the mcpLogoImg export, and the
untyped form-values bag are left alone on purpose; each is a behavior or
cross-file change that does not belong in a move.
Pure rename, no behavior change. create_mcp_server.tsx and its test move
to CreateMCPServer, the two importers and one stale e2e comment follow,
and the local/filename-pascal-case suppression drops now that the file
passes the rule on its own.
The rename is scoped to this one component rather than the whole
directory because three PRs are currently open against its snake_case
siblings; the rest can follow once those land.
An evicted client was left for the garbage collector, but every OpenAI/Azure
SDK client is a reference cycle, so nothing freed the client or its pooled TCP
connections until a generational sweep ran. Driving 2000 azure calls through
the official image with no forced collection, live clients and open sockets
climbed from 202 to 1361 while the cache stayed at its 200-entry bound, and RSS
grew 279 MB to 456 MB against a TLS upstream.
Closing on eviction is what caused the earlier 'Cannot send a request, as the
client has been closed' regression, so an evicted client litellm created is now
closed only once a grace window has passed, by which point any request that was
already holding it has finished. A client the caller supplied is never closed,
since litellm does not own its lifecycle.
Resolves LIT-4883
gitpython arrives transitively through mlflow-skinny, which accepts
>=3.1.9,<4, so this is a lock-only move with no pyproject change.
Relocked with `uv lock --upgrade-package gitpython`; gitpython is the only
package whose version changed. `uv sync --all-groups --all-extras` and
tests/test_litellm/integrations/test_mlflow.py pass on the result.
The dashboard pins both packages exactly in `overrides`, so the lockfile
stays on whatever those pins say. Move brace-expansion from 5.0.8 to 5.0.9
and postcss from 8.5.22 to 8.5.23, both upstream patch releases, and
regenerate the lockfile.
`npm ci`, `next build`, and the 5888-test vitest suite all pass on the
updated lockfile.
* feat(teams): apply default organization to new teams from default team settings
Adds organization_id to DefaultTeamSSOParams so proxy admins can pick a
default organization in Default Team Settings. new_team applies it before
org validation whenever a team is created without an explicit
organization_id, so API, Admin UI, SCIM, SSO, and team upsert creations
all inherit it and go through the same existence and org-limit checks.
Explicit organization selections win and existing teams are untouched.
The default is validated at save time (PATCH /update/default_team_settings
returns 400 for an unknown org) and at create time, where a missing org now
surfaces as a clean 400 instead of a 500 by routing OrganizationNotFoundError
into the previously dead org_table None guard.
The Admin UI Default Team Settings tab gets a Default Organization row
backed by the shared OrganizationDropdown.
* fix(teams): validate org limits against final team state including defaults
Applies default_team_params and the legacy max_budget fallback before the
organization validation block, so _check_org_team_limits sees the values the
team will actually be persisted with. Also loads the org's budget table in
the lookup; without include_budget_table every budget comparison in
_check_org_team_limits was skipped because litellm_budget_table was None.
* test(proxy_behavior): pin org team limits as enforced on /team/new
The dead-code pins existed to turn red when include_budget_table went
live; that happened, so the scenarios now assert the 400 rejections plus
within-cap acceptance, and the unknown-org pin asserts the handler's 400
instead of the surfaced 500.
Adds a Stream responses checkbox (default on) to the playground Model
Settings popover. When unchecked, chat completions and responses API
requests are sent with stream: false and the full reply renders at
once. The non-streamed result is replayed through the existing
streaming handlers as synthesized chunks/events so MCP events, vector
store results, usage and response ids behave identically in both
modes. TTFT is suppressed when not streaming; total latency now also
reported for the responses API. The toggle is scoped to the chat and
responses endpoints, persists via sessionStorage, and is isolated from
the simplified Agent Builder chat.
Resolves LIT-3251
* fix(proxy): backfill null user_email on existing users during JWT auth
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): guard mapped-key email backfill and make null update atomic
Resolve Greptile review on the JWT user_email backfill:
- only backfill when the mapped virtual-key owner is the JWT principal, so a
mismatched admin-created mapping cannot write one user's email onto another
- make the best-effort mapped-key enrichment non-fatal so a database outage on
a cached-key request no longer fails otherwise-valid authentication
- persist the backfill with an atomic null-guarded update_many so concurrent
writers cannot overwrite an already-populated email
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep cache coherent when a concurrent backfill wins the null-email update
* fix(proxy): cache DB-persisted email after JWT backfill, not the proposed value
Resolve the Greptile finding that a successful null-guarded backfill could
cache this request's proposed email even if a concurrent ordinary user update
wrote a different email first. The helper now always re-reads the row after the
atomic update and refreshes the cache from the value the database holds, so
cache-hit auth and attribution stay consistent with the persisted record.
Annotate the Prisma and model_copy dict literals to keep the LIT002 budget within its ceiling.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
disable_team_logging cleared only metadata["callback_settings"], but callbacks
registered through POST /team/{team_id}/callback and the Admin UI live in
metadata["logging"], and request-time resolution stops at that slot without
ever reading callback_settings. The endpoint reported success while the team
kept sending request and response data to its third-party destination.
Empty the logging slot alongside the existing callback_settings reset, and
refresh the cached team object so the change applies to keys that are already
in flight rather than at the next cache expiry. The same refresh is added to
add_team_callbacks, which has the symmetric problem of a newly registered
callback staying dormant until the entry expires.
Resolves LIT-5101
The strict-priority e2e (added with the zero-increment limiter fix) can
never pass on stage: the proxy there does not run the
dynamic_rate_limiter_v3 callbacks + priority_reservation settings the
module requires, confirmed by zero limiter log lines across every
gateway and backend pod during the 2026-08-02 run. Config lives in the
infra repo; LIT-5118 tracks adding it.
The throughput SLO test failed the same run with 65.9% of requests dying
at the ELB as 502/503 before reaching a pod. The per-replica SLO rework
fixed the RPS-floor assertion but cannot help when stage idles at one
warm gateway replica; LIT-5119 tracks pre-scaling the fleet for the load
phase.
Both skips name their ticket, and the coverage registry returns the two
cells to the gap list while they are in place.
A single read of key_info.spend races the batched spend writer: deltas
earned before a reset flush to the DB up to ~60s later
(proxy_batch_write_at) and land on the row after the reset zeroed it.
The stage runs on Jul 30 and Aug 2 failed
test_key_budget_reset_at_advances_after_window exactly this way, with
spend back at the driven total while budget_reset_at had advanced and
calls flowed again.
Replace the single reads in rung 3 (spend zeroed after reset) and rung 4
(roomy window keeps spend) with _poll_key_spend, which re-reads to a 90s
deadline covering one full flush-plus-reset cycle. A reset that never
zeroes the row keeps spend pinned and still times out, so the regression
guard keeps its teeth.
* fix(team-callbacks): report API-registered callbacks from GET /team/{team_id}/callback
POST /team/{team_id}/callback writes metadata["logging"] while the GET read
metadata["callback_settings"], so every team configured through the API or the
Admin UI got back an empty list. c620d76fe4 migrated the writer to the new key
and left this reader on the old one.
Resolve the read the same way request-time resolution does in
_get_dynamic_logging_metadata: a logging slot that is present wins outright and
callback_settings stays as the deprecated fallback, so the endpoint reports what
a request would really do rather than the union of both shapes. An empty logging
list therefore reports no callbacks, matching a request that fires none.
Decrypt callback_vars for the response and mask the credential keys. Ciphertext
would be unusable to the caller, and a value encrypted under a key that is no
longer classified as sensitive would otherwise come back as a raw blob.
Resolves LIT-5093
* Update litellm/proxy/management_endpoints/team_callback_endpoints.py
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(team-callbacks): mask callback vars that fail to decrypt
decrypt_callback_vars passes a value through untouched when it cannot be
decrypted, which happens to existing rows after a salt-key rotation. Under a
key that is not classified as sensitive that blob reached the caller as opaque
ciphertext it could not use or tell apart from a real value, so mask anything
still carrying the encrypted prefix.
Raised by Greptile on the first commit.
---------
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A keyless internal user signing in to the Admin UI was redirected off the
post-login landing to /ui/connect, which renders nothing but the MCP apps panel,
so a plain gateway sign-in ended on an MCP OAuth surface the user never asked
for. The landing now renders the keys dashboard for every role. The key lookup
that existed only to make that routing decision goes with it, along with the
useKeys enabled flag it was the sole caller of and the role-hydration hold that
guarded its one-frame dashboard flash
The gateway DCR consent flow moves the other way. Its /authorize handed the
browser to /ui/chat/integrations, whose layout hard-blocks when enable_chat_ui
is off, which is the default, and client-side redirects to /ui/ without the
query string; that destroys the connect_flow handle and strands the MCP client
until the 600s flow cookie expires. It now lands on /ui/connect, which reads
connect_flow and connect_client, mounts the consent banner and puts the apps
panel in connect mode. /ui/chat/integrations keeps its connect-mode handling
this release so flows sealed before the deploy still finish
Resolves LIT-5104
Resolves LIT-4911
About 35,000 fixes ruff marks safe across 32 rules (UP006/UP045/UP007
modern annotations, UP032 f-strings, SIM114/SIM118, RET501, and
friends), removal of the 1,296 typing imports the rewrite orphaned, and
hand fixes for what the fixers could not see: five star-import
freeloaders of typing names, two F823 late-import annotations, the
/get/config/list introspection crash on types.UnionType, redundant
function-local RoleMappings imports in ui_sso.py that shadowed the
module-level name once the annotation lost its quotes, and one FURB168
tautology.
B009/B010/PIE804/RUF019 are excluded on purpose: their safe fixes
rewrite getattr/setattr/**-splat/key-in-dict escape hatches into forms
basedpyright then rejects (283 new errors measured), so their budgets
stay at base values.
ruff-strict-budget.json drops by 39,579 this commit (39,968 across the
branch) with 28 rules at an actual 0 and 9 more sharply down.
type-discipline-budget.json ratchets LIT002/LIT006/LIT009 down; LIT001
moves to the now-honest total: the checker matches the spelling `set`
but not the alias `Set`, so the 160 typing.Set annotations rewritten to
set[...] were always mutable-set annotations and only now count.