The litellm-helm values file shipped the stock helm create boilerplate for
resources: an empty default plus a commented 100m/128Mi example it invites
operators to uncomment. 128Mi is roughly 32x below what the proxy needs at
DB-connected steady state, and it was the only sizing figure this chart ever
showed, so operators who followed it were sized for OOMKills.
Point the example at the documented 1 CPU / 4Gi per worker instead, link the
production sizing guidance, and note why the default stays unset. The
migration job's commented block carried the same trap with a 100m/100Mi
example; drop those numbers rather than substitute proxy figures that do not
transfer to a job that migrates and exits.
The defaults are deliberately left at {} so no existing release changes shape
on upgrade; rendered output is unchanged.
Second slice of the same sweep, covering src/components. Same rule as the
first: every removal is an unused import, an unused interface or type alias,
or a local const whose only mention was its own declaration.
The modelGroupOptions computation in add_auto_router_tab goes whole rather
than losing only its binding, since a Set and two arrays allocated per render
and then discarded is no better than the dead const was.
ToolDetail is deliberately left alone. Its unread teamsData traces back to a
useQuery that still issues a /team/list request, so removing it drops a
network call; that is a behavior change and belongs in a slice that gets QA'd,
not this one.
Stacked on litellm_dead_locals_1_app_routes; review that one first.
Part of LIT-5162.
Dropping the binding but keeping the initializer left two statements that
compute a value and throw it away: a ternary in ChatUI returning rawSelected
from both branches under a comment about resolving server IDs, and an
isAdminRole call in the prompts panel that also kept its import alive.
Both computations were already unreachable in effect; remove them whole.
* test(e2e): retry provider-transient statuses at the transport with bounded backoff
The Anthropic passthrough cost test failed a full-suite run on a real 529
overloaded_error. Passthrough routes forward provider responses verbatim
and bypass the router's num_retries, so provider blips reach the harness
only on those paths. Following standard practice, the retry is scoped to
the dependency boundary instead of rerunning tests: only the enumerated
transient statuses (500/502/503/504/529, the set production SDKs retry by
default) are retried, with bounded exponential backoff and a printed line
per retry so flakiness stays visible in run logs.
429 is deliberately excluded: the quota suites assert the proxy's own
rate-limit and budget 429s, and a transport that absorbed them would break
those tests. Network errors and timeouts are not retried either, so a hang
surfaces as a hang. request_with_retry takes injected callables, and the
new harness tests pin the contract with protocol fakes, no monkeypatching
* test(e2e): narrow the transport retry to 529, the one status the proxy cannot emit
Greptile's review is right that status-only classification could absorb an
intermittently failing proxy: at the transport a 500/502/503/504 from the
proxy is indistinguishable from one it relayed, and the proxy is the system
under test. 529 is the only status litellm provably never originates
(Anthropic's overload signal, forwarded verbatim on passthrough) and the
only transient observed across the full-suite runs, so the set shrinks to
exactly that. The canary tests now also pin 500/502/503/504 as never
retried
mcpTokenStore was the only OAuth path writing straight to window.sessionStorage;
useMcpOAuthFlow, useToolsOAuthFlow, the callback page and the edit-screen UI state
all already go through secureStorage. Align it so the OAuth surface has one storage
format instead of two.
The stored payload also carried a refresh_token that nothing ever read back. All
three read sites take access_token only, and nothing reads the mcp-session-token:
keys directly, so the field was write-only. Drop it from the store and from the four
callers that populated it. The client-forwarded modes (true_passthrough and
oauth_delegate) re-authorize rather than refresh, and authorization_code is
unaffected because it persists through storeMCPOAuthUserCredential on the backend,
which keeps its own refresh token.
Entries written before this change decode to null and are treated as absent, which
surfaces the normal Authorize prompt; they are session-scoped and expire in an hour.
Add two regression tests that decode the stored value before asserting, so neither
can pass merely because the payload is no longer plain text.
The Locust throughput SLO test is a different testing category from
functional e2e (variance-driven, historically flaky, currently
skip-annotated against LIT-5119) and erodes trust in the suite as a
release gate; it comes out of the default collection along with its
exclusive plumbing (locustfile, load-mock registration fixtures,
run_chat_load). Re-implementation as its own pipeline is tracked in
LIT-5163. The weekly session-anomaly test never ran in the suite (opt-in
via E2E_WEEKLY_ANOMALY, driven by its own workflow) and stays, as do the
markerless aggregation unit tests.
The vllm passthrough test read-times-out (60s) against the shared
vllm-cpu backend in every run on the per-SHA e2e stack; it is removed
until LIT-5164 establishes whether that is backend capacity or a
passthrough defect. Its registry cells return to the gap list, which is
the honest state
Backend parity check: a membership-granted org admin key gets 401 on
/v1/tool/list (route absent from org_admin_allowed_routes), so the
capability map denying the formatted Org Admin runtime value is the
intended behavior, now pinned by a test
@typescript-eslint/no-unused-vars is disabled in the dashboard eslint
config, so unused locals accumulated with nothing to catch them. This is
the first slice: symbols under src/app that no code reads.
Every removal is an unused import, an unused interface or type alias, or a
local const whose only mention was its own declaration. Nothing else on the
touched lines changes, so no behavior moves with it.
Part of LIT-5162.
Module-level names in litellm/__init__.py are the SDK's documented config
surface: users assign litellm.api_key and friends directly, and the proxy
rebinds them via setattr from litellm_settings. The package ships py.typed,
so the Final sweep made every such documented assignment a mypy error
("Cannot assign to final name") in downstream codebases. Strip Final from
the module scope of that file, keep it on function locals, and teach LIT010
that the config surface's module scope is exempt so the gate stays green
without suppression comments
* test(e2e): cover legacy text /completions endpoint
The /completions (and /v1/completions) text-completion route had zero e2e
coverage despite being the second-busiest endpoint in production; everything
'completions' in the suite was chat. Add a text-completion endpoint test that
registers an OpenAI instruct deployment, drives /v1/completions through the
gateway, and asserts real generated text. Adds text_completions() + the
completion request/result models to EndpointsClient, the 'completions' endpoint
to the coverage registry vocab, and the registry cell.
* test(e2e): assert /v1/completions choices shape, not just joined text
Assert the response carries a choices array and the first choice has real text,
so a malformed response (no choices) and a clean-but-empty completion are
distinct failures. Drop the unused text property / id / model fields (model only
what the test reads).
Moves the proxy extra's cryptography floor from 48.0.1 to 49.0.0 and widens the
ceiling to <51, then holds the lock at 50.0.0 with a uv override
mlflow caps cryptography at <50 even in its newest release, so publishing a
plain >=50.0.0,<51.0 range would make `pip install "litellm[proxy,mlflow]"`
unresolvable for downstream consumers. Publishing >=49.0.0,<51.0 keeps that
combination installable (it resolves to 49.0.0), while the
override-dependencies entry, which is a uv workspace setting and never reaches
published metadata, keeps our own lock and Docker images on 50.0.0
mlflow only uses PBKDF2HMAC, AESGCM, Fernet and InvalidTag from cryptography;
none of those are affected by the 49 or 50 breaking changes, so overriding its
ceiling is safe in practice
Lock delta is cryptography 48.0.1 -> 50.0.0, the mlflow trio 3.14.0 -> 3.15.0
and msal 1.36.0 -> 1.37.0
cryptography 49 dropped its x86_64 macOS and 32-bit Windows wheels. Linux CI
and the Docker images are unaffected; developers on Intel Macs will build from
source
Anthropic's structured outputs reject any `additionalProperties` value other
than `false` ("output_format.schema: For 'object' type, 'additionalProperties:
true' is not supported. Please set 'additionalProperties' to false")
`filter_anthropic_output_schema` only added the key when it was absent, so an
explicit `true` (or a sub-schema) was copied verbatim into output_format.schema
and 400'd. Coerce it for object schemas instead, at every recursion depth,
matching what the Anthropic Python/TypeScript SDKs do
The permissive tool-use path (map_response_format_to_anthropic_tool, used for
vertex_ai) is deliberately left alone
Fixes#35808
* test(e2e): cover vendor strategy gaps for chat contract, image edits, auth, team activity
Resolves the first slice of LIT-4778 (vendor API testing strategy): image edits happy path, chat multi-turn + validation + sanitization, LLM-route auth header matrix, and /team/daily/activity structure
* test(e2e): expand vendor API strategy coverage across endpoints
Adds validation cases on existing endpoint suites, plus vector stores, search,
bedrock native, realtime HTTP secrets/calls, responses retrieve, files/batches
contract, and chat stream SSE. Registers coverage cells for LIT-4778
* test(e2e): finish vendor strategy open items
Audio transcription negatives, vector-store file attach/poll/search,
OpenAI moderation category matrix across chat/messages/responses, and
smoke model matrix for chat (LIT-4778)
* test(e2e): harden vendor strategy suite against live env edges
Fix stream [DONE] tracking, XSS no-crash contract, realtime model routing,
vector store list/search models, responses validation, and provider-denied
Bedrock paths so the suite is stable against a live proxy
* test(e2e): rename suites, drop vendor_contract, fix greptile gaps
Move shared status helpers into e2e_http, rename chat auth headers and
chat security suites, remove vendor_contract and dev_config files_settings,
and tighten transcription validation plus vector-store search assertions
* test(e2e): route bedrock stream disconnects through e2e_http
Catch mid-stream RequestException in the shared harness so bedrock native
tests do not import requests directly
* fix(e2e): address greptile and veria review on vendor strategy suite
Store search tool keys as os.environ refs and resolve them in SearchAPIRouter.
Tighten validation helpers and assertions so 5xx/empty/unrelated failures no longer pass coverage cells
* fix(e2e): drop search_api_router os.environ expansion from vendor suite
Keep the PR test-only. Search tools register without an api_key so the
proxy falls back to its own PERPLEXITY/TAVILY env, same pattern as a2a.
* test(e2e): drop search e2e suite from vendor strategy PR
Remove the /v1/search coverage file and its registry rows so this PR
no longer carries search endpoint testing.
Replace Any-typed payload dicts, record shapes, and provider request/response
seams with TypedDicts, Protocols, and precise annotations in the ten litellm/
files carrying the highest combined basedpyright reportAny + reportExplicitAny
counts. No behavior changes.
Adds a regression test covering the managed-id list path so a prisma client
missing the managed tables keeps returning a fail-closed empty page.
api.ts owned the base-vs-origin precedence, the trailing-slash trim and the
base+path join inline. That logic belongs with the rest of base resolution and
was only reachable through a fetch client, so it could not be tested directly.
Extract resolveRequestUrl into resolveApiBase.ts with its own unit tests.
api.ts now only wires the shared resolver into openapi-fetch's Request option.
No behaviour change: same precedence, same trimming, same output.
* chore(build): move the Admin UI toolchain to Node 24
Node 18 and Node 20 both reached end of life (2025-04-30 and 2026-04-30), and
the release images along with every CI lane were still building on them. Node 24
is the current LTS through 2028-04-30, so this moves the four UI build images,
the CircleCI lanes, and the four GitHub Actions workflows onto it
Node 24 also ships npm 11.17, which is the first line that implements the
min-release-age setting this repo already carries in its .npmrc files. On npm 10
the key is parsed and discarded, so the release-age gate has had no effect
regardless of its value. Tightening the dashboard's engines range and turning on
engine-strict makes an unsupported npm fail loudly rather than skip the gate
quietly, and a new step in the UI build workflow probes an impossible cooldown
so an inert setting cannot pass unnoticed again
Node 24's bundled undici tightened its brand check on RequestInit.signal, which
rejects the AbortSignal jsdom installs and broke the two cases in
src/lib/http/api.test.ts that rebase a request onto a runtime base url. Under
jsdom the Request global comes from Node while AbortSignal comes from jsdom;
tests/jsdomFetchEnv.ts delegates to the jsdom environment and then restores
Node's native AbortController and AbortSignal so both come from one realm.
Upgrading jsdom does not address this, as jsdom still does not own Request
The workflows now read ui/litellm-dashboard/.nvmrc instead of repeating a
literal, so the Node version has a single source of truth, and ui/Dockerfile is
pinned by digest to match the other three build images. The lockfile changes are
npm 11 normalising the engines range and dropping optional peer entries it no
longer records
* fix(build): point every Admin UI build script at .nvmrc
The enterprise Docker path was left on Node 18. docker/build_admin_ui.sh runs
only when enterprise/enterprise_ui/enterprise_colors.json is present, which it
never is in the OSS tree, so neither CI nor a default image build reaches it;
it pinned nvm to v18.17.0 and then built the dashboard, which now requires Node
24, so a customized enterprise image would have failed EBADENGINE
All three UI build scripts now resolve the version from
ui/litellm-dashboard/.nvmrc rather than carrying their own pin, so the Node
version has a single home across Docker, CI, and local builds. build_ui.sh was
on v20 and build_ui_custom_path.sh on v18.17.0
Also drops the dependency-cooldown probe from the UI build workflow. The
engines floor plus engine-strict already fails an unsupported npm loudly at
install time, so the probe was redundant, and treating any nonzero exit from a
live registry call as proof of enforcement made it unsound besides
api.ts read globalThis.location when the module loaded, which froze the base
URL at import and pinned its test file to jsdom. The creation-time baseUrl and
the middleware's runtime rebase were also two mechanisms doing overlapping
work, and the rebase hand-copied eleven RequestInit fields on every call.
Pass openapi-fetch's Request option instead, so the constructor applies
whatever getRequestBaseUrl() returns at the moment the request is built.
registerBaseUrlGetter is now the single source of the base URL, rebaseUrl and
rebaseRequest are deleted, and the request is constructed once, so the init
openapi-fetch assembled reaches the platform Request untouched. The abort
signal is no longer copied by hand.
This preserves behaviour rather than approximating it: getProxyBaseUrl() falls
back to location.origin, so the runtime base was never empty in a browser and
the old middleware already rebased every request, discarding the creation-time
value each time.
setupTests.ts gates its DOM-only tail behind a window check; setup files run
for every environment, so that tail previously stopped any node-environment
test file from loading.
api.test.ts now runs under @vitest-environment node with its assertions intact
and no location stub, plus regressions for per-call base resolution and abort
forwarding. api.sameOrigin.test.ts covers the browser fallback to the page
origin, which needs a DOM environment.
* fix(proxy): propagate user_email and bind api_key on JWT auth paths
Standard JWT auth built UserAPIKeyAuth with user_id but never user_email, and the first auto-registered request early-returned a key with token set but api_key unset, so spend-log attribution logged user_api_key_user_email and user_api_key_hash as null. Bind api_key to the token hash on the auto-registered key, copy user_email from the resolved user object on both the standard and auto-register JWT paths, and warn when enable_jwt_auth/litellm_jwtauth are placed at the config top level where they are silently ignored.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover misplaced top-level JWT config warning
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
Internal users saw the Tool Policies page but its /v1/tool/list call always
returned 401. This adds a single source of truth for which roles may trigger
which UI fetches (utils/capabilities.ts) plus a useCan hook, and wires the
Tool Policies route through it: the nav item, the page, and the query all
read the same capability, so the sidebar hides the entry, deep links render
an admin-only notice, and the query never fires. The tools list call also
moves onto a queryOptions factory
* feat(ui): add admin-configurable user banner
Proxy admins can publish a markdown announcement that renders as a
dismissible banner on every dashboard page for all authenticated users,
editable from Admin Settings > UI Settings without a redeploy. Backed by
new /get/user_banner and /update/user_banner endpoints persisting to the
existing LiteLLM_UISettings table
* fix(ui): re-surface dismissed banner on identical republish
Stamp a server-side revision on every banner update and fold it into
the client dismissal signature, so unpublishing and republishing the
same message reaches users who dismissed the earlier run
* fix(ui): stamp banner revision as an opaque uuid instead of a counter
Two overlapping admin updates could read the same prior revision and
both persist the same incremented value, letting an identical republish
collide with a previously dismissed signature. A server-generated uuid
per update makes every publication identity unique by construction with
no read-modify-write
* refactor(ui): drop the server-side banner cache
Reads go straight to the single-row table; the dashboard already
throttles fetches client-side, so the cache only added staleness
windows under concurrent updates and multiple workers
* refactor(ui): move banner storage behind a domain repository and drop the store_model_in_db gate
UserBannerRepository owns the row shape instead of the endpoint
reaching through the generic .table bridge, and publishing no longer
depends on the unrelated STORE_MODEL_IN_DB flag; a connected database
remains the only requirement
The reload-failure warning built its message with an f-string, so the
interpolation ran on every failed reload whether or not the warning level was
enabled. `test_logging_calls_do_not_build_their_message_eagerly` scans the whole
litellm package and asserts zero offenders, so this one call has been reddening
`misc / Run tests` on litellm_internal_staging for every branch cut from it
Passing the reason as a %-style argument defers the interpolation to
`record.getMessage()`, which only runs once the record passes the level check