* feat(router): add auto_router/quality_router for quality-tier routing (#25987)
* feat(router): add auto_router/quality_router for quality-tier routing
Adds a new auto-router type that routes a request to a model at a target
quality tier. The quality tier is inferred by re-using the existing
ComplexityRouter's classification, then mapped through an admin-configured
complexity_to_quality table. Each candidate model declares its own
quality_tier in model_info.litellm_routing_preferences.
Resolution strategy: exact tier match, else round up to the next higher
tier, else fall back to default_model.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): add capability-based filtering
Each deployment can declare a `capabilities: List[str]` field in
`model_info.litellm_routing_preferences` (e.g. ["vision",
"function_calling"]). Requests can pass `litellm_capabilities` in
`request_kwargs` to require specific capabilities — the router will only
route to deployments whose declared capabilities are a superset.
Resolution still walks tier (exact → round up), but at each tier filters
by capability before picking. Falls back to default_model only when it
also satisfies the required capabilities; otherwise raises rather than
silently routing to a model that lacks a required capability.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): expose routing decision in response headers
For transparency, expose the QualityRouter's routing decision in the
proxy response headers:
x-litellm-quality-router-model → picked model_name (e.g. "haiku-vision")
x-litellm-quality-router-tier → resolved quality tier (e.g. "1")
x-litellm-quality-router-complexity → ComplexityTier name (e.g. "SIMPLE")
Mechanism: the pre-routing hook stashes the decision in
request_kwargs["metadata"]["quality_router_decision"]. After the call
returns, Router.set_response_headers lifts the decision into
response._hidden_params["additional_headers"] alongside the existing
x-litellm-model-group / x-litellm-model-id headers. Existing metadata
keys (trace_id, user_id, etc.) are preserved.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): replace capabilities with keyword override
Drops the capability-based filtering in favor of a keyword-based override
for v0:
- RoutingPreferences.keywords: List[str] (replaces capabilities) — each
deployment can declare substring keywords.
- If any declared keyword (case-insensitive) appears in the user message,
the router short-circuits the complexity-classification flow and routes
to the matching deployment.
- Tiebreaker for overlapping keyword matches: quality_tier DESC, then
cheapest model_info.input_cost_per_token ASC. Unpriced models lose ties
to priced ones.
Decision metadata + headers now expose the override:
x-litellm-quality-router-via → "keyword" | "quality_tier"
x-litellm-quality-router-keyword → matched keyword (only on keyword route)
x-litellm-quality-router-complexity → complexity tier (only on tier route)
Removes:
- request_kwargs["litellm_capabilities"] reading
- _model_capabilities, _model_supports_capabilities,
_first_capable_model_at_tier, capability filter in
_resolve_model_for_quality_tier
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): add explicit `order` to RoutingPreferences
Adds an explicit priority field to RoutingPreferences for resolving
collisions deterministically:
RoutingPreferences.order: Optional[int] # lower wins; unset = +inf
Used as the PRIMARY tiebreaker in two places:
1. Keyword overlap: when multiple deployments declare the same matching
keyword, sort by (order ASC, quality_tier DESC, input_cost_per_token
ASC, model_name ASC). Explicit always beats implicit.
2. Tier resolution: when multiple deployments share a quality tier,
`_resolve_model_for_quality_tier` picks the one with the lowest
order. The tier list is now sorted at index-build time.
This lets admins make routing decisions explicit when the natural
quality-and-price ordering would pick the wrong model.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): reorder tiebreak to (quality, order, price)
Changes the tiebreak ordering so quality_tier always wins first, then
explicit `order` is used to break ties within the same tier, then price
breaks the rest:
1. quality_tier DESC ← best model wins first
2. order ASC ← explicit priority within a tier
3. input_cost_per_token ASC
4. model_name ASC
Previously `order` was the primary key — that meant a tier-2 model with
`order=1` would beat a tier-3 model with no `order`, which is the wrong
default. Now `order` only resolves collisions among same-tier candidates.
Tier resolution (within a single tier) keeps the same key minus quality:
(order ASC, cost ASC, name).
Test renames + flips:
- test_explicit_order_overrides_quality_tier → test_quality_wins_over_explicit_order
- new: test_order_breaks_tie_within_same_quality_tier
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* fix(quality_router): resolve Greptile review feedback
Addresses four P1 findings from PR review plus test coverage:
1. set_model_list missing quality_routers reset
- Hot-reloading the Router would leave stale QualityRouter instances
pointing at the old model_list. `set_model_list` now clears
`self.quality_routers` alongside the other indices.
2. Round-down fallback before default_model
- `_resolve_model_for_quality_tier` now rounds DOWN to the closest
lower tier after round-up fails, before falling back to
`default_model`. Degrades gracefully rather than jumping straight
off-tier.
3. RoutingPreferences validation bypass
- `_build_tier_index` now instantiates `RoutingPreferences(**prefs)`
so invalid shapes (e.g. non-int quality_tier) raise a clear
ValueError instead of silently succeeding.
4. Config-ordering dependency
- `_tier_to_models` is now built lazily on first access. Previously,
eager construction in `__init__` meant a QualityRouter deployment
had to appear AFTER all its referenced models in config.yaml,
because `Router._create_deployment` populates `model_list`
incrementally. Any `available_models` defined after the router
entry would silently be reported as missing.
Also adds 6 new tests covering each fix:
- test_invalid_quality_tier_type_raises_clear_error
- test_router_can_be_instantiated_before_its_targets_exist
- test_set_model_list_clears_quality_routers_registry
- test_rounds_down_when_no_higher_tier_exists
- test_rounds_down_prefers_closest_lower_tier
- test_prefers_round_up_over_round_down
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* style: apply black 24.10.0 formatting to pre-existing offenders
Unblocks the LiteLLM Linting check for this PR — these 12 files are already
failing `black --check` on main (the lint workflow only runs on PRs, so main
drifts). No behavior changes; formatting-only.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Update litellm/router.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Support /v1/responses in complexity router (#26137)
* feat(proxy): add --reload flag for uvicorn hot reload (dev only)
Opt-in CLI flag, off by default, no env var. Only affects the uvicorn
run path; gunicorn/hypercorn paths and prod (which doesn't pass the
flag) are unaffected.
* Feature/add audio support for scaleway (#26110)
* feat(scaleway): add SCALEWAY to LlmProviders enum
* feat(scaleway): add audio transcription config and dispatch wiring
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(scaleway): add behavior tests for audio transcription config
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(scaleway): advertise audio_transcriptions in endpoint-support JSON
* docs(scaleway): document audio transcription support
* fix(scaleway): address PR review — plain-text response_format + missing-key fail-fast
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(scaleway): cover new response paths, drop gettysburg.wav coupling
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Prompt Compression - add it to the proxy (#25729)
* refactor: new agentic loop event hook
simplifies how to create logic for tool based multi llm calls
* fix: compress - make it work on anthropic input as well
* fix(compress.py): working prompt compression for claude code
ensures claude code messages can run through proxy easily
* docs: add agentic loop hook guide
* docs: add agentic_loop_hook to sidebar
* fix: fix multiple arguments error
* fix: fix tool call loop for compression on streaming /v1/messages
* fix: fix linting errors
* fix: fix ci/cd errors
* feat(litellm_pre_call_utils.py): use claude code session for litellm session id
allows claude code logs to be stitched together, making it easy to know they were all part of the same conversation
* fix: suppress incorrect mypy warning rE: module
* revert: drop PR's changes to litellm/proxy/_experimental/out/
Restores the 34 HTML files under _experimental/out/ to their pre-PR
paths (X/index.html -> X.html). All renames are R100 (content
unchanged); no other files are touched.
* fix: address greptile review comments on PR #25729
- Skip ``kwargs["tools"] = []`` injection when compression is a no-op —
Anthropic Messages rejects empty tool arrays on requests that did not
originally declare tools.
- Move agentic-loop safety guards (fingerprint cycle / max depth) out of
the per-callback try/except so they propagate instead of being swallowed
by the generic exception handler. Extracted _check_agentic_loop_safety.
- Gate generic ``x-<vendor>-session-id`` capture behind the
LITELLM_CAPTURE_VENDOR_SESSION_HEADERS env var (off by default) to
preserve backwards compatibility; explicit x-litellm-* headers are
unaffected.
- Fix monkeypatch target in pre-call-hook test to patch the actual
module-level binding
(litellm.integrations.compression_interception.handler.compress).
- Add regression tests for empty-tools skip and opt-in session capture.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* revert: drop LITELLM_CAPTURE_VENDOR_SESSION_HEADERS flag
Generic x-<vendor>-session-id header capture is a new feature and only
runs *after* the explicit x-litellm-trace-id / x-litellm-session-id
checks, so it does not change behavior for any existing caller that was
already using the LiteLLM headers — no backwards-incompatibility to gate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor(compress): replace input_type with CallTypes call_type
Drop the bespoke ``CompressionInputType`` literal and use the existing
``litellm.types.utils.CallTypes`` enum instead. ``litellm.compress()``
now takes ``call_type: Union[CallTypes, str]`` (default
``CallTypes.completion``) — no new concept to learn, and the enum is
already the way the rest of the codebase talks about request shapes.
Supported values: ``completion`` / ``acompletion`` (OpenAI chat-completions
shape) and ``anthropic_messages`` (Anthropic structured content blocks).
Updated: compress(), the compression_interception handler, tests, docs,
and the two eval scripts.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* Support /v1/responses in complexity router
Adds cross-format support to the complexity router via the guardrail
translation handler dispatch. Adds get_structured_messages to base
translation plus OpenAI chat, Responses, and Anthropic handlers.
Auto-router helper _extract_text_from_messages handles tool-call and
multimodal messages. Widens async_pre_routing_hook messages type to
Dict[str, Any].
Fixes https://github.com/BerriAI/litellm/issues/25134
* chore: apply black formatting
* fix: fallback to trying each handler when route inference fails
---------
Co-authored-by: Ryan Crabbe <ryan@berri.ai>
Co-authored-by: nhyy244 <106547304+nhyy244@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: cover _is_quality_router_deployment and init_quality_router_deployment
* fix: reset auto_routers on set_model_list to prevent hot-reload ValueError
* style: apply black formatting to websearch_interception and agentic_streaming_iterator
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Ryan Crabbe <ryan@berri.ai>
Co-authored-by: nhyy244 <106547304+nhyy244@users.noreply.github.com>
Run auth_ui_unit_tests against a per-job cimg/postgres:16.0 sidecar
with DATABASE_URL pointing at localhost:5432, matching the pattern
used by e2e_ui_testing. Seed the schema via 'litellm --skip_server_startup
--use_prisma_db_push' so each run starts on a clean DB with the current
schema.prisma.
Six tests in test_hooks.py were written against an older API and had been
failing in CI. Updated:
- test_resolve_session_key_* (4 tests): _resolve_session_key now requires
at least SIGNAL_GATE_MIN_MESSAGES messages before deriving a hash (it
returns None on shorter convos to match the signal-processing gate).
Switched the tests to use _long_messages() so they hit the hash path.
- test_post_call_success_hook_* (2 tests): the hook was migrated from
async_post_call_success_hook (mutates response._hidden_params) to
async_post_call_response_headers_hook (returns a headers dict) because
the former fires too late for streaming responses. Rewrote the tests
against the new API; added a metadata-not-dict noop case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The previous revert left three files under
_experimental/out/_next/static/8Gn6tA2K4jsPxzCCMwObH/ still staged
(git checkout origin/main -- only restores files present in main, it
does not delete files that exist only on the PR branch). Remove them
so the PR no longer touches _experimental/out/ at all.
Anthropic retired claude-3-haiku-20240307 on 2026-04-20, causing the
test_anthropic_messages_litellm_router_non_streaming_with_logging
test to 404. Update the model references in this file to the current
pinned haiku version.
The UI rebuild bundled into this PR is not needed for the Dockerfile
change — the image simply copies whatever _experimental/out/ is in the
tree. Regenerating here conflicts with the release-time refresh policy
and adds ~100 files of review noise / merge-conflict risk for any
concurrent UI PR.
* fix: /health/readiness returns 503 when DB is unreachable due to handle_db_exception re-raising
handle_db_exception() re-raises the Prisma exception inside _db_health_readiness_check's
except block, which propagates out to health_readiness() and gets wrapped in a 503.
The health endpoint never reached the reconnect path and the service never recovered.
Fix:
- Remove handle_db_exception() call from _db_health_readiness_check — that helper is
for API request handlers (allow_requests_on_db_unavailable flag), not health checks
- Replace raw disconnect()+connect() with attempt_db_reconnect(), which uses the proper
lock, cooldown, escalation, and heavy-reconnect (recreate_prisma_client) machinery
* test: update health readiness tests for handle_db_exception removal
- Remove tests that expected handle_db_exception to re-raise (old buggy behaviour)
- Remove tests asserting disconnect()/connect() calls (replaced by attempt_db_reconnect)
- Add regression tests covering the 503 loop fix:
- transport errors never raise (ClientNotConnectedError, httpx.ConnectError, etc.)
- reconnect success path returns 'connected'
- reconnect failure path returns 'disconnected' without raising
- non-transport errors return 'disconnected', skip reconnect
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
The post-call hook was hardcoding tool_results=[] on every Turn, so the
failure detector never saw tool errors and the bandit only learned from
satisfaction — never from negative tool outcomes.
Added _recent_tool_results(messages): walks the request messages from the
tail and collects the contiguous run of role=='tool' entries — those are
the results from the most recent assistant tool_calls round. Normalizes
each to {content, is_error}, the only fields signals._detect_failure /
_detect_exhaustion read.
Tests: 6 new covering empty input, trailing-run extraction, is_error
propagation, boundary at first non-tool message, no-trailing-tool case,
and the end-to-end path from hook -> Turn.tool_results.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Add supported providers to prompt caching doc
* Move Z.ai / GLM to cache_control marker list
* Mark xAI models as supporting prompt caching
* Narrow xAI prompt caching flag to models with documented cache pricing
* Add prompt caching flag to grok-4, grok-4-0709, grok-4-latest
---------
Co-authored-by: Michael Riad Zaky <michaelr@Michaels-MacBook-Air.local>
- Use 'auto_router/adaptive_router' prefix in example yaml, docs, and
README — the old 'adaptive_router/...' and 'openai/gpt-4o-mini' values
silently skipped adaptive-router init because detection requires the
'auto_router/adaptive_router' prefix.
- Read x-litellm-min-quality-tier from request headers (and the
'min_quality_tier' metadata key as fallback) in async_pre_routing_hook.
Previously the documented header was defined but never extracted, so
the quality-floor feature was inert.
- Evict expired entries from _session_states. The cache grew without
bound — added a parallel expiry map (same TTL as _owner_cache) and an
opportunistic bulk sweep when the cache crosses a size threshold.
- Align adaptive-router migration SQL with Prisma schema: all count
columns and the 'clean_credit_awarded' / 'last_processed_turn' fields
are NOT NULL in the data model, so the migration now declares them
NOT NULL. Fixes test_aaaasschema_migration_check.
Tests: 8 new covering header/metadata/precedence/invalid-value paths for
min_quality_tier and TTL-based eviction of _session_states.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: new agentic loop event hook
simplifies how to create logic for tool based multi llm calls
* fix: compress - make it work on anthropic input as well
* fix(compress.py): working prompt compression for claude code
ensures claude code messages can run through proxy easily
* docs: add agentic loop hook guide
* docs: add agentic_loop_hook to sidebar
* fix: fix multiple arguments error
* fix: fix tool call loop for compression on streaming /v1/messages
* fix: fix linting errors
* fix: fix ci/cd errors
* feat(litellm_pre_call_utils.py): use claude code session for litellm session id
allows claude code logs to be stitched together, making it easy to know they were all part of the same conversation
* fix: suppress incorrect mypy warning rE: module
* revert: drop PR's changes to litellm/proxy/_experimental/out/
Restores the 34 HTML files under _experimental/out/ to their pre-PR
paths (X/index.html -> X.html). All renames are R100 (content
unchanged); no other files are touched.
* fix: address greptile review comments on PR #25729
- Skip ``kwargs["tools"] = []`` injection when compression is a no-op —
Anthropic Messages rejects empty tool arrays on requests that did not
originally declare tools.
- Move agentic-loop safety guards (fingerprint cycle / max depth) out of
the per-callback try/except so they propagate instead of being swallowed
by the generic exception handler. Extracted _check_agentic_loop_safety.
- Gate generic ``x-<vendor>-session-id`` capture behind the
LITELLM_CAPTURE_VENDOR_SESSION_HEADERS env var (off by default) to
preserve backwards compatibility; explicit x-litellm-* headers are
unaffected.
- Fix monkeypatch target in pre-call-hook test to patch the actual
module-level binding
(litellm.integrations.compression_interception.handler.compress).
- Add regression tests for empty-tools skip and opt-in session capture.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* revert: drop LITELLM_CAPTURE_VENDOR_SESSION_HEADERS flag
Generic x-<vendor>-session-id header capture is a new feature and only
runs *after* the explicit x-litellm-trace-id / x-litellm-session-id
checks, so it does not change behavior for any existing caller that was
already using the LiteLLM headers — no backwards-incompatibility to gate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor(compress): replace input_type with CallTypes call_type
Drop the bespoke ``CompressionInputType`` literal and use the existing
``litellm.types.utils.CallTypes`` enum instead. ``litellm.compress()``
now takes ``call_type: Union[CallTypes, str]`` (default
``CallTypes.completion``) — no new concept to learn, and the enum is
already the way the rest of the codebase talks about request shapes.
Supported values: ``completion`` / ``acompletion`` (OpenAI chat-completions
shape) and ``anthropic_messages`` (Anthropic structured content blocks).
Updated: compress(), the compression_interception handler, tests, docs,
and the two eval scripts.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
- Use LoggingCallbackManager.add_litellm_callback instead of
litellm.callbacks.append (required by callback_manager_test)
- init_adaptive_router_deployment now uses model_name_to_deployment_indices
for O(k) lookup instead of scanning model_list
- Rephrase comment in set_model_list to avoid the 'in self.model_list'
substring that the linear-scan test greps for
- Whitelist _finalize_adaptive_router_if_configured in
test_no_linear_scans_in_router — prefix match on 'auto_router/adaptive_router'
has no supporting index; runs once at init
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
prisma --version invokes the Schema Engine, which has no binary for the
Wolfi base image (only debian). In the baseline this was silenced by a
trailing || true wrapping the whole prisma chain; removing that wrapper
uncovered the failure on arm64 builds. The main Dockerfile does not
call prisma --version at all, so drop it here to match — prisma generate
is sufficient to validate the toolchain.
Revert the .dockerignore ui/ exclusion and remove the UI Drift Guard
workflow. _experimental/out/ refresh is already handled by the release
runbook; the global .dockerignore change also broke Dockerfile.custom_ui
(explicit COPY ./ui/litellm-dashboard) and the enterprise-colors inline
rebuild path in Dockerfile, Dockerfile.database, and Dockerfile.dev.
Dockerfile.non_root itself is unchanged functionally — still stages the
UI from the checked-in _experimental/out/. Only the companion workflow
and global dockerignore exclusion are dropped.
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Has been cancelled
Router coverage check flagged this method as untested. Adds two cases:
- initializes AdaptiveRouter from model_list and is idempotent on re-entry
- no-op when no adaptive deployments are configured
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unrelated timestamp and version drift was showing in the PR diff. This PR adds no new deps — keep uv.lock identical to main.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Remove _experimental/out/ changes from this PR — these are auto-generated Next.js build outputs, not part of the adaptive router feature.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Has been cancelled
Vertex generateContent returns INVALID_ARGUMENT if cachedContent is sent
with system_instruction, tools, or toolConfig; those belong on CachedContent.
Fixes#26014
Made-with: Cursor
Five small, individually-verified cleanups collected into one commit:
- Drop 'prisma migrate diff --from-empty ... > /dev/null 2>&1 || true'
from the builder. Stdout/stderr/exit-status all discarded; nothing
reads the output. Dead line.
- Drop 'mkdir -p /app/.cache/npm' from the same RUN. npm is gone.
- Drop the runtime's redundant 'sed -i' + 'chmod +x' on the entrypoint
scripts. The builder already does the same three lines, and the
runtime copies /app from the builder via COPY --from=builder, so
the normalized files (and exec bits, which buildkit preserves) are
already in place.
- Drop NPM_CONFIG_CACHE and NPM_CONFIG_PREFER_OFFLINE from the runtime
ENV — nothing reads them after Task 2.2 removed npm.
- Drop '/.npm' and '/tmp/.npm' from the runtime's mkdir + chown. These
directories only existed as npm's writable dirs for the non-root
user; npm is gone.
.dockerignore: add 'ui/'. After Task 2.1 the non_root image sources
its UI bytes from litellm/proxy/_experimental/out/, so the whole
ui/litellm-dashboard/ source tree is dead weight when the blanket
'COPY . .' pulls it into /app. Verified (with ripgrep) that no Python
code under litellm/ opens any file under ui/. All string references to
'ui/...' are URL paths, not filesystem paths.
Final image size: 6.57GB baseline -> 1.96GB. API parity and UI visual
regression match baseline across all 12 API scenarios and 10 UI
routes. Trivy HIGH/CRITICAL: 6 -> 2, no new CVEs introduced.
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
Mount /app/.cache/uv as a BuildKit type=cache on both 'uv sync' steps.
The cache persists across builds on the same builder (and, when used
with type=gha in CI, across CI runs) so repeat builds don't re-download
every wheel.
Side-effect: because the cache lives outside the image layer, the
~742MB of downloaded wheel archives that were previously baked into
/app/.cache/uv drop out of the final image. Compressed image size
goes from ~5.0GB to ~3.7GB, and the 'USER nobody' prisma-generate
layer is 1.7GB vs 2.4GB.
Warm-build timing: a uv-sync-invalidating edit now takes ~1m30s vs
~2m39s without the cache mount, on this dev VM.
API parity and UI visual regression continue to match baseline.
Trivy HIGH/CRITICAL: 6 at baseline -> 2 now, no new CVEs.
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
After Task 2.1 removed the in-image Next.js build, the builder stage no
longer needs a full C/C++ + Clang toolchain. Keep gcc + python3-dev
(required to compile ml-dtypes 0.4.1 from source — no wheel published
for Python 3.13 yet). Drop everything else.
Removed from apk: clang, llvm, lld, linux-headers, build-base,
openssl-dev, npm. Removed NVM_DIR env and /root/.nvm from PATH
(no nvm-based Node install anymore).
Kept: python3, python3-dev, gcc, bash, coreutils, curl, openssl,
libsndfile, nodejs. gcc (15.2) serves both C and C++; the separate
g++ package doesn't exist in Wolfi.
Image size unchanged (builder stage doesn't end up in the runtime);
cold builds slightly slower due to ml-dtypes source compile, but that
will be recovered in the next task via a BuildKit uv cache mount.
API parity and UI visual regression both match baseline, Trivy
HIGH/CRITICAL CVE count unchanged from opt-2 (4 CVEs, none new).
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
npm was installed in the runtime only to globally install vulnerability
patched versions of tar/glob/brace-expansion/minimatch/diff and to
in-place rewrite npm's own bundled package.json. Both were to silence
CVE scanners against modules that ship with npm itself.
Since we no longer run npm anywhere in the runtime (Prisma uses the
node binary directly for migrate deploy and generate), we can just
skip installing npm in the first place. This eliminates both the
~25-line CVE-patch shuffle AND the underlying CVE surface.
Kept: nodejs (needed by prisma-python's CLI and migrate deploy).
Removed: npm apk package, all 'npm install -g', all find+sed patching,
the redundant 'apk upgrade --no-cache nodejs' (already covered by the
preceding 'apk upgrade').
Image: 4.97GB (opt-1) -> 4.97GB (opt-2); the real win is that two
CVEs (CVE-2026-33671 and GHSA-q4gf-8mx6-v5v3) drop off the Trivy
HIGH/CRITICAL list. No new CVEs introduced. API parity and UI
visual regression both match baseline.
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
The checked-in Next.js static export at litellm/proxy/_experimental/out/
is kept fresh by the UI Drift Guard CI workflow. Stage it directly
instead of re-running npm ci + npm run build inside the image.
This removes: nvm install, node 20.20.2 install, npm ci (801 pkgs),
next build, and the resulting intermediate node_modules/out tree.
Build time: ~6m25s -> ~2m (fuse-overlayfs DinD); image 6.57GB -> 5.0GB.
Behavior parity verified: API endpoints, UI screenshots (all 10 routes
pixel-perfect), and Trivy HIGH/CRITICAL CVE count (6 -> 5, one npm
GHSA removed) all match or improve over baseline.
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
Adds a CI job that rebuilds the admin UI from source and fails if the
committed static export at litellm/proxy/_experimental/out/ has drifted
from what npm run build produces. This prevents silently shipping stale
UI bytes and is a prerequisite for the non_root Dockerfile streamlining
work, which will stage the UI from _experimental/out/ directly instead
of rebuilding it inside the image.
Also regenerates litellm/proxy/_experimental/out/ to match a fresh
npm run build (Node 20.20.2) — the committed tree had drifted from
source prior to this commit.
Co-authored-by: yuneng-jiang <yuneng-berri@users.noreply.github.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
* bump litellm-proxy-extras version to 0.4.67
* bump litellm-proxy-extras pin to 0.4.67 in litellm pyproject
* regenerate uv.lock for litellm-proxy-extras 0.4.67
* bump litellm-enterprise version to 0.1.38
* bump litellm-enterprise pin to 0.1.38 in litellm pyproject
* regenerate uv.lock for litellm-enterprise 0.1.38