litellm/docker
devin-ai-integration[bot] 6f5ca84f69
feat(prometheus): cap series per metric for every labeled metric (#44420)
* feat(prometheus): cap series per metric for every labeled metric

Add prometheus_metrics_max_series_per_metric: per metric and per worker process, the first N label
sets keep a series of their own. Counters and histograms record every later label set on one series
whose labels are all "other", so totals stay exact, and gauges skip it. The cap holds with multiple
workers because it never needs to remove a series.

Add prometheus_metrics_ttl_seconds: a series idle for that long is removed and its slot is freed.
The prometheus client cannot remove a series in multi-process mode, so the TTL is ignored there with
a startup warning.

Both settings are off by default. The end_user caps are unchanged.

* fix(prometheus): share the series cap across workers of one proxy instance

Workers writing to one PROMETHEUS_MULTIPROC_DIR now agree on which label
sets get a series through an append-only admissions file per metric, so a
merged scrape stays at the cap plus `other` instead of growing with every
worker and every worker restart. The two fallback counters now pass their
label names as a keyword so the cap and prometheus_exclude_labels apply to
them, admission and child creation happen under one lock, the test fixture
restores the shared registry, and the `other` label value lives in
constants.py.

* test(prometheus): check emitted labels instead of wrapper types, close the admission match

The exclude-labels test now emits through the spend and provider budget
metrics and checks the scrape keeps all their labels. The admission match
arms end in assert_never so the match is exhaustive.

* fix(prometheus): return the exhaustive-match fallback so every admission arm returns

* fix(prometheus): pick the series tracker with isinstance so every path of _admits returns

* fix(prometheus): skip an admissions line a worker could only write part of

* fix(prometheus): frame each admissions record with newlines so a cut-off record cannot swallow the next

A record a worker could only write part of used to merge with the next worker's record, and both were skipped for one request. Each record is now written between two newlines, so the fragment is a line of its own. The clock fixture in the series tests starts from a constant instead of reading the real clock

* fix(prometheus): ignore a non-positive series cap or TTL with a warning instead of failing the logger

A cap or TTL of 0 or less raised at logger init. The proxy logs that as a non-blocking error and keeps serving, so the result was a running proxy with no Prometheus metrics at all. The setting is now ignored with a startup warning naming it, the same rule the end_user cap already follows for a non-positive value

* fix(prometheus): start the series cap over on a one-worker restart and audit it live

A proxy with one worker and an operator-set PROMETHEUS_MULTIPROC_DIR now drops litellm's admission files at boot, so a restart frees every slot there the way it already does with several workers. A cap or TTL that is not a number greater than 0 (a bool, a non-numeric string, an empty value) is ignored with the startup warning instead of breaking the logger

The integration cells drive the cap on every endpoint through the OpenAI and Anthropic SDKs and raw httpx, streaming and not, plus gauges, cache hits, failures, both workers of one instance, the TTL on one worker and its warning on two, ignored settings, excluded labels on the fallback counters, a null cap, /config/update, a concurrent burst scraped mid-flight, a provider outage, a killed worker, and restarts with one and two workers

* fix(prometheus): wipe an operator-set multiprocess directory on a one-worker boot too

* fix(prometheus): leave the multiprocess directory alone on a setup-only run

A run with --skip_server_startup starts no worker, so it no longer creates or
wipes PROMETHEUS_MULTIPROC_DIR. Wiping there deleted the samples of a proxy
already running against the same directory

* fix(prometheus): free the capped series slots when a gateway or backend container restarts

The component image entrypoint starts uvicorn without the proxy CLI and wiped only the .db sample files at container start, so the admitted-series files of the previous container survived an in-place restart. Every label set seen after the restart was then counted on `other` once the previous container had filled the cap

* test(prometheus): cover a setup-only run and a gateway image restart under the cap

Two integration cells from the audit: a `--skip_server_startup` run pointed at a live
two-worker proxy's operator directory leaves its samples alone, and the gateway image
(`docker/component_entrypoint.sh` running `python -m gateway.launch`) restarted on a kept
PROMETHEUS_MULTIPROC_DIR starts the cap over. The burst cells now wait for every counter
they assert on, since the request and failure counters of one call increment at different
points of the logging callback

* test(prometheus): prove the cap reaches the fallback counters in the X1 cell

* fix(prometheus): ignore a cleanup interval that is not a number of at least 0

A string or negative prometheus_metrics_cleanup_interval_seconds reached the
series tracker unvalidated, so the first labeled emit with a TTL on raised
TypeError inside the callback and recorded no series. The interval is now
validated the way the cap and the TTL are: an invalid value is ignored with a
warning and the default 60 seconds applies. The I2 integration cell drives a
string interval through a live proxy and reads the warning from its log

* test(prometheus): give the restart cells the boot budget of their siblings

C4 and C5 boot two proxies each and hit the file's 240 s budget on a loaded
box; C3 and D1 already carry 420 s

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 01:28:44 +00:00
..
build_from_pip fix(docker): match USE_DDTRACE case-insensitively and route build_from_pip through prod_entrypoint.sh (#39344) 2026-09-03 15:10:33 -07:00
tests fix(docker.non_root): use numeric UID 65534 for K8s runAsNonRoot (#26268) 2026-04-22 18:00:04 -07:00
.env.example docs: stop advertising sk-1234 as the master key in shipped configs and examples 2026-09-19 12:59:48 -07:00
build_admin_ui.sh chore(build): move the Admin UI toolchain to Node 24 (#35801) 2026-08-04 12:36:07 -07:00
component_entrypoint.sh feat(prometheus): cap series per metric for every labeled metric (#44420) 2026-10-06 01:28:44 +00:00
docker-compose.quickstart.yml feat(docker): one-command quickstart that starts the gateway, Postgres, and the admin UI (#43673) 2026-10-02 09:37:37 -07:00
docker-compose.tracing.yml fix(lens): align source setup with available worker images (#44476) 2026-10-03 19:42:38 -07:00
Dockerfile.database feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
Dockerfile.non_root feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
entrypoint.sh build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
install_auto_router.sh build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
prod_entrypoint.sh feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444) 2026-10-03 17:28:18 -07:00
README.md feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
tracing-config.yaml fix(tracing): unify ClickHouse storage configuration (#43941) 2026-10-02 16:31:26 +00:00

Docker Development Guide

This guide provides instructions for building and running the LiteLLM application using Docker and Docker Compose.

Just want to run LiteLLM? This guide builds from source. To run the published image instead, use docker-compose.quickstart.yml in this directory — the two-service stack (gateway + Postgres) that the Docker quickstart documents:

curl -sSLO https://github.com/BerriAI/litellm/raw/main/docker/docker-compose.quickstart.yml
printf 'LITELLM_MASTER_KEY=sk-%s\nLITELLM_SALT_KEY=sk-%s\n' "$(openssl rand -hex 32)" "$(openssl rand -hex 32)" > .env
docker compose -f docker-compose.quickstart.yml up -d

Prerequisites

  • Docker
  • Docker Compose

Building and Running the Application

To build and run the application, you will use the docker-compose.yml file located in the root of the project. This file is configured to use the Dockerfile.non_root for a secure, non-root container environment.

1. Set the Master Key

The application requires a LITELLM_MASTER_KEY for signing and validating tokens. You must set this key as an environment variable before running the application.

Create a .env file in the root of the project and add the following line:

LITELLM_MASTER_KEY=your-secret-key

Replace your-secret-key with a strong, randomly generated secret.

2. Build and Run the Containers

Once you have set the LITELLM_MASTER_KEY, you can build and run the containers using the following command:

docker compose up -d --build

This command will:

  • Build the Docker image using Dockerfile.non_root.
  • Start the litellm, litellm_db, and prometheus services in detached mode (-d).
  • The --build flag ensures that the image is rebuilt if there are any changes to the Dockerfile or the application code.

3. Verifying the Application is Running

You can check the status of the running containers with the following command:

docker compose ps

To view the logs of the litellm container, run:

docker compose logs -f litellm

4. Stopping the Application

To stop the running containers, use the following command:

docker compose down

Embedded LiteAdmin MCP

Source builds containing embedded LiteAdmin MCP can serve it at /admin/mcp on the existing LiteLLM port. This capability is unreleased. Keep your existing database, master key, and proxy configuration, then add these settings to the serving container's environment:

LITELLM_ENABLE_ADMIN_MCP=true
LITELLM_LICENSE="your-enterprise-license"
PROXY_BASE_URL=https://gateway.example.com

For the unified source deployment described above, put them in its .env file and rebuild:

docker compose up -d --build

In componentized deployments, set the flag and license on the backend container and route /admin/mcp to the backend service. The gateway component excludes this endpoint. The unified, database, non-root, and backend image builds bundle the connector

Hosting is disabled by default. Opting in requires a valid base Enterprise license; an unlicensed opt-in or invalid flag value prevents startup. Enabling it reserves /admin, so rename any MCP server alias called admin first

With native key authentication, connect with a personal proxy-admin bearer key. When enable_oauth2_proxy_auth is enabled, the existing trusted-proxy identity headers select the user instead; the MCP bearer is required by the connector but does not select the native user. The resolved user must have the stored proxy_admin role, and trusted_proxy_ranges applies to the original caller's direct peer

Embedded responses default to full; selecting LITELLM_ADMIN_RESPONSE_VIEW=compact requires subsequent saved-result reads to reach the same worker process, including within a multi-worker pod

See the LiteAdmin MCP guide for client configuration, tool restrictions, and verification

Hardened / Offline Testing

To ensure changes are safe for non-root, read-only root filesystems and restricted egress, always validate with the hardened compose file:

docker compose -f docker-compose.yml -f docker-compose.hardened.yml build --no-cache
docker compose -f docker-compose.yml -f docker-compose.hardened.yml up -d

This setup:

  • Builds from docker/Dockerfile.non_root with Prisma engines and Node toolchain baked into the image.
  • Runs the proxy as a non-root user with a read-only rootfs and only writable tmpfs mounts:
    • /app/cache (Prisma/NPM cache; backing PRISMA_BINARY_CACHE_DIR, NPM_CONFIG_CACHE, XDG_CACHE_HOME)
    • /app/migrations (Prisma migration workspace; backing LITELLM_MIGRATION_DIR)
  • Pre-builds and serves the admin UI from read-only paths:
    • /var/lib/litellm/ui (pre-restructured Next.js UI with .litellm_ui_ready marker)
    • /var/lib/litellm/assets (UI logos and assets)
  • Routes all outbound traffic through a local Squid proxy that denies egress, so Prisma migrations must use the cached CLI and engines.

You should also verify offline Prisma behaviour with:

docker run --rm --network none --entrypoint prisma ghcr.io/berriai/litellm:main-stable --version

This command should succeed (showing engine versions) even with --network none, confirming that Prisma binaries are available without network access.

Troubleshooting

  • build_admin_ui.sh: not found: This error can occur if the Docker build context is not set correctly. Ensure that you are running the docker-compose command from the root of the project.
  • Master key is not initialized: This error means the LITELLM_MASTER_KEY environment variable is not set. Make sure you have created a .env file in the project root with the LITELLM_MASTER_KEY defined.