* feat(proxy): offload spend tracking to a pod-local spend worker sidecar py-spy on the gateway showed the post-response _PROXY_track_cost_callback, spend-log and DBSpendUpdateWriter work running on the inference workers' event loop, so a DB or Redis stall backed up the request path. When LITELLM_SPEND_WORKER_ENABLED=true, _ProxyDBLogger serializes one compact typed SpendEvent per success and hands it to a SpendEventProducer that ships it over a unix socket (default) or loopback-only TCP to a sidecar started as `python -m gateway.spend_worker`. The sidecar runs the unchanged _ProxyDBLogger pipeline against the pod's PgBouncer (pooled_database_url). When the sidecar is unreachable, the buffer is full, or the gateway shuts down with events still queued or in flight, the producer applies LITELLM_SPEND_WORKER_ON_UNAVAILABLE (fallback in-process, or drop). The sidecar half-closes producers on SIGTERM and drains, the producer treats EOF as unavailable, and the gateway flushes buffered spend counters on shutdown. The sidecar honors LITELLM_LOG so its writes are visible in its own process log. Helm: both charts gain an opt-in spend-worker sidecar container sharing an emptyDir socket dir, and the componentized chart's HPA uses a ContainerResource CPU metric scoped to the gateway container so sidecar CPU does not drive inference scaling. * feat(terraform): opt-in spend-worker sidecar for the AWS and GCP gateway stacks Adds spend_worker_* inputs to both modules. On ECS Fargate the sidecar is a second, non-essential container in the gateway task; on Cloud Run it is a second container in the gateway service. Both listen on loopback TCP, share the gateway's DB/Redis/secret env, and set LITELLM_JOB_ROLE=spend_worker. Disabled by default. Plan-only tests cover both, and the terraform CI workflow now runs the gcp module too Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): retrieve a completed batch in the in-process spend path test The base now defers cost tracking for batches that are still in flight, so an in_progress batch never reaches update_database Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): rename the spend worker sidecar to collector Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): run the collector from the installed litellm package and finish in-flight fallbacks on shutdown The sidecar command becomes python -m litellm.proxy.collector so the classic image, whose runtime stage copies only the installed package, can run it. The module now assembles DATABASE_URL and the pod-local pgbouncer URL itself, replacing gateway/collector.py The componentized collector sidecar inherits gateway.volumeMounts so custom CA mounts reach it. SpendEventProducer shields an in-progress fallback from the writer task cancellation so close() no longer loses an event already handed to the in-process pipeline Helpers used across modules (address_argument, should_store_prompts_and_responses_in_spend_logs, flush_spend_counters_on_shutdown) become public so the change adds no reportPrivateUsage errors Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci(terraform): drop the gcp job duplicated by the aws/gcp matrix Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(collector): keep metrics env off the classic sidecar and reject shared loopback ports The classic chart no longer hands PROMETHEUS_METRICS_PORT and the billing metrics env to the collector container, and gives it the same /.npm scratch mount as the proxy on a read-only root. AWS and GCP now refuse a plan where the spend collector and the metrics sidecar bind the same loopback port. A regression test drives a sidecar crash mid-stream on asyncio and uvloop and checks no event is billed by both the sidecar and the in-process fallback; the producer docstring spells out why a failed drain() cannot double count Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy): format pooled_database_url after the pgbouncer rebase Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): keep the cache-hit preset key and survive dead producers on collector drain Cache hits updated the logging object after the early return, so the offloaded spend event carried preset_cache_key=None and the collector re-hashed reconstructed kwargs. Also guard write_eof() against producer transports uvloop already closed so one dead connection cannot abort the drain Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(terraform): keep the gcp collector port off the metrics sidecar health port Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): collector connects to Postgres directly under IAM or Entra token auth Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): mark the collector's DATABASE_URL as pooled when it uses the pod's pgbouncer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| charts | ||
| ci | ||
| templates | ||
| tests | ||
| .helmignore | ||
| Chart.lock | ||
| Chart.yaml | ||
| README.md | ||
| values.yaml | ||
Helm Chart for LiteLLM
Important
This is community maintained, Please make an issue if you run into a bug We recommend using Docker or Kubernetes for production deployments
Prerequisites
- Kubernetes 1.21+
- Helm 3.8.0+
If db.deployStandalone is used:
- PV provisioner support in the underlying infrastructure
If db.useStackgresOperator is used (not yet implemented):
- The Stackgres Operator must already be installed in the Kubernetes Cluster. This chart will not install the operator if it is missing.
Parameters
LiteLLM Proxy Deployment Settings
| Name | Description | Value |
|---|---|---|
replicaCount |
The number of LiteLLM Proxy pods to be deployed | 1 |
masterkeySecretName |
The name of the Kubernetes Secret that contains the Master API Key for LiteLLM. If not specified, use the generated secret name. | N/A |
masterkeySecretKey |
The key within the Kubernetes Secret that contains the Master API Key for LiteLLM. If not specified, use masterkey as the key. |
N/A |
masterkey |
The Master API Key for LiteLLM. If not specified, a random key in the sk-... format is generated on first install and reused on upgrades. |
N/A |
environmentSecrets |
An optional array of Secret object names. The keys and values in these secrets will be presented to the LiteLLM proxy pod as environment variables. See below for an example Secret object. | [] |
environmentConfigMaps |
An optional array of ConfigMap object names. The keys and values in these configmaps will be presented to the LiteLLM proxy pod as environment variables. See below for an example Secret object. | [] |
image.repository |
LiteLLM Proxy image repository | ghcr.io/berriai/litellm |
image.pullPolicy |
LiteLLM Proxy image pull policy | IfNotPresent |
image.tag |
Overrides the image tag whose default the latest version of LiteLLM at the time this chart was published. | "" |
imagePullSecrets |
Registry credentials for the LiteLLM and initContainer images. | [] |
serviceAccount.create |
Whether or not to create a Kubernetes Service Account for this deployment. The default is false because LiteLLM has no need to access the Kubernetes API. |
false |
service.type |
Kubernetes Service type (e.g. LoadBalancer, ClusterIP, etc.) |
ClusterIP |
service.port |
TCP port that the Kubernetes Service will listen on. Also the TCP port within the Pod that the proxy will listen on. | 4000 |
livenessProbe.* |
Liveness probe settings for the LiteLLM container (path, periodSeconds, timeoutSeconds, thresholds, and initial delay). |
See values.yaml |
readinessProbe.* |
Readiness probe settings for the LiteLLM container (path, periodSeconds, timeoutSeconds, thresholds, and initial delay). |
See values.yaml |
startupProbe.* |
Startup probe settings for the LiteLLM container (path, periodSeconds, timeoutSeconds, thresholds, and initial delay). |
See values.yaml |
resources.* |
CPU/memory requests and limits for the LiteLLM container. Unset by default; production deployments should set 1 CPU and 4Gi of memory per worker. | {} |
service.loadBalancerClass |
Optional LoadBalancer implementation class (only used when service.type is LoadBalancer) |
"" |
ingress.labels |
Additional labels for the Ingress resource | {} |
ingress.* |
See values.yaml for example settings | N/A |
proxyConfigMap.create |
When true, render a ConfigMap from .Values.proxy_config and mount it. |
true |
proxyConfigMap.name |
When create=false, name of the existing ConfigMap to mount. |
"" |
proxyConfigMap.key |
Key in the ConfigMap that contains the proxy config file. | "config.yaml" |
proxy_config.* |
See values.yaml for default settings. Rendered into the ConfigMap’s config.yaml only when proxyConfigMap.create=true. See example_config_yaml for configuration examples. |
N/A |
extraContainers[] |
An array of additional containers to be deployed as sidecars alongside the LiteLLM Proxy. | |
pdb.enabled |
Enable a PodDisruptionBudget for the LiteLLM proxy Deployment | false |
pdb.minAvailable |
Minimum number/percentage of pods that must be available during voluntary disruptions (choose one of minAvailable/maxUnavailable) | null |
pdb.maxUnavailable |
Maximum number/percentage of pods that can be unavailable during voluntary disruptions (choose one of minAvailable/maxUnavailable) | null |
pdb.annotations |
Extra metadata annotations to add to the PDB | {} |
pdb.labels |
Extra metadata labels to add to the PDB | {} |
| billingMetrics.enabled | Enable enterprise billable-request metering. Requires an enterprise license. | false |
| billingMetrics.endpoint | Collector that the billable-request counter is pushed to. | https://telemetry.litellm.ai |
| billingMetrics.secretName | Name of an existing Secret holding the mTLS client certificate, under the keys tls.crt and tls.key. | litellm-billing-metrics-mtls |
| billingMetrics.caSecretName | Name of an existing Secret holding a CA bundle under the key ca.crt. Only needed for a private or test collector whose server certificate is not on the public web PKI. | "" |
| billingMetrics.exportIntervalMs | How often the counter is pushed, in milliseconds. The proxy defaults to 60000 when unset. | "" |
Example proxy_config ConfigMap from values (default):
proxyConfigMap:
create: true
key: "config.yaml"
proxy_config:
general_settings:
master_key: os.environ/PROXY_MASTER_KEY
model_list:
- model_name: gpt-3.5-turbo
litellm_params:
model: gpt-3.5-turbo
api_key: eXaMpLeOnLy
Example using existing proxyConfigMap instead of creating it:
proxyConfigMap:
create: false
name: my-litellm-config
key: config.yaml
# proxy_config is ignored in this mode
Example environmentSecrets Secret
apiVersion: v1
kind: Secret
metadata:
name: litellm-envsecrets
data:
AZURE_OPENAI_API_KEY: TXlTZWN1cmVLM3k=
type: Opaque
Enterprise billable-request metering
Enterprise licenses meter billable requests by pushing a counter to LiteLLM's collector over mutual TLS. The chart does not create the client certificate; it mounts one you already hold, read-only, so the private key is never exposed through the environment. Create the Secret under the name the chart expects, then turn the block on:
kubectl create secret tls litellm-billing-metrics-mtls --cert=client.crt --key=client.key
billingMetrics:
enabled: true
Set billingMetrics.caSecretName only when the collector is a private or test one whose server certificate is not on the public web PKI; the production collector needs no CA override. The chart fails the render rather than deploying a proxy that silently never exports, so a missing secretName or an emptied endpoint surfaces at helm install time.
Database Settings
| Name | Description | Value |
|---|---|---|
db.useExisting |
Use an existing Postgres database. A Kubernetes Secret object must exist that contains credentials for connecting to the database. An example secret object definition is provided below. | false |
db.endpoint |
If db.useExisting is true, this is the IP, Hostname or Service Name of the Postgres server to connect to. |
localhost |
db.database |
If db.useExisting is true, the name of the existing database to connect to. |
litellm |
db.url |
If db.useExisting is true, the connection url of the existing database to connect to can be overwritten with this value. |
postgresql://$(DATABASE_USERNAME):$(DATABASE_PASSWORD)@$(DATABASE_HOST)/$(DATABASE_NAME) |
db.secret.name |
If db.useExisting is true, the name of the Kubernetes Secret that contains credentials. |
postgres |
db.secret.usernameKey |
If db.useExisting is true, the name of the key within the Kubernetes Secret that holds the username for authenticating with the Postgres instance. |
username |
db.secret.passwordKey |
If db.useExisting is true, the name of the key within the Kubernetes Secret that holds the password associates with the above user. |
password |
db.useStackgresOperator |
Not yet implemented. | false |
db.deployStandalone |
Deploy a standalone, single instance deployment of Postgres, using the Bitnami postgresql chart. This is useful for getting started but doesn't provide HA or (by default) data backups. | true |
postgresql.* |
If db.deployStandalone is true, configuration passed to the Bitnami postgresql chart. See the Bitnami Documentation for full configuration details. See values.yaml for the default configuration. |
See values.yaml |
postgresql.auth.* |
If db.deployStandalone is true, care should be taken to ensure the default password and postgres-password values are NOT used. |
NoTaGrEaTpAsSwOrD |
postgresql.image.* |
If db.deployStandalone is true, the image for the bundled Postgres. Pinned to a docker.io/bitnamilegacy build because Bitnami retired the versioned tags under docker.io/bitnami. |
bitnamilegacy/postgresql:16.2.0-debian-12-r6 |
redis.image.* |
If redis.enabled is true, the image for the bundled Redis. Pinned to a docker.io/bitnamilegacy build for the same reason. |
bitnamilegacy/redis:7.2.4-debian-12-r9 |
Bundled Postgres image
Bitnami removed the versioned tags from docker.io/bitnami and republished the archived builds under docker.io/bitnamilegacy, so the image defaults that ship inside the postgresql and redis subcharts no longer pull. The chart pins both to the bitnamilegacy copies of the exact builds those subchart versions were released with, which keeps the on-disk data directory layout unchanged for existing installs.
Keep postgresql.image.tag pinned. docker.io/bitnami/postgresql still publishes a floating latest, and pointing the bundled Postgres at a different major version starts the server against a data directory it cannot read (database files are incompatible with server). There is no in-place way back, so crossing a major version means dumping the database with the old image and restoring it into the new one. The chart refuses to render when the tag is empty or latest.
Those images no longer receive updates. For anything beyond getting started, run Postgres outside the chart and point at it with db.useExisting.
Example Postgres db.useExisting Secret
apiVersion: v1
kind: Secret
metadata:
name: postgres
data:
# Password for the "postgres" user
postgres-password: <some secure password, base64 encoded>
username: litellm
password: <some secure password, base64 encoded>
type: Opaque
Examples for environmentSecrets and environemntConfigMaps
# Use config map for not-secret configuration data
apiVersion: v1
kind: ConfigMap
metadata:
name: litellm-env-configmap
data:
SOME_KEY: someValue
ANOTHER_KEY: anotherValue
# Use secrets for things which are actually secret like API keys, credentials, etc
# Base64 encode the values stored in a Kubernetes Secret: $ pbpaste | base64 | pbcopy
# The --decode flag is convenient: $ pbpaste | base64 --decode
apiVersion: v1
kind: Secret
metadata:
name: litellm-env-secret
type: Opaque
data:
SOME_PASSWORD: cDZbUGVXeU5e0ZW # base64 encoded
ANOTHER_PASSWORD: AAZbUGVXeU5e0ZB # base64 encoded
Source: GitHub Gist from troyharvey
Migration Job Settings
The migration job supports both ArgoCD and Helm hooks to ensure database migrations run at the appropriate time during deployments.
| Name | Description | Value |
|---|---|---|
migrationJob.enabled |
Enable or disable the schema migration Job | true |
migrationJob.backoffLimit |
Backoff limit for Job restarts | 4 |
migrationJob.ttlSecondsAfterFinished |
TTL for completed migration jobs | 120 |
migrationJob.annotations |
Additional annotations for the migration job pod | {} |
migrationJob.extraContainers |
Additional containers to run alongside the migration job | [] |
migrationJob.hooks.argocd.enabled |
Enable ArgoCD hooks for the migration job (uses PreSync hook with BeforeHookCreation delete policy) | true |
migrationJob.hooks.helm.enabled |
Enable Helm hooks for the migration job (uses pre-install,pre-upgrade hooks with before-hook-creation delete policy) | false |
migrationJob.hooks.helm.weight |
Helm hook execution order (lower weights executed first). Optional - defaults to "1" if not specified. | N/A |
Accessing the Admin UI
When browsing to the URL published per the settings in ingress.*, you will
be prompted for Admin Configuration. The Proxy Endpoint is the internal
(from the litellm pod's perspective) URL published by the <RELEASE>-litellm
Kubernetes Service. If the deployment uses the default settings for this
service, the Proxy Endpoint should be set to http://<RELEASE>-litellm:4000.
The Proxy Key is the value specified for masterkey or, if a masterkey
was not provided to the helm command line, the masterkey is a randomly
generated string in the sk-... format stored in the <RELEASE>-litellm-masterkey Kubernetes Secret.
The key is generated once on the first install; later helm upgrade runs reuse the
value already in that Secret, so upgrading never rotates the master key.
kubectl -n litellm get secret <RELEASE>-litellm-masterkey -o jsonpath="{.data.masterkey}"
Admin UI Limitations
At the time of writing, the Admin UI is unable to add models. This is because
it would need to update the config.yaml file which is a exposed ConfigMap, and
therefore, read-only. This is a limitation of this helm chart, not the Admin UI
itself.