* feat(proxy): push-based OTLP billable-request metering for enterprise deployments
Adds opt-in, license-gated metering that counts 2xx HTTP requests to LLM
inference, MCP, and A2A endpoints and exports them over mutual TLS to a global
OpenTelemetry Collector for request-based billing.
A pure ASGI middleware (BillableRequestMetricsMiddleware) classifies each
request by route and records one count per 2xx response via an injected
recorder. The recorder (BillingMetricsRecorder) owns a dedicated OTEL meter
provider and an OTLP/gRPC exporter authenticated with client certificates, kept
isolated from the global meter provider so a customer's own OTEL metrics are
untouched. The recorder is built only when a valid LITELLM_LICENSE is present
and the cert material is configured; otherwise the middleware is a transparent
pass-through.
Deployment identity rides on the mTLS client certificate rather than the
payload, so the secret license key is never sent as an attribute or header; only
the license org id travels as a resource attribute for cross-checking.
Resolves LIT-4089
* fix(proxy): align billable-request metering with the global collector
- switch the exporter to OTLP/HTTP with a TLS client certificate. The
collector front end terminates mutual TLS and validates the client cert
against our CA; server verification uses the system trust store, so the
CA env var is now an optional override for private collectors
- resolve the metrics recorder on the first request via a factory instead
of at import time, so deployments that provide the license and cert env
vars through the YAML config's environment_variables export correctly
- close the metering bypass: classify /images/edits, /images/variations,
/v1/messages, /v1/videos, video remix, /v1/ocr and Gemini generateContent
as billable, and gate LLM routes to POST so GET reads (list videos, fetch
a response) do not bill. Verified live: the collector count matches the
UI usage page successful_requests exactly, with failures excluded on both
sides
* fix(proxy): wrap enterprise billing import in try-except per code-quality gate
The check_unsafe_enterprise_import gate requires every import from an
enterprise-pathed module to be guarded. Annotate the factory with the
middleware's BillingRecorder protocol so no enterprise type import is
needed at type-check time
* chore: satisfy strict lint gates in billing modules
- builtin generics per UP006 (dict/tuple instead of typing.Dict/Tuple)
- noqa the deliberate blind catch that keeps metering from breaking startup
- sort proxy_server import blocks split by the guarded enterprise import
* fix(proxy): bill provider passthrough, search, and rag routes
Route-inventory audit against LiteLLMRoutes.llm_api_routes found more
SpendLogs-producing surfaces the classifier missed: provider passthrough
(/bedrock, /vertex-ai, /cohere and the rest of mapped_pass_through_routes),
/v1/search and vector-store search, and the rag ingest/query routes. All are
counted by the dashboard usage page, so missing them undercounts billing.
The passthrough prefix list is read from LiteLLMRoutes so new providers are
picked up without touching this module. /langfuse is excluded: it forwards
observability traffic and writes no SpendLogs row. Known limitation recorded
in the PR: /v1/realtime is a websocket flow the HTTP middleware does not see
* fix(proxy): bill MCP and A2A requests by protocol transport routes only
The billable-request classifier matched the whole /v1/mcp prefix, so
management and discovery reads such as GET /v1/mcp/tools and GET
/v1/mcp/server counted as billable MCP requests, while real MCP tool
calls on the /{server}/mcp and /toolset/{name}/mcp aliases were missed
because their route handlers rewrite the ASGI scope only after this
middleware has already classified the original path. Classify MCP by the
concrete transport surface (the /mcp streamable-HTTP and SSE sub-app plus
the single-segment server and toolset aliases) and exclude the /v1/mcp
management API. Apply the same shape to A2A, which had the identical
issue: only the /message/send invoke route bills, not /v1/a2a/discover or
the .well-known agent-card reads.
* fix(proxy): harden billable-request classification and recorder lifecycle
Exact-match Anthropic /v1/messages so OpenAI Assistants thread-message
routes no longer bill, add Google Interactions create routes, guard
recorder.record() so a broken exporter can never fail a served request,
lock lazy recorder resolution against concurrent first requests, and
disable metering on empty-string env config instead of accepting a
blank endpoint
* chore(ui): regenerate eslint metrics after staging merge
* docs(proxy): state the lower-bound billing contract in middleware comments
* fix(proxy): bill mcp-rest tool calls and bare a2a agent invokes
POST /mcp-rest/tools/call executes a tool and fires the same MCP spend
logging as the /mcp transport, and POST /a2a/{agent_id} is the JSON-RPC
invoke route whose method (message/send or message/stream) travels in
the body; both returned 2xx without being recorded
* fix(proxy): flush billable-request counts on proxy shutdown
PeriodicExportingMetricReader buffers up to one export interval of
counts; without a final flush every restart silently dropped them. The
factory registers the recorder it builds and proxy_shutdown_event pops
and flushes it, bounded by a 5s timeout so a dead collector cannot
stall shutdown
* fix(proxy): stop billing bare a2a task RPCs and close the shutdown race
POST /a2a/{agent_id} multiplexes JSON-RPC methods off the request body. Only
message/send and message/stream write a SpendLogs row; tasks/get, tasks/cancel
and the pushNotificationConfig RPCs are forwarded upstream and write none.
Classifying the bare path as billable counted those task RPCs and pushed the
metric above the dashboard's successful-request count. Since a path-only
classifier cannot read the body, the bare route no longer bills; the explicit
/message/send routes still do. Missing a bare-path invoke undercounts, which is
the only direction this metric is allowed to drift. The /mcp transport keeps
billing every method because its list path logs a SpendLogs row too.
The billing middleware also sat outside InFlightRequestsMiddleware, and it
records after the inner app returns. A request could therefore be counted as
drained while its record() had not yet run, letting proxy_shutdown_event flush
and stop the exporter underneath it. Registering it before the in-flight
tracker nests it inside, so wait_for_drain covers the record
* test(proxy): stub the OTLP exporter in the recorder-build test
test_premium_with_full_config_builds_recorder built a real MeterProvider, so
the shutdown flush resolved collector.example and opened a TLS connection from
a unit test. The exporter is now stubbed, and a getaddrinfo spy asserts nothing
resolves the collector host so the stub cannot be quietly dropped later
* fix(helm): truncate the helm.sh/chart label to 63 bytes
Kubernetes caps a label value at 63 bytes and .Chart.Version is unbounded. CI
publishes branch builds as 0.0.0-branch-<branch>-<sha>, so helm.sh/chart
rendered as a 64 byte value and the API server rejected every labeled resource
with "must be no more than 63 bytes", including the migrations Job. The
litellm-helm chart already guards this through a litellm.chart helper; this
adds the same helper here.
Swept the rest of the chart for label and name values built from unbounded
input. .Chart.Version appeared only in this label. The remaining candidates all
derive from .Release.Name, which helm itself caps at 53 characters, so they
cannot overflow; three of them are selector labels feeding immutable Deployment
matchLabels, where adding trunc would risk churn for no gain. They are left
alone deliberately.
Verified with a new helm-unittest suite, tests/chart_label_tests.yaml, which
overrides chart.version per test:
helm unittest -f 'tests/*.yaml' helm/litellm # 13 passed
helm unittest -f 'tests/*.yaml' helm/litellm-helm # 54 passed
The truncation cases fail against the previous helper. Reproduced the original
overflow by rendering with the real branch version and measuring the label:
helm template rel helm/litellm -f helm/litellm/tests/values/required.yaml \
| grep helm.sh/chart # 64 bytes before, 63 after
* feat(proxy): accept inline PEM for the billing-metrics mTLS credentials
LITELLM_BILLING_METRICS_CLIENT_CERT, _CLIENT_KEY and _CA_CERT took a filesystem
path. ECS injects Secrets Manager values as environment content and cannot mount
them as files, so a licensed deployment there could not turn metering on.
Each variable now takes either a path or the PEM itself. Inline PEM, detected by
the "-----BEGIN" prefix, is written once when the recorder is built into a 0700
temp dir as a 0600 file, and the config points at that path. The OTLP exporter
still only ever sees paths. A write failure disables metering through the
existing failure-as-None path rather than raising, and path-valued variables are
passed through untouched, so nothing changes for deployments that mount files.
The mixed case works too: mount the CA, inject the client credentials
* feat(helm): add first-class billingMetrics values to the componentized chart
Turning enterprise billable-request metering on meant hand-rolling the env vars
and the cert volume through gateway.extraEnv and gateway.volumes. This adds a
top-level billingMetrics block, off by default, consumed only by the gateway
since that is the component serving billable traffic.
When enabled it renders LITELLM_BILLING_METRICS_ENDPOINT plus the two cert paths
and mounts secretName read-only at /etc/litellm/billing-mtls. caSecretName is
optional and only needed for private collectors whose server certificate is not
on the public web PKI; when set it mounts at /etc/litellm/billing-mtls-ca and
adds the CA env var. exportIntervalMs is passed through only when set.
Enabling without secretName or with an empty endpoint fails the render with a
named message rather than producing a gateway that silently never exports.
The generic gateway.volumes, gateway.volumeMounts and gateway.extraEnv paths are
untouched and still compose with this, so existing overlays keep working.
The chart has no values.schema.json and no README, so there is nothing further to
update. Verified with a new helm-unittest suite:
helm unittest -f 'tests/*.yaml' helm/litellm # 23 passed
helm unittest -f 'tests/*.yaml' helm/litellm-helm # 54 passed
* feat(terraform): billing-metrics variables for the aws and gcp templates
* feat(helm): add billingMetrics values to the classic chart
The componentized chart just gained a first-class billingMetrics block; this
mirrors it in litellm-helm so enabling enterprise billable-request metering no
longer means hand-rolling the env vars and the cert volume through envVars and
volumes.
When enabled the proxy Deployment renders LITELLM_BILLING_METRICS_ENDPOINT plus
the two cert paths, and mounts secretName read-only at /etc/litellm/billing-mtls.
secretName defaults to litellm-billing-metrics-mtls, the conventional name, so
enabling the block is enough once that Secret exists. caSecretName is optional
and only needed for private collectors whose server certificate is not on the
public web PKI; when set it mounts at /etc/litellm/billing-mtls-ca and adds the
CA env var. exportIntervalMs is passed through only when set.
The env entries render after envVars and extraEnvVars, so a user-supplied
LITELLM_BILLING_METRICS_ENDPOINT cannot silently redirect the export under
Kubernetes last-wins duplicate-env semantics; this is the same ordering the
migrations Job relies on for DISABLE_SCHEMA_UPDATE.
Enabling with an emptied secretName or endpoint fails the render with a named
message rather than producing a proxy that silently never exports.
The generic volumes, volumeMounts, envVars and extraEnvVars paths are untouched
and still compose with this, so existing overlays keep working. The chart has no
values.schema.json; README parameters and a setup section are updated.
helm unittest -f 'tests/*.yaml' helm/litellm-helm # 68 passed (54 + 14 new)
helm lint helm/litellm-helm # 0 failed
* test(helm): pin that the migrations job never mounts the billing cert
The componentized chart's suite asserts the backend Deployment stays clear of the
billing wiring, since only the gateway serves billable traffic. The classic chart
has no backend, but it does have a second pod: the migrations Job, which renders
its own env from envVars and extraEnvVars. Nothing today wires the billing
include into it, and nothing stopped a future edit from doing so.
Asserts absence of the env, and that the Job grows no volumes or volumeMounts at
all. Both are notExists rather than notContains because the Job renders neither
key by default, so a notContains would fail on an unknown path instead of
checking the absence it looks like it is checking.
* fix(helm): meter the backend too, it serves the MCP transport
Scoping billingMetrics to the gateway was wrong. Applying each component's own
route allowlist to the proxy app shows the split is 75 billable routes on the
gateway and one on the backend: /{mcp_server_name}/mcp, the named-server MCP
transport, which writes a SpendLogs row on success. Metering only the gateway
would have silently dropped every MCP transport call from the counter, an
undercount proportional to a customer's MCP traffic.
The backend deployment now renders the same env and mounts the same read-only
cert secret. The migrations job still gets neither; it runs prisma and serves no
traffic, and a test pins that.
helm unittest -f 'tests/*.yaml' helm/litellm # 25 passed
helm unittest -f 'tests/*.yaml' helm/litellm-helm # 69 passed
This also aligns the chart with the terraform templates, which inject the
credentials into both components.
* fix(proxy): never log billing credential values when they fail to resolve
Accepting inline PEM turned the cert env vars into secret-bearing values, but
the disable warning still echoed them. A value that is neither a readable path
nor `-----BEGIN`-prefixed PEM, for example a key with a preamble or a malformed
secret, fell through to the path branch and was written to the proxy logs
verbatim, exposing the client certificate or private key to anyone who can read
them.
The warning now names the offending environment variables and tells the operator
what a valid value looks like, without ever printing one
* Revert "fix(helm): truncate the helm.sh/chart label to 63 bytes"
This reverts commit
|
||
|---|---|---|
| .. | ||
| examples/default | ||
| .terraform.lock.hcl | ||
| bootstrap.tf | ||
| cloudrun.tf | ||
| cloudsql.tf | ||
| gcs.tf | ||
| iam.tf | ||
| load_balancer.tf | ||
| locals.tf | ||
| network.tf | ||
| outputs.tf | ||
| README.md | ||
| redis.tf | ||
| secrets.tf | ||
| variables.tf | ||
| versions.tf | ||
LiteLLM on GCP (Cloud Run)
The button above opens the DeployStack installer in Cloud Shell, walks you through TUTORIAL.md, and runs terraform apply once you've answered the prompts. The rest of this README is the manual / advanced path.
Deploys the componentized LiteLLM proxy on GCP:
- VPC + Private Services Access range + a Serverless VPC Access connector so Cloud Run can reach private IPs
- Cloud SQL for PostgreSQL — primary instance + cross-zone read replica, password auth via Secret Manager
- Memorystore (Redis) for caching + rate limiting, private IP only
- GCS bucket — private, versioned, uniform IAM; exposed as
GCS_BUCKET_NAME - Secret Manager entries for
LITELLM_MASTER_KEYandDATABASE_PASSWORD - Cloud Run v2 services for
gateway(port 4000),backend(port 4001), andui(port 3000), all using a shared runtime service account - Cloud Run Job (
litellm-migrations) that runsprisma migrate deployfrom the dedicatedghcr.io/berriai/litellm-migrationsimage - External global HTTP(S) load balancer with serverless NEGs and a URL
map mirroring the helm-chart ingress path routing:
- LLM data-plane prefixes →
gateway - UI asset paths →
ui - Everything else →
backend
- LLM data-plane prefixes →
Image pulls
There are four images: litellm-gateway, litellm-backend, litellm-ui,
and litellm-migrations (slim image used only by the one-off Cloud Run
Job — runs prisma migrate deploy against the writer DB and exits).
Bump them together when bumping LiteLLM.
Required override. The image_registry default (ghcr.io/berriai)
does not work as-is — Cloud Run only accepts images from Artifact
Registry, [region.]gcr.io, or docker.io, and rejects ghcr.io URIs
at apply time. Every deploy (including HCP Terraform 1-click) must
supply either image_registry pointed at an Artifact Registry remote
repo backed by GHCR, or full per-component *_image URIs against
images you've already mirrored. The default is present only so
terraform plan succeeds during local iteration.
One-time setup (per project): create a remote repo and let Cloud Run pull through it.
gcloud artifacts repositories create litellm \
--repository-format=docker \
--location=us-central1 \
--mode=remote-repository \
--remote-repo-config-desc="GitHub Container Registry passthrough" \
--remote-docker-repo=https://ghcr.io
Then point the stack at it via image_registry:
image_registry = "us-central1-docker.pkg.dev/my-gcp-project/litellm/berriai"
image_tag = "v1.86.0-dev"
The four litellm-<component>:${image_tag} URIs are composed from those
two vars. Set gateway_image / backend_image / ui_image /
migrations_image only if you need a per-component override (custom
build, different tag).
Two further notes:
-
The runtime SAs the stack creates do not need
roles/artifactregistry.reader— Cloud Run pulls images using the per-project serverless agent (service-<project-num>@serverless-robot-prod.iam.gserviceaccount.com), not the runtime SA. -
For a fully air-gapped option, mirror the images into a regular AR repository instead of a remote repo:
for c in gateway backend ui migrations; do docker pull ghcr.io/berriai/litellm-$c:<tag> docker tag ghcr.io/berriai/litellm-$c:<tag> \ us-central1-docker.pkg.dev/$PROJECT/litellm/$c:<tag> docker push us-central1-docker.pkg.dev/$PROJECT/litellm/$c:<tag> donethen set
image_registry = "us-central1-docker.pkg.dev/$PROJECT/litellm"(drop the/berriaisuffix — the mirrored layout has no org segment).
Database authentication
LiteLLM's init_iam_db_url_from_env() mints AWS RDS tokens via boto3 —
it doesn't speak GCP IAM. To IAM-auth against Cloud SQL from Cloud Run you'd
need the Cloud SQL Auth Proxy as a sidecar, which complicates the service
spec. This stack therefore uses password authentication:
- A random password is generated and stored in Secret Manager
(
<name>-db-password). - Each Cloud Run service receives the password as
DATABASE_PASSWORDviavalue_source.secret_key_ref. - The container's entrypoint shim assembles
DATABASE_URL(andDATABASE_URL_READ_REPLICA) fromDATABASE_HOST/DATABASE_PASSWORDbefore exec'ing uvicorn — so the password never appears in the service spec or in logs.
If you need GCP-native IAM auth later, add cloud-sql-proxy as a sidecar
container under template.template.containers (Cloud Run v2 supports
multiple containers) and replace the password-based URL with the proxy's
Unix socket.
Configuring the proxy
proxy_config
Mirrors the helm chart's gateway.config.proxy_config. The map is
YAML-encoded and uploaded to a dedicated GCS bucket as config.yaml, then
mounted read-only into the gateway and backend at /etc/litellm via Cloud
Run v2's gcsfuse volume. CONFIG_FILE_PATH points at the mount path. A
hash of the YAML rides along as an env var so an edit to proxy_config
forces a new Cloud Run revision; without it the new file would sit in the
bucket unread until the next unrelated revision rollover. The migrations
job doesn't get the config (it only runs prisma migrate deploy).
proxy_config = {
model_list = [
{
model_name = "gpt-4o"
litellm_params = {
model = "openai/gpt-4o"
api_key = "os.environ/OPENAI_API_KEY"
}
},
]
general_settings = {
master_key = "os.environ/LITELLM_MASTER_KEY"
database_url = "os.environ/DATABASE_URL"
}
}
LiteLLM resolves os.environ/<NAME> references against the container
environment. Provider API keys belong in *_extra_secrets and are
referenced from the YAML by env-var name.
Extra env / secrets
Non-sensitive env vars:
gateway_extra_env = {
LANGFUSE_HOST = "https://us.cloud.langfuse.com"
}
Sensitive values — create the secret in Secret Manager first, then reference its resource ID:
echo -n "sk-proj-..." | gcloud secrets create openai-api-key --data-file=-
gateway_extra_secrets = {
OPENAI_API_KEY = "projects/my-gcp-project/secrets/openai-api-key"
}
The Cloud Run runtime SA auto-gains roles/secretmanager.secretAccessor on
every secret referenced. Pass the bare secret resource ID only —
projects/.../secrets/openai-api-key, never the version-suffixed form
projects/.../secrets/openai-api-key/versions/3. The Cloud Run
secret_key_ref binding and the stack's IAM secret_id grant both
reject the version suffix; version is always resolved as latest. If
you need a pinned version, edit local.gateway_extra_secret_kv in
cloudrun.tf directly to set version = "3" for the entry in question.
OpenTelemetry v2
OTel v2 (https://docs.litellm.ai/docs/observability/opentelemetry_v2) is
opt-in and gated entirely on otel_endpoint. Empty (default) and nothing
OTel-related lands in the container env. Set it and both gateway and
backend gain LITELLM_OTEL_V2=true plus the OTEL_* block, with
OTEL_SERVICE_NAME stamped per component (${tenant}-litellm-${env}-gateway
and -backend) so spans land tagged with the right hop. Any OTEL_* key
set in gateway_extra_env / backend_extra_env overrides the default for
that service (Cloud Run rejects duplicate env names, so the override is
predictable).
otel_endpoint = "https://otel.example.com:4318"
otel_exporter = "otlp_http" # or otlp_grpc
otel_environment_name = "prod" # default: var.env
otel_headers_secret = "projects/my-gcp-project/secrets/otel-headers"
OTEL_HEADERS is wired as a Secret Manager secret_key_ref since it
typically carries the collector's auth token; create the secret with the
literal header string, e.g. Authorization=Bearer <token>.
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT defaults to
no_content; flip otel_capture_message_content = "prompt_and_completion"
only after auditing what lands in the backend, since prompts and
completions are typically sensitive.
Behavior matches the AWS stack 1:1; the only naming differences are
otel_headers_secret (a Secret Manager resource ID) vs AWS's
otel_headers_secret_arn (a Secrets Manager ARN).
Enterprise billing metrics
License-gated request metering is opt-in and gated entirely on
billing_metrics_endpoint. Empty (default) and no billing env is added to
the container, so existing deployments are unchanged. Set it and both
gateway and backend export billable-request counts over OTLP/HTTP,
authenticating to the collector with the mTLS client certificate issued for
your deployment.
The proxy accepts the certificate, key, and CA bundle as either a file path
or literal PEM content. This stack takes the PEM, writes each one to its own
Secret Manager entry, grants the runtime service account
roles/secretmanager.secretAccessor on them, and injects them as Cloud Run
secret env vars LITELLM_BILLING_METRICS_CLIENT_CERT / _CLIENT_KEY (and
_CA_CERT when set), so no volume mount is needed.
billing_metrics_endpoint = "https://telemetry.litellm.ai/v1/metrics"
export TF_VAR_billing_metrics_client_cert_pem="$(cat client.crt)"
export TF_VAR_billing_metrics_client_key_pem="$(cat client.key)"
billing_metrics_ca_cert_pem is only for private or test collectors whose
CA is not in the system trust store; leave it empty against
telemetry.litellm.ai. Metering requires an enterprise license, so pair
this with litellm_license. To tune the export cadence, set
LITELLM_BILLING_METRICS_EXPORT_INTERVAL_MS through gateway_extra_env /
backend_extra_env
Behavior matches the AWS stack 1:1; the variable names are identical
Tenant deployment
Every resource the stack creates is named ${tenant}-litellm-${env} (or
that plus a per-resource suffix), so multiple tenants and multiple
environments coexist in the same project as long as the (tenant, env)
pair differs:
tenant |
env |
Example resource name |
|---|---|---|
acme |
stage |
acme-litellm-stage-gateway |
acme |
prod |
acme-litellm-prod-master-key |
globex |
dev |
globex-litellm-dev-license |
For a per-tenant instance via the example root, the only inputs that change are the tenant slug, env, and the two pre-issued secrets:
cd terraform/litellm/gcp/examples/default
export TF_VAR_litellm_master_key="sk-..." # the tenant's master key
export TF_VAR_litellm_license="lic-..." # their LITELLM_LICENSE
terraform apply \
-var "project_id=my-gcp-project" \
-var "region=us-central1" \
-var "tenant=acme" \
-var "env=stage"
To run many tenants from a single config, call the module with
for_each instead of one root per tenant — only possible because the
module declares no provider block (see "Using as a module").
Both litellm_master_key and litellm_license are optional:
- Omit
litellm_master_key→ the stack auto-generates a randomsk-…value (trial/dev path). - Omit
litellm_license→ no license secret is created and gateway/ backend run withoutLITELLM_LICENSE(OSS-only).
Use TF_VAR_* env vars rather than tfvars files for these — values
written to a tfvars file end up in terraform.tfstate and any committed
example files.
Quick start
cd terraform/litellm/gcp/examples/default
cp terraform.tfvars.example terraform.tfvars
# Edit: project, region, tenant, env, image_registry, proxy_config, gateway_extra_secrets.
terraform init
terraform apply
examples/default/ is a thin root that configures the google /
google-beta providers and calls the module (../../). It exposes a
curated variable surface; for advanced knobs (per-component
CPU/memory/instances, Cloud SQL tier/edition, Memorystore tier,
per-component image pins) set them on the module "litellm" block in
examples/default/main.tf, or call the module from your own config — see
"Using as a module" below.
That single apply provisions everything, runs the prisma schema migration via
the Cloud Run job (auto-triggered by bootstrap.tf), and only then starts the
gateway/backend services. When it returns, the stack is serving traffic.
terraform output lb_url
# UI login: admin / <master key>
gcloud secrets versions access latest --secret="$(terraform output -raw master_key_secret_id)"
The migration_run_command output is preserved for break-glass manual re-runs.
Prerequisite: gcloud must be authenticated (gcloud auth login) and the
required APIs must be enabled (run, sqladmin, redis, secretmanager,
vpcaccess, compute, servicenetworking, storage, artifactregistry).
TLS
terraform plan refuses to provision an HTTP-only LB by default — TLS
is the supported posture. Two paths:
Production / staging — set lb_domains:
terraform applyonce withallow_plaintext_lb = true(intentional chicken-and-egg escape hatch) to provision the LB and read the anycast IP fromterraform output -raw lb_ip.- Point each DNS name you want to serve from at that IP.
- Set
lb_domains = ["proxy.example.com"]and removeallow_plaintext_lb; re-apply.
Result: a 443 forwarding rule with a Google-managed cert covering each
listed domain; the 80 forwarding rule is rewritten to serve a permanent
301 redirect to HTTPS, so HTTP clients are automatically upgraded. The
managed cert sits in PROVISIONING for ~15-60 min on first apply until
DNS propagation completes — gcloud compute ssl-certificates describe <tenant>-litellm-<env>-cert shows the state.
Trial / dev — explicitly opt into HTTP-only:
Set allow_plaintext_lb = true and leave lb_domains = []. Without the
flag, plan fails with a clear error pointing at the precondition.
Intended for short-lived trial / dev stacks only.
Using as a module
The directory itself is a module with no provider block — the caller
owns provider config. You can call it directly with for_each (many
tenants from one config), count, depends_on, or providers configured
to impersonate a service account / target a different project:
provider "google" {
project = "my-gcp-project"
region = "us-central1"
}
provider "google-beta" {
project = "my-gcp-project"
region = "us-central1"
}
module "litellm" {
source = "github.com/BerriAI/litellm//terraform/litellm/gcp?ref=<tag>"
project = "my-gcp-project"
region = "us-central1"
tenant = "acme"
env = "prod"
# ...any of the inputs in variables.tf...
}
Both the default google and google-beta configs are inherited by the
module automatically through the call; declare both in the caller.
Labels: the module stamps its own litellm-stack and managed-by labels
onto every label-supporting resource (Cloud Run services and the
migrations job, Cloud SQL writer and reader, Memorystore, Secret Manager
entries, GCS buckets, the LB global address and forwarding rules) and
merges var.labels on top. Use the labels input for per-deployment
labels; mirrors the AWS stack's tags input.
for_each shares one provider config. The module's versions.tf declares
google / google-beta without configuration_aliases, so it only ever
receives the caller's single default (unaliased) google / google-beta
providers. That's deliberate — it keeps the one-command path simple — but it
means a for_each over the module runs every instance against the same
project, region, and credentials. Use for_each for many tenants in one
project (distinct tenant/env); it cannot fan out across projects or regions
on its own. To deploy into separate projects/regions, give each its own root
with its own provider config (one examples/default-style root per project),
or fork the module to add configuration_aliases and pass per-instance
providers = { ... }.
Storage and database retention
Two opt-in tripwires guard against accidental data loss on
terraform destroy:
cloudsql_deletion_protection(Cloud SQL writer + reader; defaulttrue) — destroy fails with a clear error rather than dropping the database.gcs_force_destroy(GCS bucket holding request log archives,/v1/filescontent, and the GCS cache backend; defaultfalse) —terraform destroyagainst a non-empty bucket fails.
Flip cloudsql_deletion_protection to false or gcs_force_destroy to
true only for ephemeral / CI stacks where you accept losing the data.
Redis encryption
Memorystore runs with transit_encryption_mode = "SERVER_AUTHENTICATION",
so the proxy connects via rediss://. The instance's self-signed CA cert
(server_ca_certs[0].cert) is shipped to gateway + backend as
REDIS_CA_PEM_B64; their entrypoint shell decodes it to /tmp/redis-ca.pem
before uvicorn starts and points REDIS_SSL_CA_CERTS at that path. No
extra config needed — but if you ever swap Memorystore for an external
Redis, override REDIS_HOST/REDIS_PORT and either drop these env vars
or point them at your own CA.
Files
| File | What's in it |
|---|---|
versions.tf |
Terraform + required_providers constraints (module declares no provider config) |
examples/default/ |
Thin root: google / google-beta providers + a call to the module. The one-command deploy path. |
variables.tf |
All input variables |
locals.tf |
Path-prefix lists (mirror of helm/.../ingress.yaml) + proxy_config helpers |
network.tf |
VPC, subnet, PSA range, Serverless VPC connector |
secrets.tf |
Secret Manager entries + random master_key |
cloudsql.tf |
Cloud SQL writer + read replica + app user + password secret |
redis.tf |
Memorystore Redis (private IP) |
gcs.tf |
GCS bucket + objectAdmin binding |
iam.tf |
Runtime SA + Cloud SQL client + Secret Manager accessor |
cloudrun.tf |
3 Cloud Run services + Cloud Run Job for migrations |
load_balancer.tf |
External HTTPS LB, serverless NEGs, URL map for path routing |
outputs.tf |
LB IP, service URLs, secret IDs, migration execute command |