litellm/terraform/litellm/README.md
Yassin Kortam 1cff02f50e
refactor: convert AWS and GCP Terraform stacks into reusable modules … (#28103)
* refactor: convert AWS and GCP Terraform stacks into reusable modules with examples/default entry point

- Remove `provider` blocks from both AWS and GCP stack roots so the modules
  can be consumed with `count`, `for_each`, `depends_on`, assumed-role or
  aliased providers — patterns that are forbidden when a module owns its own
  provider configuration
- Add `examples/default/` thin-root wrappers for both stacks that wire the
  provider (AWS) / providers (google + google-beta) and call the module with
  a curated variable surface, preserving the one-command deploy experience
- Move `terraform.tfvars.example` files into `examples/default/` alongside
  the new roots; update example comments to reflect the curated variable surface
- Thread `local.tags` (containing `litellm:stack`, `managed-by`, and
  `var.tags`) explicitly onto every taggable AWS resource since the module no
  longer controls the provider's `default_tags`; GCP resource labels already
  flow through the module's `labels` input
- Add `examples/default/variables.tf` and `outputs.tf` for both stacks,
  exposing the most-used knobs and re-exporting all module outputs
- Commit provider lock files for both examples so `terraform init` is
  reproducible without a network fetch
- Update top-level and per-stack READMEs to document the module-first design,
  the `for_each` multi-tenant pattern, and the `examples/default/` quick-start path

* docs(terraform): address review — state-migration guide, tag dedupe, for_each note

- Add 'Migrating an existing deployment' section to AWS & GCP READMEs
  documenting the required terraform state mv step (resource addresses now
  gain a module.litellm. prefix under the examples/default root)
- Remove redundant managed-by tag from the AWS example providers.tf;
  reserve default_tags there for org-wide tags only
- Document the for_each single-provider limitation for GCP (no
  configuration_aliases) in the README and example main.tf

Resolves LIT-3504

* docs(terraform/gcp): note expected SSL cert replacement in state-migration guide

The managed SSL cert is named with a hash of lb_domains, so TLS-enabled
stacks that migrated from the old un-hashed name will see one
create_before_destroy cert replacement after terraform state mv — not a
clean 'No changes'. Document that this single replacement is expected and
safe.

* docs(terraform): drop state-migration guides

The AWS/GCP stacks have never been published, so there are no existing
deployments to migrate from the old root-module layout. Remove the
'Migrating an existing deployment' sections from both READMEs.

* docs(terraform): call out image-registry override required for GCP 1-click

The GCP stack's default image_registry points at ghcr.io, which Cloud
Run won't authenticate against, so any real deploy (HCP Terraform
no-code or otherwise) must override it. Document that as a hard
requirement on the GCP README rather than a side note, and add a
top-level HCP Terraform 1-click section enumerating the required
inputs per stack and the migration-task caveat for HCP-hosted runners.

* feat(terraform/aws): mount proxy_config from S3 and wire OpenTelemetry v2

proxy_config

Drop the inline LITELLM_PROXY_CONFIG_B64 env var. Upload the YAML to S3
at config/litellm-config.yaml; gateway and backend container entrypoints
download it to /tmp/litellm-config.yaml via boto3 before exec'ing
uvicorn. The S3 object etag is wired into the task definition so a
config edit produces a new task-def revision and a rolling redeploy. The
existing s3_access policy already grants the task role s3:GetObject on
this bucket, so no IAM changes were needed for the mount itself.

OpenTelemetry v2

New variables otel_endpoint, otel_exporter, otel_service_name, and
otel_headers_secret_arn. Setting otel_endpoint to a non-empty value adds
LITELLM_OTEL_V2=true plus OTEL_EXPORTER / OTEL_ENDPOINT /
OTEL_SERVICE_NAME / OTEL_ENVIRONMENT_NAME to the shared env block; an
optional Secrets Manager ARN backs OTEL_HEADERS for collectors that need
an auth header. Execution role auto-gains GetSecretValue on that ARN.
Empty endpoint = nothing added, so existing deployments are unchanged.

* feat(terraform/gcp): add DeployStack one-click installer

Wires up a Cloud Shell "Open in Cloud Shell" badge backed by the
GoogleCloudPlatform DeployStack flow so examples/default can be
installed from a click in the README without a local terraform setup.

- examples/default/deploystack.json drives project/region collection
  plus prompts for tenant, env, image_tag, and allow_plaintext_lb.
  Complex inputs (proxy_config, *_extra_secrets, lb_domains) and
  sensitive vars (litellm_master_key, litellm_license, ui_password)
  stay tfvars / env only so they never land in a committed file.
- examples/default/TUTORIAL.md is a Cloud Shell walkthrough that
  enables required APIs, creates the GHCR-passthrough Artifact
  Registry repo, optionally exports the TF_VAR_* secrets, runs
  `deploystack install`, and shows how to fetch the master key plus
  migrate from plaintext LB to TLS.
- Renames var.project to var.project_id across the module and the
  examples/default wrapper to match the variable DeployStack injects
  from `collect_project: true`. Breaking rename for anyone with a
  `project = ...` line in terraform.tfvars; the fix is one line.

* feat(terraform/gcp): mount proxy_config from GCS and wire OpenTelemetry v2

proxy_config

Drop the inline LITELLM_PROXY_CONFIG_B64 env var and the python-decode
startup fragment. Upload the YAML to a dedicated GCS bucket as
config.yaml, then mount it read-only into the gateway and backend at
/etc/litellm via Cloud Run v2's gcsfuse volume. CONFIG_FILE_PATH points
at the mount; an md5 of the YAML rides along as PROXY_CONFIG_HASH so a
config-only edit forces a new Cloud Run revision (gcsfuse only surfaces
new objects on container restart, so without the hash an updated
proxy_config would sit in the bucket unread).

The config bucket is separate from the data-plane bucket so the runtime
SA can hold objectViewer here (read-only at runtime) while keeping
objectAdmin on the data-plane bucket. Both bucket and IAM binding are
gated on proxy_config != {}; an empty config skips bucket creation and
mounts nothing.

OpenTelemetry v2

LITELLM_OTEL_V2=true is now wired into shared_env_kv unconditionally so
both the gateway and backend boot with the integration enabled. It's
dormant until otel_endpoint is non-empty; setting it injects
OTEL_EXPORTER / OTEL_ENDPOINT / OTEL_ENVIRONMENT_NAME plus a
per-component OTEL_SERVICE_NAME (\${tenant}-litellm-\${env}-{gateway,backend})
so spans land tagged with the right hop. otel_headers_secret takes a
Secret Manager resource ID for OTEL_HEADERS (collector auth); the
runtime SA auto-gains roles/secretmanager.secretAccessor on it.
otel_capture_message_content defaults to no_content matching the litellm
default. Any OTEL_* key set in *_extra_env wins over the defaults so
Cloud Run doesn't reject the apply on the duplicate-env-name check.

* refactor(terraform): make AWS and GCP stacks behave identically

Bring both modules to the same surface and the same runtime behavior so
swapping clouds (or reading either README) is symmetric.

Labels and tags. GCP previously stamped var.labels onto only the two GCS
buckets, leaving Cloud Run, Cloud SQL, Memorystore, Secret Manager, and
the LB resources unlabeled; the variable description claimed full
coverage. Now the module computes local.labels (litellm-stack +
managed-by + var.labels, mirroring AWS's local.tags) and threads it onto
every label-supporting resource: Cloud Run services and the migrations
job, Cloud SQL writer and reader (via user_labels), Memorystore, Secret
Manager entries (master_key, license, ui_password, db_password), both
GCS buckets, the global LB address, and the http/https forwarding rules.
GCP keys use 'litellm-stack' instead of AWS's 'litellm:stack' because
GCP label keys forbid colons; var.labels now defaults to {}.

OpenTelemetry v2 is opt-in on both stacks. AWS already gated everything
on otel_endpoint; GCP previously stamped LITELLM_OTEL_V2=true into
shared_env unconditionally and only ungated the OTEL_* block. Both
stacks now do the same thing: leave otel_endpoint empty and nothing
OTel-related lands in the container env; set it and gateway and backend
get LITELLM_OTEL_V2=true plus OTEL_EXPORTER, OTEL_ENDPOINT,
OTEL_ENVIRONMENT_NAME, OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT,
and a per-component OTEL_SERVICE_NAME (${tenant}-litellm-${env}-gateway
or -backend) so spans land tagged with the right hop. AWS picks up the
richer GCP surface: otel_environment_name (defaults to var.env),
otel_capture_message_content (defaults to no_content), and *_extra_env
override filtering so a caller-set OTEL_* key wins over the default for
that service (ECS allows duplicates, but the filter gives the same
predictable last-wins shape Cloud Run enforces). var.otel_service_name
on AWS is gone, replaced by the per-component naming.

uvicorn workers. GCP gains gateway_num_workers, matching AWS; threads
into the gateway args as --workers ${var.gateway_num_workers}.

Docs reflect the parity: each README's OTel section, the GCP 'Using as
a module' Labels paragraph, and a new feature-parity table in the
top-level README that lays out the AWS/GCP input mapping side by side.

* fix(terraform/aws): expose skip_final_snapshot through the default example

The example wrapper already exposed `s3_force_destroy` so ephemeral / CI
stacks could destroy the S3 bucket without manual cleanup, but the matching
Aurora knob (`skip_final_snapshot`) was hidden behind the module surface.
That meant a `terraform destroy` on a trial stack still produced a
`<cluster>-final-<short-sha>` snapshot, with no opt-out short of editing the
module call.

Adds `var.skip_final_snapshot` to the example (default `false`, preserving
the data-loss tripwire) and threads it through to the module input,
mirroring the existing `s3_force_destroy` pattern. Documented alongside it
in the tfvars example.

Verified by deploying the example end-to-end against a clean AWS account
(VPC + Aurora w/ IAM auth + Redis + ALB + 3 ECS services), confirming all
services reach steady state and the data plane serves traffic, then running
`terraform destroy` with `skip_final_snapshot = true` to a clean teardown
(93 destroyed, no Aurora snapshot left behind, no leftover billable
resources).

---------

Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
Co-authored-by: yassin-berriai <yassin.kortam@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-06-06 12:57:44 -07:00

17 KiB

LiteLLM Terraform stacks

Two self-contained, reusable Terraform modules that deploy the componentized LiteLLM proxy — the gateway, backend, and UI as three independent containers (see helm/litellm/ for the canonical chart with the same split).

Each module declares no provider block of its own, so it can be called with count / for_each / depends_on and the caller controls region, assume-role / impersonation, aliases, and default_tags. A ready-to-run root that wires the provider lives at <stack>/examples/default/ — that's the one-command deploy path. To embed a stack in your own config, call the module by source:

module "litellm" {
  source = "github.com/BerriAI/litellm//terraform/litellm/aws?ref=<tag>"
  # ... inputs ...
}
Stack Compute Database (writer + reader) Cache Object store Public entrypoint
aws/ ECS Fargate Aurora Postgres (IAM auth) ElastiCache S3 Application LB
gcp/ Cloud Run Cloud SQL Postgres (password auth) Memorystore GCS External HTTPS LB

Each stack creates its own VPC and managed data stores — from <stack>/examples/default/, drop in a tfvars file and run terraform apply. Both stacks support a typed proxy_config input (mirrors helm/litellm's gateway.config.proxy_config) and per-component extra env vars / secret-manager refs.

Components

The proxy is split into three deployables:

Component Default image Port Role
gateway ghcr.io/berriai/litellm-gateway:main-stable 4000 LLM data plane (/v1/chat/completions, /v1/embeddings, …)
backend ghcr.io/berriai/litellm-backend:main-stable 4001 Management API (/key/*, /user/*, /team/*, /model/*, …)
ui ghcr.io/berriai/litellm-ui:main-stable 3000 Static Next.js dashboard served by nginx

The load balancer routes gateway path prefixes (mirrored verbatim from gateway/routes/allowlist.py) to the gateway, UI asset paths (/, /litellm-asset-prefix/*, /_next/*, /favicon.ico) to the UI, and everything else to the backend.

Architecture

AWS (terraform/litellm/aws/)

                        ┌───────────────────────────────────────┐
                        │            Public Internet            │
                        └─────────────────┬─────────────────────┘
                                          │ HTTP/80
                          ┌───────────────▼───────────────┐
                          │   Application Load Balancer   │
                          │   (path-routing listener)     │
                          └─┬─────────────┬─────────────┬─┘
                            │             │             │
            UI assets, /    │  /v1/chat,  │   /key/*    │
            /_next/*, …     │  /v1/embed, │   /user/*   │
                            │  …          │   …         │
              ┌─────────────▼───┐  ┌──────▼──────┐  ┌───▼──────────────┐
              │    ECS Service  │  │ ECS Service │  │   ECS Service    │
              │       (ui)      │  │  (gateway)  │  │    (backend)     │
              │   Fargate :3000 │  │ Fargate:4000│  │  Fargate :4001   │
              └─────────────────┘  └──────┬──────┘  └────────┬─────────┘
                                          │                  │
              ┌─── private subnets (one per AZ) ──────────────────────┐
              │                                                       │
              │   ┌────────────────────────┐    ┌────────────────┐   │
              │   │  Aurora Postgres       │    │  ElastiCache   │   │
              │   │  cluster (IAM auth)    │    │  Redis (1 node)│   │
              │   │  ┌───────┐  ┌───────┐  │    └────────────────┘   │
              │   │  │writer │  │reader │  │                         │
              │   │  └───────┘  └───────┘  │    ┌────────────────┐   │
              │   └────────────────────────┘    │  S3 bucket     │   │
              │                                  │  (versioned)   │   │
              │   ┌────────────────────────┐    └────────────────┘   │
              │   │  Secrets Manager       │                         │
              │   │  • LITELLM_MASTER_KEY  │    ┌────────────────┐   │
              │   │  • DB master password  │    │ One-off ECS    │   │
              │   │  • user-supplied API   │    │ task: prisma   │   │
              │   │    keys (referenced)   │    │ migrate deploy │   │
              │   └────────────────────────┘    └────────────────┘   │
              │                                                       │
              └─── VPC ───────────────────────────────────────────────┘
                          │ NAT gateway in one public subnet
                          ▼
                    egress to LLM providers

GCP (terraform/litellm/gcp/)

                        ┌───────────────────────────────────────┐
                        │            Public Internet            │
                        └─────────────────┬─────────────────────┘
                                          │ HTTP/80
                          ┌───────────────▼───────────────┐
                          │ External HTTPS Load Balancer  │
                          │   (global, URL map routing)   │
                          └─┬─────────────┬─────────────┬─┘
                            │             │             │
                            │ Serverless NEGs (one per service)
                            │             │             │
              ┌─────────────▼───┐  ┌──────▼──────┐  ┌───▼──────────────┐
              │   Cloud Run     │  │  Cloud Run  │  │    Cloud Run     │
              │      (ui)       │  │  (gateway)  │  │    (backend)     │
              │      :3000      │  │   :4000     │  │      :4001       │
              └─────────────────┘  └──────┬──────┘  └────────┬─────────┘
                                          │                  │
                                          │ Serverless VPC Access connector
              ┌─── VPC (private services access range) ──────────────────┐
              │                                                          │
              │   ┌────────────────────────┐    ┌──────────────────┐    │
              │   │  Cloud SQL Postgres    │    │  Memorystore     │    │
              │   │  ┌───────┐  ┌───────┐  │    │  Redis           │    │
              │   │  │writer │  │reader │  │    └──────────────────┘    │
              │   │  └───────┘  └───────┘  │                            │
              │   └────────────────────────┘    ┌──────────────────┐    │
              │                                  │  GCS bucket      │    │
              │   ┌────────────────────────┐    │  (versioned)     │    │
              │   │  Secret Manager        │    └──────────────────┘    │
              │   │  • LITELLM_MASTER_KEY  │                            │
              │   │  • DB password         │    ┌──────────────────┐    │
              │   │  • user-supplied API   │    │ Cloud Run Job:   │    │
              │   │    keys (referenced)   │    │ prisma migrate   │    │
              │   └────────────────────────┘    │ deploy           │    │
              │                                  └──────────────────┘    │
              └──────────────────────────────────────────────────────────┘

Images

Both stacks take per-component image references as variables. The defaults point at the public ghcr.io/berriai/litellm-<component>:main-stable images, so the stack is runnable end-to-end without pre-flight setup — pin to a specific tag for production:

  • AWS can pull from any registry the task execution role can reach. The role gets AmazonECSTaskExecutionRolePolicy attached, which grants ECR pull permissions for repositories in the same account.

  • GCP Cloud Run can only pull from Artifact Registry or gcr.io-style registries. To use images hosted elsewhere, mirror them into Artifact Registry first.

Migrations

LiteLLM's proxy runs prisma migrate deploy at startup, but on first apply the gateway/backend can race the empty database. Both stacks expose a one-off migration task that runs python litellm/proxy/prisma_migration.py against the backend image:

  • AWS: an aws_ecs_task_definition (litellm-migrations). Run with aws ecs run-task — the command is printed in terraform output.
  • GCP: a google_cloud_run_v2_job (litellm-migrations). Run with gcloud run jobs execute — the command is printed in terraform output.

Run the migration job once after the first terraform apply and before the gateway/backend services start serving traffic.

Feature parity between stacks

The two modules expose the same conceptual surface; concrete inputs differ only where the underlying cloud forces it.

Capability AWS input(s) GCP input(s)
Tenant + env naming tenant, env tenant, env
Pre-shared master key / license litellm_master_key, litellm_license litellm_master_key, litellm_license
UI admin password ui_password ui_password
Per-deployment tags / labels tags (map(string)) labels (map(string))
TLS posture acm_certificate_arn, allow_plaintext_alb lb_domains, allow_plaintext_lb
Force destroy of object store s3_force_destroy gcs_force_destroy
Database deletion protection skip_final_snapshot cloudsql_deletion_protection
proxy_config (typed YAML map) proxy_config proxy_config
Extra plain env per component gateway_extra_env, backend_extra_env gateway_extra_env, backend_extra_env
Extra secret-backed env gateway_extra_secrets, backend_extra_secrets (ARNs) gateway_extra_secrets, backend_extra_secrets (resource IDs)
Uvicorn --workers on gateway gateway_num_workers gateway_num_workers
OpenTelemetry v2 (opt-in) otel_endpoint, otel_exporter, otel_environment_name, otel_capture_message_content, otel_headers_secret_arn otel_endpoint, otel_exporter, otel_environment_name, otel_capture_message_content, otel_headers_secret

Each module stamps its own stack-identity tag (litellm:stack on AWS, litellm-stack on GCP — GCP label keys forbid colons) plus managed-by = "terraform" onto every taggable / labelable resource and merges var.tags / var.labels on top. Provider default_tags on AWS merge on top of all of these.

OTel is opt-in on both clouds: leave otel_endpoint empty and nothing OTel-related is added to the container env; set it and both gateway and backend get LITELLM_OTEL_V2=true plus the full OTEL_* block, with OTEL_SERVICE_NAME stamped per component (<tenant>-litellm-<env>-gateway and -backend). Any OTEL_* key set in gateway_extra_env / backend_extra_env wins for that service.

What's not included

  • TLS certificates / custom domains. Both stacks expose plain-HTTP load balancers; bring your own ACM cert (AWS) or managed cert (GCP) and wire it into the LB resource.
  • Remote state backends. Default local state — add an s3 or gcs backend block to versions.tf when graduating to a team environment.
  • Observability beyond the cloud provider's defaults (CloudWatch logs on AWS, Cloud Logging on GCP). Wire your own Prometheus / Datadog / Langfuse via the *_extra_env variables, or turn on OTel v2 (see the parity table above).

HCP Terraform no-code (1-click) deploy

Both stacks are publishable as no-code modules in HCP Terraform's private registry. The end-user flow is: open the no-code launch URL, fill in a few inputs, hit Create workspace, and HCP runs plan/apply against your cloud account using a variable-set of credentials (static keys or dynamic-credentials OIDC).

Required overrides the launcher must supply per stack:

  • AWS (terraform/litellm/aws): region, azs, tenant, env. The image vars (gateway_image, backend_image, ui_image, migrations_image) can be left at their defaults — the GHCR images are anonymous-readable and ECS Fargate pulls them without extra credentials.

  • GCP (terraform/litellm/gcp): project, tenant, env, and one of:

    • image_registry pointed at an Artifact Registry remote repository backed by https://ghcr.io (e.g. us-central1-docker.pkg.dev/<project>/litellm/berriai), so Cloud Run pulls the four upstream litellm-* images through it; or
    • all four per-component *_image URIs pointing at images mirrored into a regular Artifact Registry repo.

    The defaults (ghcr.io/berriai) cause Cloud Run admission to reject the service spec — Cloud Run only authenticates against Artifact Registry, [region.]gcr.io, or docker.io. See terraform/litellm/gcp/README.md#image-pulls for the gcloud artifacts repositories create … --mode=remote-repository command that sets up the passthrough repo (one-time, per project).

What still requires a manual step regardless of HCP no-code:

  • The one-off migration task. The stacks auto-run it via local-exec during terraform apply, but that requires the aws / gcloud CLI on the runner. HCP-hosted runners don't have them; use an HCP agent pool with a custom image that includes the relevant CLI, or run the command printed in the migration_run_command output by hand after the first apply.