litellm/terraform/provider
devin-ai-integration[bot] 0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
..
docs feat(terraform): add display_name to litellm_model resource and model data sources (#42987) 2026-09-24 17:06:30 -05:00
examples feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00
litellm fix(provider): accept 2xx status codes in unified_access_group create (#42461) 2026-09-28 13:54:37 -07:00
tools feat(anthropic): workload identity federation and pluggable identity sources (#44448) 2026-10-03 17:08:30 -07:00
.gitignore feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00
.goreleaser.yml docs(terraform/provider): the provider now ships at the LiteLLM version (#37912) 2026-08-21 22:09:38 -07:00
CHANGELOG.md fix(provider): accept 2xx status codes in unified_access_group create (#42461) 2026-09-28 13:54:37 -07:00
go.mod chore(deps): bump grpc and golang.org/x modules in the terraform provider 2026-08-04 16:09:33 -07:00
go.sum chore(deps): bump grpc and golang.org/x modules in the terraform provider 2026-08-04 16:09:33 -07:00
LICENSE feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00
main.go feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00
Makefile feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00
README.md fix(terraform): send litellm_key model_max_budget as per-model BudgetConfig objects (#40450) 2026-09-09 19:46:26 -07:00
RELEASING.md ci: follow the default branch in development tooling 2026-09-07 14:34:45 -07:00
terraform-registry-manifest.json feat(terraform): vendor terraform-provider-litellm as source of truth with endpoint drift CI (#32241) 2026-07-07 09:16:59 -07:00

LiteLLM Terraform Provider

This Terraform provider allows you to manage LiteLLM resources through Infrastructure as Code. It provides support for managing models, teams, team members, API keys, users, organizations, budgets, tags, projects, guardrails, prompts, agents, search tools, access groups, fallbacks, MCP servers, credentials and vector stores via the LiteLLM REST API, along with read-only data sources for each of them.

Source of truth

This directory (terraform/provider/ in BerriAI/litellm) is the source of truth for the provider. BerriAI/terraform-provider-litellm is a thin release mirror that the public Terraform Registry ingests from; do not open PRs there. Changes land here, where CI builds the provider, runs its tests, and statically audits every endpoint the provider calls against the proxy's generated OpenAPI schema (tools/endpointaudit/), so the provider cannot drift from the LiteLLM API silently. The same audit runs in reverse as a coverage gate: every management endpoint in the schema must be covered by a resource or data source, or carry a documented entry in tools/endpointaudit/coverage_allowlist.txt, and stale allowlist entries fail CI. Releases are published by mirroring this directory into the split repo and tagging it, which triggers the goreleaser workflow there (see RELEASING.md)

Versioning

The provider version is the LiteLLM version. Every LiteLLM release (dev, rc and stable) publishes the provider at the same version as the proxy, built from the same commit, so 1.99.0 of the provider is the one that shipped with 1.99.0 of the proxy and was audited against that proxy's API. Pin the provider to the line your proxy runs:

version = "~> 1.99.0"

Pre-release versions (1.99.0-rc.1, 1.99.0-dev.1) are published too; Terraform only selects one when it is pinned exactly.

Versions 0.1.0 through 0.4.0 predate this scheme and sit on their own line. They stay in the registry, but a ~> 0.4 constraint will never pick up another release: re-pin to the LiteLLM version to keep receiving updates.

Features

  • Manage LiteLLM model configurations
  • Associate models with specific teams
  • Create and manage teams
  • Configure team members and their permissions
  • Set usage limits and budgets
  • Control access to specific models
  • Specify model modes (e.g., completion, embedding, image generation)
  • Manage API keys with fine-grained controls
  • Support for reasoning effort configuration in the model resource

Requirements

Using the Provider

To use the LiteLLM provider in your Terraform configuration, you need to declare it in the terraform block:

terraform {
  required_providers {
    litellm = {
      source  = "BerriAI/litellm"
      version = "~> 1.99.0" # the LiteLLM version your proxy runs
    }
  }
}

provider "litellm" {
  api_base = var.litellm_api_base
  api_key  = var.litellm_api_key
}

Then, you can use the provider to manage LiteLLM resources. Here's an example of creating a model configuration:

resource "litellm_model" "gpt4" {
  model_name          = "gpt-4-proxy"
  custom_llm_provider = "openai"
  model_api_key       = var.openai_api_key
  model_api_base      = "https://api.openai.com/v1"
  base_model          = "gpt-4"
  tier                = "paid"
  mode                = "chat"
  reasoning_effort    = "medium"  # Optional: "low", "medium", or "high"
  
  input_cost_per_million_tokens  = 30.0
  output_cost_per_million_tokens = 60.0
}

For full details on the litellm_model resource, see the model resource documentation.

Here's an example of creating an API key with various options:

resource "litellm_key" "example_key" {
  models               = ["gpt-4", "claude-3.5-sonnet"]
  max_budget           = 100.0
  user_id              = "user123"
  team_id              = "team456"
  max_parallel_requests = 5
  tpm_limit            = 1000
  rpm_limit            = 60
  budget_duration      = "monthly"
  key_alias            = "prod-key-1"
  duration             = "30d"
  metadata             = {
    environment = "production"
  }
  allowed_cache_controls = ["no-cache", "max-age=3600"]
  soft_budget          = 80.0
  aliases              = {
    "gpt-4" = "gpt4"
  }
  config               = {
    default_model = "gpt-4"
  }
  permissions          = {
    can_create_keys = "true"
  }
  model_max_budget     = jsonencode({
    "gpt-4" = {
      budget_limit = 50.0
      time_period  = "30d"
    }
  })
  model_rpm_limit      = {
    "claude-3.5-sonnet" = 30
  }
  model_tpm_limit      = {
    "gpt-4" = 500
  }
  guardrails           = ["content_filter", "token_limit"]
  blocked              = false
  tags                 = ["production", "api"]
}

The litellm_key resource supports the following options:

  • models: List of allowed models for this key
  • max_budget: Maximum budget for the key
  • user_id and team_id: Associate the key with a user and team
  • max_parallel_requests: Limit concurrent requests
  • tpm_limit and rpm_limit: Set tokens and requests per minute limits
  • budget_duration: Specify budget duration (e.g., "monthly", "weekly")
  • key_alias: Set a friendly name for the key
  • duration: Set the key's validity period
  • metadata: Add custom metadata to the key
  • allowed_cache_controls: Specify allowed cache control directives
  • soft_budget: Set a soft budget limit
  • aliases: Define model aliases
  • config: Set configuration options
  • permissions: Specify key permissions
  • model_max_budget, model_rpm_limit, model_tpm_limit: Set per-model limits
  • guardrails: Apply specific guardrails to the key
  • blocked: Flag to block/unblock the key
  • tags: Add tags for organization and filtering

For full details on the litellm_key resource, see the key resource documentation.

Available Resources

  • litellm_model: Manage model configurations. Documentation
  • litellm_team: Manage teams. Documentation
  • litellm_team_member: Manage team members. Documentation
  • litellm_team_member_add: Add multiple members to teams. Documentation
  • litellm_key: Manage API keys. Documentation
  • litellm_mcp_server: Manage MCP (Model Context Protocol) servers. Documentation
  • litellm_credential: Manage credentials for secure authentication. Documentation
  • litellm_vector_store: Manage vector stores for embeddings and RAG. Documentation
  • litellm_jwt_key_mapping: Map JWT claim values to virtual keys for per-client budgets and limits. Documentation

Available Data Sources

  • litellm_credential: Retrieve information about existing credentials. Documentation
  • litellm_vector_store: Retrieve information about existing vector stores. Documentation

Development

Project Structure

The project is organized as follows:

terraform-provider-litellm/
├── litellm/
│   ├── provider.go
│   ├── resource_model.go
│   ├── resource_model_crud.go
│   ├── resource_team.go
│   ├── resource_team_member.go
│   ├── resource_key.go
│   ├── resource_key_utils.go
│   ├── types.go
│   └── utils.go
├── main.go
├── go.mod
├── go.sum
├── Makefile
└── ...

Building the Provider

  1. Clone the repository:
git clone https://github.com/your-username/terraform-provider-litellm.git
  1. Enter the repository directory:
cd terraform-provider-litellm
  1. Build and install the provider:
make install

Development Commands

The Makefile provides several useful commands for development:

  • make build: Builds the provider
  • make install: Builds and installs the provider
  • make test: Runs the test suite
  • make fmt: Formats the code
  • make vet: Runs go vet
  • make lint: Runs golangci-lint
  • make clean: Removes build artifacts and installed provider

Testing

To run the tests:

make test

Contributing

Contributions are welcome! Please read our contributing guidelines first.

License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Notes

  • Always use environment variables or secure secret management solutions to handle sensitive information like API keys and AWS credentials.
  • Refer to the comprehensive documentation in the docs/ directory for detailed usage examples and configuration options.
  • Keep the provider version in step with the LiteLLM version your proxy runs; see Versioning.
  • The provider now supports AWS cross-account access with aws_session_name and aws_role_name parameters in the model resource.
  • All example configurations have been consolidated into the documentation for better organization and maintenance.