Find a file
Ishaan Jaff 7697b1c397
fix(sso): direct PKCE token exchange + Redis wiring for multi-instance SSO (#22923)
* fix(sso): add direct PKCE token exchange and Redis cache wiring for multi-instance SSO

When PKCE is enabled, bypass fastapi-sso and perform direct token exchange so
code_verifier is correctly included. Store PKCE verifiers as dict in cache
for proper JSON serialization in Redis. Wire user_api_key_cache to Redis when
available so PKCE verifiers are shared across ECS tasks/pods.

Also adds clearer error messages when PKCE is required but not configured.

* refactor(sso): extract PKCE token exchange into SSOAuthenticationHandler methods

- Move import httpx/jwt to module level (top of file, not inside function)
- Extract inline PKCE token exchange + userinfo logic into two static methods:
  _pkce_token_exchange() and _get_pkce_userinfo()
- get_generic_sso_response PKCE path is now a single method call
- Fix double-logging in except block for non-PKCE errors
- Use %-style log formatting (no f-strings in log calls)

* fix: address greptile review feedback

- Fix access_token missing in PKCE path: read from combined_response directly
  instead of generic_sso.access_token (which is only set by verify_and_process)
- Fix PKCE error hint firing when PKCE is already enabled: only show
  'set GENERIC_CLIENT_USE_PKCE=true' advice when code_verifier was absent
- Fix unguarded KeyError on access_token: check for error field in HTTP 200
  responses before accessing token_response['access_token']
- Fix silent empty userinfo: raise ProxyException when both userinfo endpoint
  and id_token fallback produce no user data
- Fix backward-incompatible Redis wiring: only attach Redis to user_api_key_cache
  when GENERIC_CLIENT_USE_PKCE=true, preserving existing in-memory behaviour

* fix: address second round of greptile review feedback

- Fix PKCE error hint: check env var directly (not code_verifier presence) to
  distinguish 'PKCE not configured' from 'PKCE enabled but cache miss'
- Fix misleading Redis TTL comment in proxy_server.py

* fix: address third round of greptile review feedback

- Fix CRITICAL log firing on every non-PKCE callback: only log when PKCE is enabled
- Remove unused pkce_env_value intermediate variable
- Prefer reusing redis_usage_cache over creating separate RedisCache instance
  (avoids losing advanced connection options like SSL, timeouts, db)

* fix: address fourth round of greptile review feedback

- Strip OAuth token credentials from response_convertor input to prevent
  access_token/id_token appearing in restricted-group error messages
- Reuse single httpx.AsyncClient for both token exchange and userinfo requests
  to avoid a second TCP/TLS handshake per SSO callback
- Revert Redis wiring to user_api_key_cache: PKCE code already uses
  redis_usage_cache directly; wiring would route all API-key lookups through
  Redis unnecessarily. Add startup warning instead when PKCE+Redis mismatch.
- Move _OAUTH_TOKEN_FIELDS to module level

* fix remaining PKCE test assertion for dict-format verifier storage

* sanitize PKCE cache log to not expose verifier content

* address greptile review feedback (greploop iteration 3)

* address greptile review feedback (greploop iteration 4)

* address greptile review feedback (greploop iteration 5)

* simplify _get_pkce_userinfo: remove shared-client complexity, use async with directly

* address greptile review feedback (greploop iteration 6)

* address greptile review feedback (greploop iteration 7)

* address greptile review feedback (greploop iteration 8)

* address greptile review feedback (greploop iteration 9)

* address greptile review feedback (greploop iteration 10)

* address greptile review feedback (greploop iteration 11)

* address greptile review feedback (greploop iteration 12)

* fix misleading comment on user_api_key_cache TTL line

* address greptile review feedback (greploop iteration 13)

* address greptile review feedback (greploop iteration 14)

* address greptile review feedback (greploop iteration 15)

* address greptile review feedback (greploop iteration 16)

* address greptile review feedback (greploop iteration 17)

* address greptile review feedback (greploop iteration 18)

* address greptile review feedback (greploop iteration 19)

* address greptile review feedback (greploop iteration 20)

* address greptile review feedback (greploop iteration 21)

* address greptile review feedback (greploop iteration 22)

* address greptile review feedback (greploop iteration 23)

* address greptile review feedback (greploop iteration 24)

* address greptile review feedback (greploop iteration 25)

* address greptile review feedback (greploop iteration 26)

* address greptile review feedback (greploop iteration 27)

* address greptile review feedback (greploop iteration 28)

* address greptile review feedback (greploop iteration 29)

* address greptile review feedback (greploop iteration 30)

* address greptile review feedback (greploop iteration 31)

* address greptile review feedback (greploop iteration 32)

* address greptile review feedback (greploop iteration 33)

* address greptile review feedback (greploop iteration 34)

* address greptile review feedback (greploop iteration 35)

* address greptile review feedback (greploop iteration 37)

- read GENERIC_CLIENT_USE_PKCE env var once in prepare_token_exchange_parameters
- include actual decode error in jwt.decode failure exception message
- add GENERIC_CLIENT_USE_PKCE=true to no-state regression test

* defer PKCE verifier deletion until after all downstream processing

Move _delete_pkce_verifier to after response_convertor and
process_sso_jwt_access_token complete. If JWT processing raises,
the verifier stays in cache so the user can retry without restarting
the full OAuth flow.

* address greptile review feedback (greploop iteration 38)

- fix strict-mode cache miss error message to differentiate
  cross-instance routing failures (Redis configured) from single-instance
  issues (TTL expiry, pod restart) when only in-memory cache is available
- add comment above _get_pkce_userinfo call explaining that bearer
  credentials are always sourced from token_response in the merge step

* fix null JSON response body in _pkce_token_exchange

- Guard against HTTP 200 with body null: response.json() returns None
  for JSON null, and calling .get() on None raises AttributeError.
  Now raises a clean ProxyException with a clear error message.
- Fix misleading userinfo warning: was always saying "empty dict" but
  also fires for JSON null responses; updated to say "empty or null".
- Add HTTP status code assertion to cache miss test.

* address greptile review feedback (greploop iteration 39)

- fix credential leakage: directly assign received_response from
  combined_response instead of relying on nonlocal mutation; Pyright
  was flagging the old guard as unreachable, meaning credential stripping
  might not execute — now it always runs unconditionally
- add test for legacy plain-string cache format backward compat branch
- add test for HTTP 200 with no error field and no access_token (else branch)
- add test for HTTP 200 with JSON null body (new AttributeError guard)

* fix _OAUTH_TOKEN_FIELDS merge loop to preserve userinfo values on absent fields

When the token endpoint omits a bearer-credential field entirely (field
absent from token_response), the previous code deleted it from merged even
if userinfo provided a valid value. Now:
- non-null in token_response → restore authoritative token endpoint value
- explicit null in token_response → remove key from merged (clean absence)
- field absent from token_response → leave userinfo value unchanged

* use HTTP 401 for PKCE missing config errors

GENERIC_CLIENT_ID and GENERIC_TOKEN_ENDPOINT missing when PKCE is
enabled are auth-flow failures, not server errors. Use 401 instead
of 500 to avoid triggering false-positive server error alerts in
monitoring systems.

* address greptile review feedback (greploop iteration 40)

- fix duplicate error logging: demote first format-error log to DEBUG
  so the detailed ERROR in strict-mode branch is not duplicated
- add HTTP status code assertions to all PKCE ProxyException tests
  for better regression protection against accidental code changes

* add credential absence assertions to test_pkce_token_exchange_basic_auth

Verify that client_id and client_secret are NOT double-sent in the POST
body when Basic Auth is used (include_client_id=False with client_secret).
Catches regressions where credentials leak into both Auth header and body.

* address greptile review feedback (greploop iteration 41)

- add Bearer token header assertion to test_pkce_token_exchange_credentials_in_body
- add cache query assertions to both non-strict mode tests to confirm
  the cache was accessed before the warning path triggers

* address greptile review feedback (greploop iteration 42)

- assert null id_token is absent from merged result in basic auth test
- add test for HTTP 200 empty/null userinfo body with no id_token fallback

* use caplog to verify warning logs in non-strict cache miss tests

The two non-strict mode tests now use pytest's caplog fixture to assert
that a warning is actually emitted, not just that the code continues
without raising. This catches regressions where the warning silently
disappears.

* remove dead-code response=None guard in _pkce_token_exchange

* clean up stale pkce verifier cache entries in non-strict mode

* fix test: configure async_delete_cache as AsyncMock and assert cleanup called

* add sentinel guard so pkce-no-redis warning only fires once across hot-reloads

* fix misleading comments: code_verifier init and bearer-credential merge docs

* add best-effort cleanup in strict-mode for corrupt/empty cache entries

* add redirect_uri assertion, userinfo body in non-200 log, sentinel comment
2026-03-12 12:41:39 -07:00
.circleci [Fix] Update test_bad_database_url to match new startup error message 2026-03-11 16:25:35 -07:00
.claude Mcp user permissions (#21462) 2026-02-18 18:53:59 -08:00
.devcontainer chore: setting devcontainer for develop 2025-09-27 12:51:44 +09:00
.github ci: exclude enterprise/ from black --check in linting workflow 2026-03-12 14:27:00 -03:00
.semgrep/rules Merge branch 'main' into litellm_oss_staging_02_11_2026 2026-02-12 20:04:46 +05:30
ci_cd CircleCI test stability (#23055) 2026-03-07 15:19:39 -08:00
cookbook Revert "[Feature] Add /public/supported_endpoints endpoint" 2026-02-26 17:21:43 -08:00
db_scripts fix(migrate_keys.py): add script for migrating keys to new db 2025-07-16 10:18:36 -07:00
deploy merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
dist build: update dependencies 2025-11-01 12:58:39 -07:00
docker Fix CVEs: bump tar/minimatch/pypdf + harden Docker SBOM patching (#23082) 2026-03-07 18:31:27 -08:00
docs/my-website fix(openai): drop all reasoning_effort for gpt-5.4 + tools, including 'none' 2026-03-12 16:22:40 -03:00
enterprise bump: litellm-enterprise 0.1.33 → 0.1.34 2026-03-09 11:12:05 +00:00
litellm fix(sso): direct PKCE token exchange + Redis wiring for multi-instance SSO (#22923) 2026-03-12 12:41:39 -07:00
litellm-js build(deps): bump hono from 4.10.6 to 4.12.7 in /litellm-js/spend-logs (#23312) 2026-03-11 14:13:33 +05:30
litellm-proxy-extras merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
scripts [Feat] Add Tool Policies for AI Gateway (#22732) 2026-03-03 20:22:20 -08:00
tests fix(sso): direct PKCE token exchange + Redis wiring for multi-instance SSO (#22923) 2026-03-12 12:41:39 -07:00
ui/litellm-dashboard Merge branch 'main' into litellm_oss_staging_03_11_2026 2026-03-12 16:21:28 -03:00
.dockerignore fix critical CVE vulnerabliltes (#20683) 2026-02-07 22:23:01 -08:00
.env.example Add new model provider Novita AI (#7582) (#9527) 2025-05-12 21:49:30 -07:00
.flake8 chore: list all ignored flake8 rules explicit 2023-12-23 09:07:59 +01:00
.git-blame-ignore-revs Add my commit to .git-blame-ignore-revs 2024-05-12 10:21:10 -07:00
.gitattributes ignore ipynbs 2023-08-31 16:58:54 -07:00
.gitguardian.yaml [Fix] CI/CD - litellm_security_tests (#18567) 2026-01-01 14:20:04 -08:00
.gitignore Add observatory test workflow for RC/stable releases 2026-03-01 15:30:09 -03:00
.pre-commit-config.yaml docs(index.md): update release note with rc patch 2025-06-17 22:55:50 -07:00
.trivyignore litellm_fix(security): allowlist Next.js CVEs for 7 days (#20169) 2026-01-31 10:25:57 -08:00
AGENTS.md [Feat] UI Polish - MCP Servers page - show transport type (#23051) 2026-03-07 13:05:46 -08:00
ARCHITECTURE.md [Docs] Litellm architecture fixes 2 (#19252) 2026-01-16 14:52:16 -08:00
CLAUDE.md merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
codecov.yaml fix comment 2024-10-23 15:44:27 +05:30
CONTRIBUTING.md UI contributing and trouble shooting docs 2026-02-07 15:11:49 -08:00
dev_config.yaml [Feat] UI - Add Open in New Tab on leftnav Bar (#22731) 2026-03-03 19:56:55 -08:00
docker-compose.hardened.yml [Feature] Download Prisma binaries at build time instead of at runtime for Security Restricted environments (#17695) 2025-12-16 21:25:53 +05:30
docker-compose.yml fix(docker-compose.yml): move to docker.litellm.ai 2025-12-16 08:50:34 +05:30
Dockerfile Fix CVEs: bump tar/minimatch/pypdf + harden Docker SBOM patching (#23082) 2026-03-07 18:31:27 -08:00
GEMINI.md docs: cleanup README and improve agent guides (#17003) 2025-11-23 21:53:53 -08:00
index.yaml add 0.2.3 helm 2024-08-19 23:59:58 +08:00
LICENSE refactor: creating enterprise folder 2024-02-15 12:54:13 -08:00
license_cache.json fix failing tests 2026-02-21 15:48:26 -08:00
Makefile ci: add matrix-based parallel test workflow (#19942) 2026-02-12 19:39:05 +05:30
mcp_servers.json Add ScrapeGraph MCP server configuration (#18923) 2026-01-11 21:57:46 +05:30
model_prices_and_context_window.json Merge branch 'main' into litellm_oss_staging_03_11_2026 2026-03-12 10:43:08 -03:00
package-lock.json fix pkg lock 2025-11-22 11:51:15 -08:00
package.json Fix CVEs: bump tar/minimatch/pypdf + harden Docker SBOM patching (#23082) 2026-03-07 18:31:27 -08:00
poetry.lock chore: regenerate poetry.lock to match pyproject.toml (#23405) 2026-03-12 01:09:56 +00:00
policy_templates.json feat: Add Canadian PII protection (PIPEDA) (#22951) 2026-03-06 18:27:31 -08:00
prometheus.yml build(docker-compose.yml): add prometheus scraper to docker compose 2024-07-24 10:09:23 -07:00
provider_endpoints_support.json merge: resolve conflicts with upstream staging (bedrock + mcp tests) 2026-03-12 13:40:16 -03:00
proxy_server_config.yaml fix team budget checks 2026-01-31 15:28:33 -08:00
pyproject.toml bump: version 0.4.53 → 0.4.54 2026-03-11 18:07:58 -07:00
pyrightconfig.json Agents - support agent registration + discovery (A2A spec) (#16615) 2025-11-14 18:23:30 -08:00
README.md Merge pull request #20509 from ryan-crabbe/docs/mcp-trailing-slash 2026-02-24 16:38:29 -08:00
render.yaml build(render.yaml): fix health check route 2024-05-24 09:45:28 -07:00
requirements.txt bump: version 0.4.53 → 0.4.54 2026-03-11 18:07:58 -07:00
ruff.toml feat(anthropic): add Files API support for SDK 2026-03-11 12:45:19 -03:00
schema.prisma merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
security.md Corrected docs updates sept 2025 (#14916) 2025-09-25 15:49:19 -07:00
taplo.toml fix(agentcore): simplify agentcore streaming (#17141) 2026-01-19 05:20:24 -08:00
uv.lock fix(ollama): set finish_reason to tool_calls and remove broken capability check (#18924) 2026-01-14 03:52:26 +05:30

🚅 LiteLLM

Call 100+ LLMs in OpenAI format. [Bedrock, Azure, OpenAI, VertexAI, Anthropic, Groq, etc.]

Deploy to Render Deploy on Railway

LiteLLM Proxy Server (AI Gateway) | Hosted Proxy | Enterprise Tier

PyPI Version Y Combinator W23 Whatsapp Discord Slack

Group 7154 (1)

Use LiteLLM for

LLMs - Call 100+ LLMs (Python SDK + AI Gateway)

All Supported Endpoints - /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a, /messages and more.

Python SDK

pip install litellm
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic  
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

AI Gateway (Proxy Server)

Getting Started - E2E Tutorial - Setup virtual keys, make your first request

pip install 'litellm[proxy]'
litellm --model gpt-4o
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

Docs: LLM Providers

Agents - Invoke A2A Agents (Python SDK + AI Gateway)

Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI

Python SDK - A2A Protocol

from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

client = A2AClient(base_url="http://localhost:10001")

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello!"}],
            "messageId": uuid4().hex,
        }
    )
)
response = await client.send_message(request)

AI Gateway (Proxy Server)

Step 1. Add your Agent to the AI Gateway

Step 2. Call Agent via A2A SDK

from a2a.client import A2ACardResolver, A2AClient
from a2a.types import MessageSendParams, SendMessageRequest
from uuid import uuid4
import httpx

base_url = "http://localhost:4000/a2a/my-agent"  # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer sk-1234"}    # LiteLLM Virtual Key

async with httpx.AsyncClient(headers=headers) as httpx_client:
    resolver = A2ACardResolver(httpx_client=httpx_client, base_url=base_url)
    agent_card = await resolver.get_agent_card()
    client = A2AClient(httpx_client=httpx_client, agent_card=agent_card)

    request = SendMessageRequest(
        id=str(uuid4()),
        params=MessageSendParams(
            message={
                "role": "user",
                "parts": [{"kind": "text", "text": "Hello!"}],
                "messageId": uuid4().hex,
            }
        )
    )
    response = await client.send_message(request)

Docs: A2A Agent Gateway

MCP Tools - Connect MCP servers to any LLM (Python SDK + AI Gateway)

Python SDK - MCP Bridge

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm

server_params = StdioServerParameters(command="python", args=["mcp_server.py"])

async with stdio_client(server_params) as (read, write):
    async with ClientSession(read, write) as session:
        await session.initialize()

        # Load MCP tools in OpenAI format
        tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")

        # Use with any LiteLLM model
        response = await litellm.acompletion(
            model="gpt-4o",
            messages=[{"role": "user", "content": "What's 3 + 5?"}],
            tools=tools
        )

AI Gateway - MCP Gateway

Step 1. Add your MCP Server to the AI Gateway

Step 2. Call MCP tools via /chat/completions

curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer sk-1234' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

Use with Cursor IDE

{
  "mcpServers": {
    "LiteLLM": {
      "url": "http://localhost:4000/mcp/",
      "headers": {
        "x-litellm-api-key": "Bearer sk-1234"
      }
    }
  }
}

Docs: MCP Gateway


How to use LiteLLM

You can use LiteLLM through either the Proxy Server or Python SDK. Both gives you a unified interface to access multiple LLMs (100+ LLMs). Choose the option that best fits your needs:

LiteLLM AI Gateway LiteLLM Python SDK
Use Case Central service (LLM Gateway) to access multiple LLMs Use LiteLLM directly in your Python code
Who Uses It? Gen AI Enablement / ML Platform Teams Developers building LLM projects
Key Features Centralized API gateway with authentication and authorization, multi-tenant cost tracking and spend management per project/user, per-project customization (logging, guardrails, caching), virtual keys for secure access control, admin dashboard UI for monitoring and management Direct Python library integration in your codebase, Router with retry/fallback logic across multiple deployments (e.g. Azure/OpenAI) - Router, application-level load balancing and cost tracking, exception handling with OpenAI-compatible errors, observability callbacks (Lunary, MLflow, Langfuse, etc.)

LiteLLM Performance: 8ms P95 latency at 1k RPS (See benchmarks here)

Jump to LiteLLM Proxy (LLM Gateway) Docs
Jump to Supported LLM Providers

Stable Release: Use docker images with the -stable tag. These have undergone 12 hour load tests, before being published. More information about the release cycle here

Support for more providers. Missing a provider or LLM Platform, raise a feature request.

OSS Adopters

Stripe Google ADK Greptile OpenHands

Netflix

OpenAI Agents SDK

Supported Providers (Website Supported Models | Docs)

Provider /chat/completions /messages /responses /embeddings /image/generations /audio/transcriptions /audio/speech /moderations /batches /rerank
Abliteration (abliteration) ✅
AI/ML API (aiml) ✅ ✅ ✅ ✅ ✅
AI21 (ai21) ✅ ✅ ✅
AI21 Chat (ai21_chat) ✅ ✅ ✅
Aleph Alpha ✅ ✅ ✅
Amazon Nova ✅ ✅ ✅
Anthropic (anthropic) ✅ ✅ ✅ ✅
Anthropic Text (anthropic_text) ✅ ✅ ✅ ✅
Anyscale ✅ ✅ ✅
AssemblyAI (assemblyai) ✅ ✅ ✅ ✅
Auto Router (auto_router) ✅ ✅ ✅
AWS - Bedrock (bedrock) ✅ ✅ ✅ ✅ ✅
AWS - Sagemaker (sagemaker) ✅ ✅ ✅ ✅
Azure (azure) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure AI (azure_ai) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure Text (azure_text) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Baseten (baseten) ✅ ✅ ✅
Bytez (bytez) ✅ ✅ ✅
Cerebras (cerebras) ✅ ✅ ✅
Clarifai (clarifai) ✅ ✅ ✅
Cloudflare AI Workers (cloudflare) ✅ ✅ ✅
Codestral (codestral) ✅ ✅ ✅
Cohere (cohere) ✅ ✅ ✅ ✅ ✅
Cohere Chat (cohere_chat) ✅ ✅ ✅
CometAPI (cometapi) ✅ ✅ ✅ ✅
CompactifAI (compactifai) ✅ ✅ ✅
Custom (custom) ✅ ✅ ✅
Custom OpenAI (custom_openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Dashscope (dashscope) ✅ ✅ ✅
Databricks (databricks) ✅ ✅ ✅
DataRobot (datarobot) ✅ ✅ ✅
Deepgram (deepgram) ✅ ✅ ✅ ✅
DeepInfra (deepinfra) ✅ ✅ ✅
Deepseek (deepseek) ✅ ✅ ✅
ElevenLabs (elevenlabs) ✅ ✅ ✅ ✅ ✅
Empower (empower) ✅ ✅ ✅
Fal AI (fal_ai) ✅ ✅ ✅ ✅
Featherless AI (featherless_ai) ✅ ✅ ✅
Fireworks AI (fireworks_ai) ✅ ✅ ✅
FriendliAI (friendliai) ✅ ✅ ✅
Galadriel (galadriel) ✅ ✅ ✅
GitHub Copilot (github_copilot) ✅ ✅ ✅ ✅
GitHub Models (github) ✅ ✅ ✅
Google - PaLM ✅ ✅ ✅
Google - Vertex AI (vertex_ai) ✅ ✅ ✅ ✅ ✅
Google AI Studio - Gemini (gemini) ✅ ✅ ✅
GradientAI (gradient_ai) ✅ ✅ ✅
Groq AI (groq) ✅ ✅ ✅
Heroku (heroku) ✅ ✅ ✅
Hosted VLLM (hosted_vllm) ✅ ✅ ✅
Huggingface (huggingface) ✅ ✅ ✅ ✅ ✅
Hyperbolic (hyperbolic) ✅ ✅ ✅
IBM - Watsonx.ai (watsonx) ✅ ✅ ✅ ✅
Infinity (infinity) ✅
Jina AI (jina_ai) ✅
Lambda AI (lambda_ai) ✅ ✅ ✅
Lemonade (lemonade) ✅ ✅ ✅
LiteLLM Proxy (litellm_proxy) ✅ ✅ ✅ ✅ ✅
Llamafile (llamafile) ✅ ✅ ✅
LM Studio (lm_studio) ✅ ✅ ✅
Maritalk (maritalk) ✅ ✅ ✅
Meta - Llama API (meta_llama) ✅ ✅ ✅
Mistral AI API (mistral) ✅ ✅ ✅ ✅
Moonshot (moonshot) ✅ ✅ ✅
Morph (morph) ✅ ✅ ✅
Nebius AI Studio (nebius) ✅ ✅ ✅ ✅
NLP Cloud (nlp_cloud) ✅ ✅ ✅
Novita AI (novita) ✅ ✅ ✅
Nscale (nscale) ✅ ✅ ✅
Nvidia NIM (nvidia_nim) ✅ ✅ ✅
OCI (oci) ✅ ✅ ✅
Ollama (ollama) ✅ ✅ ✅ ✅
Ollama Chat (ollama_chat) ✅ ✅ ✅
Oobabooga (oobabooga) ✅ ✅ ✅ ✅ ✅ ✅ ✅
OpenAI (openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
OpenAI-like (openai_like) ✅
OpenRouter (openrouter) ✅ ✅ ✅
OVHCloud AI Endpoints (ovhcloud) ✅ ✅ ✅
Perplexity AI (perplexity) ✅ ✅ ✅
Petals (petals) ✅ ✅ ✅
Predibase (predibase) ✅ ✅ ✅
Recraft (recraft) ✅
Replicate (replicate) ✅ ✅ ✅
Sagemaker Chat (sagemaker_chat) ✅ ✅ ✅
Sambanova (sambanova) ✅ ✅ ✅
Snowflake (snowflake) ✅ ✅ ✅
Text Completion Codestral (text-completion-codestral) ✅ ✅ ✅
Text Completion OpenAI (text-completion-openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Together AI (together_ai) ✅ ✅ ✅
Topaz (topaz) ✅ ✅ ✅
Triton (triton) ✅ ✅ ✅
V0 (v0) ✅ ✅ ✅
Vercel AI Gateway (vercel_ai_gateway) ✅ ✅ ✅
VLLM (vllm) ✅ ✅ ✅
Volcengine (volcengine) ✅ ✅ ✅
Voyage AI (voyage) ✅
WandB Inference (wandb) ✅ ✅ ✅
Watsonx Text (watsonx_text) ✅ ✅ ✅
xAI (xai) ✅ ✅ ✅
Xinference (xinference) ✅

Read the Docs

Run in Developer mode

Services

  1. Setup .env file in root
  2. Run dependant services docker-compose up db prometheus

Backend

  1. (In root) create virtual environment python -m venv .venv
  2. Activate virtual environment source .venv/bin/activate
  3. Install dependencies pip install -e ".[all]"
  4. pip install prisma
  5. prisma generate
  6. Start proxy backend python litellm/proxy/proxy_cli.py

Frontend

  1. Navigate to ui/litellm-dashboard
  2. Install dependencies npm install
  3. Run npm run dev to start the dashboard

Enterprise

For companies that need better security, user management and professional support

Talk to founders

This covers:

  • ✅ Features under the LiteLLM Commercial License:
  • ✅ Feature Prioritization
  • ✅ Custom Integrations
  • ✅ Professional Support - Dedicated discord + slack
  • ✅ Custom SLAs
  • ✅ Secure access with Single Sign-On

Contributing

We welcome contributions to LiteLLM! Whether you're fixing bugs, adding features, or improving documentation, we appreciate your help.

Quick Start for Contributors

This requires poetry to be installed.

git clone https://github.com/BerriAI/litellm.git
cd litellm
make install-dev    # Install development dependencies
make format         # Format your code
make lint           # Run all linting checks
make test-unit      # Run unit tests
make format-check   # Check formatting only

For detailed contributing guidelines, see CONTRIBUTING.md.

Code Quality / Linting

LiteLLM follows the Google Python Style Guide.

Our automated checks include:

  • Black for code formatting
  • Ruff for linting and code quality
  • MyPy for type checking
  • Circular import detection
  • Import safety checks

All these checks must pass before your PR can be merged.

Support / talk with founders

Why did we build this

  • Need for simplicity: Our code started to get extremely complicated managing & translating calls between Azure, OpenAI and Cohere.

Contributors