* fix(sso): add direct PKCE token exchange and Redis cache wiring for multi-instance SSO When PKCE is enabled, bypass fastapi-sso and perform direct token exchange so code_verifier is correctly included. Store PKCE verifiers as dict in cache for proper JSON serialization in Redis. Wire user_api_key_cache to Redis when available so PKCE verifiers are shared across ECS tasks/pods. Also adds clearer error messages when PKCE is required but not configured. * refactor(sso): extract PKCE token exchange into SSOAuthenticationHandler methods - Move import httpx/jwt to module level (top of file, not inside function) - Extract inline PKCE token exchange + userinfo logic into two static methods: _pkce_token_exchange() and _get_pkce_userinfo() - get_generic_sso_response PKCE path is now a single method call - Fix double-logging in except block for non-PKCE errors - Use %-style log formatting (no f-strings in log calls) * fix: address greptile review feedback - Fix access_token missing in PKCE path: read from combined_response directly instead of generic_sso.access_token (which is only set by verify_and_process) - Fix PKCE error hint firing when PKCE is already enabled: only show 'set GENERIC_CLIENT_USE_PKCE=true' advice when code_verifier was absent - Fix unguarded KeyError on access_token: check for error field in HTTP 200 responses before accessing token_response['access_token'] - Fix silent empty userinfo: raise ProxyException when both userinfo endpoint and id_token fallback produce no user data - Fix backward-incompatible Redis wiring: only attach Redis to user_api_key_cache when GENERIC_CLIENT_USE_PKCE=true, preserving existing in-memory behaviour * fix: address second round of greptile review feedback - Fix PKCE error hint: check env var directly (not code_verifier presence) to distinguish 'PKCE not configured' from 'PKCE enabled but cache miss' - Fix misleading Redis TTL comment in proxy_server.py * fix: address third round of greptile review feedback - Fix CRITICAL log firing on every non-PKCE callback: only log when PKCE is enabled - Remove unused pkce_env_value intermediate variable - Prefer reusing redis_usage_cache over creating separate RedisCache instance (avoids losing advanced connection options like SSL, timeouts, db) * fix: address fourth round of greptile review feedback - Strip OAuth token credentials from response_convertor input to prevent access_token/id_token appearing in restricted-group error messages - Reuse single httpx.AsyncClient for both token exchange and userinfo requests to avoid a second TCP/TLS handshake per SSO callback - Revert Redis wiring to user_api_key_cache: PKCE code already uses redis_usage_cache directly; wiring would route all API-key lookups through Redis unnecessarily. Add startup warning instead when PKCE+Redis mismatch. - Move _OAUTH_TOKEN_FIELDS to module level * fix remaining PKCE test assertion for dict-format verifier storage * sanitize PKCE cache log to not expose verifier content * address greptile review feedback (greploop iteration 3) * address greptile review feedback (greploop iteration 4) * address greptile review feedback (greploop iteration 5) * simplify _get_pkce_userinfo: remove shared-client complexity, use async with directly * address greptile review feedback (greploop iteration 6) * address greptile review feedback (greploop iteration 7) * address greptile review feedback (greploop iteration 8) * address greptile review feedback (greploop iteration 9) * address greptile review feedback (greploop iteration 10) * address greptile review feedback (greploop iteration 11) * address greptile review feedback (greploop iteration 12) * fix misleading comment on user_api_key_cache TTL line * address greptile review feedback (greploop iteration 13) * address greptile review feedback (greploop iteration 14) * address greptile review feedback (greploop iteration 15) * address greptile review feedback (greploop iteration 16) * address greptile review feedback (greploop iteration 17) * address greptile review feedback (greploop iteration 18) * address greptile review feedback (greploop iteration 19) * address greptile review feedback (greploop iteration 20) * address greptile review feedback (greploop iteration 21) * address greptile review feedback (greploop iteration 22) * address greptile review feedback (greploop iteration 23) * address greptile review feedback (greploop iteration 24) * address greptile review feedback (greploop iteration 25) * address greptile review feedback (greploop iteration 26) * address greptile review feedback (greploop iteration 27) * address greptile review feedback (greploop iteration 28) * address greptile review feedback (greploop iteration 29) * address greptile review feedback (greploop iteration 30) * address greptile review feedback (greploop iteration 31) * address greptile review feedback (greploop iteration 32) * address greptile review feedback (greploop iteration 33) * address greptile review feedback (greploop iteration 34) * address greptile review feedback (greploop iteration 35) * address greptile review feedback (greploop iteration 37) - read GENERIC_CLIENT_USE_PKCE env var once in prepare_token_exchange_parameters - include actual decode error in jwt.decode failure exception message - add GENERIC_CLIENT_USE_PKCE=true to no-state regression test * defer PKCE verifier deletion until after all downstream processing Move _delete_pkce_verifier to after response_convertor and process_sso_jwt_access_token complete. If JWT processing raises, the verifier stays in cache so the user can retry without restarting the full OAuth flow. * address greptile review feedback (greploop iteration 38) - fix strict-mode cache miss error message to differentiate cross-instance routing failures (Redis configured) from single-instance issues (TTL expiry, pod restart) when only in-memory cache is available - add comment above _get_pkce_userinfo call explaining that bearer credentials are always sourced from token_response in the merge step * fix null JSON response body in _pkce_token_exchange - Guard against HTTP 200 with body null: response.json() returns None for JSON null, and calling .get() on None raises AttributeError. Now raises a clean ProxyException with a clear error message. - Fix misleading userinfo warning: was always saying "empty dict" but also fires for JSON null responses; updated to say "empty or null". - Add HTTP status code assertion to cache miss test. * address greptile review feedback (greploop iteration 39) - fix credential leakage: directly assign received_response from combined_response instead of relying on nonlocal mutation; Pyright was flagging the old guard as unreachable, meaning credential stripping might not execute — now it always runs unconditionally - add test for legacy plain-string cache format backward compat branch - add test for HTTP 200 with no error field and no access_token (else branch) - add test for HTTP 200 with JSON null body (new AttributeError guard) * fix _OAUTH_TOKEN_FIELDS merge loop to preserve userinfo values on absent fields When the token endpoint omits a bearer-credential field entirely (field absent from token_response), the previous code deleted it from merged even if userinfo provided a valid value. Now: - non-null in token_response → restore authoritative token endpoint value - explicit null in token_response → remove key from merged (clean absence) - field absent from token_response → leave userinfo value unchanged * use HTTP 401 for PKCE missing config errors GENERIC_CLIENT_ID and GENERIC_TOKEN_ENDPOINT missing when PKCE is enabled are auth-flow failures, not server errors. Use 401 instead of 500 to avoid triggering false-positive server error alerts in monitoring systems. * address greptile review feedback (greploop iteration 40) - fix duplicate error logging: demote first format-error log to DEBUG so the detailed ERROR in strict-mode branch is not duplicated - add HTTP status code assertions to all PKCE ProxyException tests for better regression protection against accidental code changes * add credential absence assertions to test_pkce_token_exchange_basic_auth Verify that client_id and client_secret are NOT double-sent in the POST body when Basic Auth is used (include_client_id=False with client_secret). Catches regressions where credentials leak into both Auth header and body. * address greptile review feedback (greploop iteration 41) - add Bearer token header assertion to test_pkce_token_exchange_credentials_in_body - add cache query assertions to both non-strict mode tests to confirm the cache was accessed before the warning path triggers * address greptile review feedback (greploop iteration 42) - assert null id_token is absent from merged result in basic auth test - add test for HTTP 200 empty/null userinfo body with no id_token fallback * use caplog to verify warning logs in non-strict cache miss tests The two non-strict mode tests now use pytest's caplog fixture to assert that a warning is actually emitted, not just that the code continues without raising. This catches regressions where the warning silently disappears. * remove dead-code response=None guard in _pkce_token_exchange * clean up stale pkce verifier cache entries in non-strict mode * fix test: configure async_delete_cache as AsyncMock and assert cleanup called * add sentinel guard so pkce-no-redis warning only fires once across hot-reloads * fix misleading comments: code_verifier init and bearer-credential merge docs * add best-effort cleanup in strict-mode for corrupt/empty cache entries * add redirect_uri assertion, userinfo body in non-200 log, sentinel comment |
||
|---|---|---|
| .circleci | ||
| .claude | ||
| .devcontainer | ||
| .github | ||
| .semgrep/rules | ||
| ci_cd | ||
| cookbook | ||
| db_scripts | ||
| deploy | ||
| dist | ||
| docker | ||
| docs/my-website | ||
| enterprise | ||
| litellm | ||
| litellm-js | ||
| litellm-proxy-extras | ||
| scripts | ||
| tests | ||
| ui/litellm-dashboard | ||
| .dockerignore | ||
| .env.example | ||
| .flake8 | ||
| .git-blame-ignore-revs | ||
| .gitattributes | ||
| .gitguardian.yaml | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| .trivyignore | ||
| AGENTS.md | ||
| ARCHITECTURE.md | ||
| CLAUDE.md | ||
| codecov.yaml | ||
| CONTRIBUTING.md | ||
| dev_config.yaml | ||
| docker-compose.hardened.yml | ||
| docker-compose.yml | ||
| Dockerfile | ||
| GEMINI.md | ||
| index.yaml | ||
| LICENSE | ||
| license_cache.json | ||
| Makefile | ||
| mcp_servers.json | ||
| model_prices_and_context_window.json | ||
| package-lock.json | ||
| package.json | ||
| poetry.lock | ||
| policy_templates.json | ||
| prometheus.yml | ||
| provider_endpoints_support.json | ||
| proxy_server_config.yaml | ||
| pyproject.toml | ||
| pyrightconfig.json | ||
| README.md | ||
| render.yaml | ||
| requirements.txt | ||
| ruff.toml | ||
| schema.prisma | ||
| security.md | ||
| taplo.toml | ||
| uv.lock | ||
🚅 LiteLLM
Call 100+ LLMs in OpenAI format. [Bedrock, Azure, OpenAI, VertexAI, Anthropic, Groq, etc.]
LiteLLM Proxy Server (AI Gateway) | Hosted Proxy | Enterprise Tier
Use LiteLLM for
LLMs - Call 100+ LLMs (Python SDK + AI Gateway)
All Supported Endpoints - /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a, /messages and more.
Python SDK
pip install litellm
from litellm import completion
import os
os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"
# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])
# Anthropic
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])
AI Gateway (Proxy Server)
Getting Started - E2E Tutorial - Setup virtual keys, make your first request
pip install 'litellm[proxy]'
litellm --model gpt-4o
import openai
client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}]
)
Agents - Invoke A2A Agents (Python SDK + AI Gateway)
Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI
Python SDK - A2A Protocol
from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4
client = A2AClient(base_url="http://localhost:10001")
request = SendMessageRequest(
id=str(uuid4()),
params=MessageSendParams(
message={
"role": "user",
"parts": [{"kind": "text", "text": "Hello!"}],
"messageId": uuid4().hex,
}
)
)
response = await client.send_message(request)
AI Gateway (Proxy Server)
Step 1. Add your Agent to the AI Gateway
Step 2. Call Agent via A2A SDK
from a2a.client import A2ACardResolver, A2AClient
from a2a.types import MessageSendParams, SendMessageRequest
from uuid import uuid4
import httpx
base_url = "http://localhost:4000/a2a/my-agent" # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer sk-1234"} # LiteLLM Virtual Key
async with httpx.AsyncClient(headers=headers) as httpx_client:
resolver = A2ACardResolver(httpx_client=httpx_client, base_url=base_url)
agent_card = await resolver.get_agent_card()
client = A2AClient(httpx_client=httpx_client, agent_card=agent_card)
request = SendMessageRequest(
id=str(uuid4()),
params=MessageSendParams(
message={
"role": "user",
"parts": [{"kind": "text", "text": "Hello!"}],
"messageId": uuid4().hex,
}
)
)
response = await client.send_message(request)
MCP Tools - Connect MCP servers to any LLM (Python SDK + AI Gateway)
Python SDK - MCP Bridge
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm
server_params = StdioServerParameters(command="python", args=["mcp_server.py"])
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
# Load MCP tools in OpenAI format
tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")
# Use with any LiteLLM model
response = await litellm.acompletion(
model="gpt-4o",
messages=[{"role": "user", "content": "What's 3 + 5?"}],
tools=tools
)
AI Gateway - MCP Gateway
Step 1. Add your MCP Server to the AI Gateway
Step 2. Call MCP tools via /chat/completions
curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Summarize the latest open PR"}],
"tools": [{
"type": "mcp",
"server_url": "litellm_proxy/mcp/github",
"server_label": "github_mcp",
"require_approval": "never"
}]
}'
Use with Cursor IDE
{
"mcpServers": {
"LiteLLM": {
"url": "http://localhost:4000/mcp/",
"headers": {
"x-litellm-api-key": "Bearer sk-1234"
}
}
}
}
How to use LiteLLM
You can use LiteLLM through either the Proxy Server or Python SDK. Both gives you a unified interface to access multiple LLMs (100+ LLMs). Choose the option that best fits your needs:
| LiteLLM AI Gateway | LiteLLM Python SDK | |
|---|---|---|
| Use Case | Central service (LLM Gateway) to access multiple LLMs | Use LiteLLM directly in your Python code |
| Who Uses It? | Gen AI Enablement / ML Platform Teams | Developers building LLM projects |
| Key Features | Centralized API gateway with authentication and authorization, multi-tenant cost tracking and spend management per project/user, per-project customization (logging, guardrails, caching), virtual keys for secure access control, admin dashboard UI for monitoring and management | Direct Python library integration in your codebase, Router with retry/fallback logic across multiple deployments (e.g. Azure/OpenAI) - Router, application-level load balancing and cost tracking, exception handling with OpenAI-compatible errors, observability callbacks (Lunary, MLflow, Langfuse, etc.) |
LiteLLM Performance: 8ms P95 latency at 1k RPS (See benchmarks here)
Jump to LiteLLM Proxy (LLM Gateway) Docs
Jump to Supported LLM Providers
Stable Release: Use docker images with the -stable tag. These have undergone 12 hour load tests, before being published. More information about the release cycle here
Support for more providers. Missing a provider or LLM Platform, raise a feature request.
OSS Adopters
Netflix |
Supported Providers (Website Supported Models | Docs)
Run in Developer mode
Services
- Setup .env file in root
- Run dependant services
docker-compose up db prometheus
Backend
- (In root) create virtual environment
python -m venv .venv - Activate virtual environment
source .venv/bin/activate - Install dependencies
pip install -e ".[all]" pip install prismaprisma generate- Start proxy backend
python litellm/proxy/proxy_cli.py
Frontend
- Navigate to
ui/litellm-dashboard - Install dependencies
npm install - Run
npm run devto start the dashboard
Enterprise
For companies that need better security, user management and professional support
This covers:
- ✅ Features under the LiteLLM Commercial License:
- ✅ Feature Prioritization
- ✅ Custom Integrations
- ✅ Professional Support - Dedicated discord + slack
- ✅ Custom SLAs
- ✅ Secure access with Single Sign-On
Contributing
We welcome contributions to LiteLLM! Whether you're fixing bugs, adding features, or improving documentation, we appreciate your help.
Quick Start for Contributors
This requires poetry to be installed.
git clone https://github.com/BerriAI/litellm.git
cd litellm
make install-dev # Install development dependencies
make format # Format your code
make lint # Run all linting checks
make test-unit # Run unit tests
make format-check # Check formatting only
For detailed contributing guidelines, see CONTRIBUTING.md.
Code Quality / Linting
LiteLLM follows the Google Python Style Guide.
Our automated checks include:
- Black for code formatting
- Ruff for linting and code quality
- MyPy for type checking
- Circular import detection
- Import safety checks
All these checks must pass before your PR can be merged.
Support / talk with founders
- Schedule Demo 👋
- Community Discord 💭
- Community Slack 💭
- Our numbers 📞 +1 (770) 8783-106 / +1 (412) 618-6238
- Our emails ✉️ ishaan@berri.ai / krrish@berri.ai
Why did we build this
- Need for simplicity: Our code started to get extremely complicated managing & translating calls between Azure, OpenAI and Cohere.