Find a file
Alex Ker 742b6be36a
fix: support served_model_name for Baseten dedicated deployments (#23382)
* fix: langfuse trace leak key on model params

* fix: pop sensitive keys from langfuse

* fixes

* fix: set oauth2_flow when building MCPServer in _execute_with_mcp_client

* fix: add oauth2_flow to NewMCPServerRequest and guard auto-detect with token_url

* fix: narrow oauth2_flow type to Literal in NewMCPServerRequest

* fix: align DefaultInternalUserParams Pydantic default with runtime fallback

The Pydantic default for user_role was INTERNAL_USER, but all runtime
provisioning paths (SSO, SCIM, JWT) fall back to INTERNAL_USER_VIEW_ONLY
when no settings are saved. This caused the UI to show "Internal User"
on fresh instances while new users actually got "Internal Viewer".

* test: add regression test for fresh-instance default role sync

Asserts that GET /get/internal_user_settings returns
INTERNAL_USER_VIEW_ONLY on a fresh DB with no saved settings,
matching the runtime fallback in SSO/SCIM/JWT provisioning.

* Update tests/test_litellm/proxy/ui_crud_endpoints/test_proxy_setting_endpoints.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Add unit tests for 5 previously untested UI dashboard files

Tests added for: UiLoadingSpinner, HashicorpVaultEmptyPlaceholder,
PageVisibilitySettings, errorUtils, and mcpToolCrudClassification.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove skip decorators from m2m tests now that oauth2_flow is set

* [Fix] Privilege escalation: restrict /key/block, /key/unblock, and max_budget updates to admins

Non-admin users (INTERNAL_USER) could call /key/block and /key/unblock on
arbitrary keys, and modify max_budget on their own keys via /key/update.
These endpoints are now restricted to proxy admins, team admins, or org admins.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore(ui): migrate DefaultUserSettings buttons from Tremor to antd

* [Infra] Merging RC Branch with Main (#23786)

* fix(test): add missing mocks for test_streamable_http_mcp_handler_mock

The test was missing mocks for extract_mcp_auth_context and set_auth_context,
causing the handler to fail silently in the except block instead of reaching
session_manager.handle_request. This mirrors the fix already applied to the
sibling test_sse_mcp_handler_mock.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ci): route OpenAI models through chat completions in pass-through tests

The test_anthropic_messages_openai_model_streaming_cost_injection test fails
because the OpenAI Responses API returns 400 for requests routed through the
Anthropic Messages endpoint. Setting LITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES=true
routes OpenAI models through the stable chat completions path instead.
Cost injection still works since it happens at the proxy level.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ci): fix assemblyai custom auth and router wildcard test flakiness

1. custom_auth_basic.py: Add user_role='proxy_admin' so the custom auth
   user can access management endpoints like /key/generate. The test
   test_assemblyai_transcribe_with_non_admin_key was hidden behind an
   earlier -x failure and was never reached before.

2. test_router_utils.py: Add flaky(retries=3) and increase sleep from 1s
   to 2s for test_router_get_model_group_usage_wildcard_routes. The async
   callback needs time to write usage to cache, and 1s is insufficient on
   slower CI hardware.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* ci: retrigger CI pipeline

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(mypy): use LitellmUserRoles enum instead of raw string in custom_auth_basic

Fixes mypy error: Argument 'user_role' has incompatible type 'str'; expected 'LitellmUserRoles | None'

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22926)

* fix: don't close HTTP/SDK clients on LLMClientCache eviction

Removing the _remove_key override that eagerly called aclose()/close()
on evicted clients. Evicted clients may still be held by in-flight
streaming requests; closing them causes:

  RuntimeError: Cannot send a request, as the client has been closed.

This is a regression from commit fb72979432. Clients that are no longer
referenced will be garbage-collected naturally. Explicit shutdown cleanup
happens via close_litellm_async_clients().

Fixes production crashes after the 1-hour cache TTL expires.

* test: update LLMClientCache unit tests for no-close-on-eviction behavior

Flip the assertions: evicted clients must NOT be closed. Replace
test_remove_key_closes_async_client → test_remove_key_does_not_close_async_client
and equivalents for sync/eviction paths.

Add test_remove_key_removes_plain_values for non-client cache entries.
Remove test_background_tasks_cleaned_up_after_completion (no more _background_tasks).
Remove test_remove_key_no_event_loop variant that depended on old behavior.

* test: add e2e tests for OpenAI SDK client surviving cache eviction

Add two new e2e tests using real AsyncOpenAI clients:
- test_evicted_openai_sdk_client_stays_usable: verifies size-based eviction
  doesn't close the client
- test_ttl_expired_openai_sdk_client_stays_usable: verifies TTL expiry
  eviction doesn't close the client

Both tests sleep after eviction so any create_task()-based close would
have time to run, making the regression detectable.

Also expand the module docstring to explain why the sleep is required.

* docs(AGENTS.md): add rule — never close HTTP/SDK clients on cache eviction

* docs(CLAUDE.md): add HTTP client cache safety guideline

* [Fix] Install bsdmainutils for column command in security scans

The security_scans.sh script uses `column` to format vulnerability
output, but the package wasn't installed in the CI environment.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: handle string callback values in prometheus multiproc setup

When callbacks are configured as a plain string (e.g., `callbacks: "my_callback"`)
instead of a list, the proxy crashes on startup with:
  TypeError: can only concatenate str (not "list") to str

Normalize each callback setting to a list before concatenating.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* bump: version 1.82.2 → 1.82.3

* fix(test): update test_startup_fails_when_db_setup_fails for opt-in enforcement

The --enforce_prisma_migration_check flag is now required to trigger
sys.exit(1) on DB migration failure, after #23675 flipped the default
behavior to warn-and-continue.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(cost_calculator): use model name for per-request custom pricing when router_model_id has no pricing

When custom pricing is passed as per-request kwargs (input_cost_per_token/output_cost_per_token),
completion() registers pricing under the model name, but _select_model_name_for_cost_calc was
selecting the router deployment hash (which has no pricing data), causing response_cost to be 0.0.

Now checks whether the router_model_id entry actually has pricing before preferring it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* Update litellm/proxy/management_endpoints/key_management_endpoints.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* fix: clear oauth2_flow when client_credentials set without token_url

* chore(ui): use antd danger prop instead of tailwind for Remove button

* feat: fetch blog posts from docs RSS feed instead of static JSON on GitHub

* fix: remove unused Any import from get_blog_posts

* [Fix] UI - Logs: Fix empty filter results showing stale data

Remove `.length > 0` check so that when a backend filter returns an
empty result set the table correctly shows no data instead of falling
back to the previous unfiltered logs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Fix] Reapply empty filter fix after merge with main

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Fix] Prevent internal users from creating invalid keys via key/generate and key/update

Internal users could exploit key/generate and key/update to create unbound
keys (no user_id, no budget) or attach keys to non-existent teams. This
adds validation for non-admin callers: auto-assign user_id on generate,
reject invalid team_ids, and prevent removing user_id on update.

Closes LIT-1884

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Fix] Remove duplicate get_team_object call in _validate_update_key_data

Move the non-admin team validation into the existing get_team_object call
site to avoid an extra DB round-trip. The existing call already fetches
the team for limits checking — we now add the LIT-1884 guard there when
team_obj is None for non-admin callers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Fix] Skip key_alias re-validation on update/regenerate when alias unchanged

When updating or regenerating a key without changing its key_alias, the
existing alias was being re-validated against current format rules. This
caused keys with legacy aliases (created before stricter validation) to
become uneditable. Now validation only runs when the alias actually changes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Fix] Update log filter test to match empty-result behavior

The test expected fallback to all logs when backend filters return empty,
but the source was intentionally changed to show empty results instead of
stale data. Updated test to match.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Feature] Disable custom API key values via UI setting

Add disable_custom_api_keys UI setting that prevents users from specifying
custom key values during key generation and regeneration. When enabled, all
keys must be auto-generated, eliminating the risk of key hash collisions
in multi-tenant environments.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Fix] Add disable_custom_api_keys to UISettings Pydantic model

Without this field on the model, GET /get/ui_settings omits the setting
from the response and field_schema, preventing the UI from reading it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: Register DynamoAI guardrail initializer and enum entry (#23752)

* fix: Register DynamoAI guardrail initializer and enum entry

Fix the "Unsupported guardrail: dynamoai" error by:
1. Adding DYNAMOAI to SupportedGuardrailIntegrations enum
2. Implementing initialize_guardrail() and registries in dynamoai/__init__.py

The DynamoAI guardrail was added in PR #15920 but never properly registered
in the initialization system. The __init__.py was missing the
guardrail_initializer_registry and guardrail_class_registry dictionaries
that the dynamic discovery mechanism looks for at module load time.

Fixes #22773

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* Update litellm/proxy/guardrails/guardrail_hooks/dynamoai/__init__.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update litellm/proxy/guardrails/guardrail_hooks/dynamoai/__init__.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* test: Add tests for DynamoAI guardrail registration

Verifies enum entry, initializer registry, class registry,
instance creation, and global registry discovery.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Haiku 4.5 <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* docs: add v1.82.3 release notes and update provider_endpoints_support.json (#23816)

* [Feature] Add disable_custom_api_keys toggle to UI Settings page

Adds a toggle switch to the admin UI Settings page so administrators can
enable/disable custom API key values without making direct API calls.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Revert "docs: add v1.82.3 release notes and update provider_endpoints_support…" (#23817)

This reverts commit 966124966f.

* [Fix] Rename toggle label to "Disable custom Virtual key values"

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [Fix] Remove "API" from custom key description text

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ui): CSV export empty on Global Usage page

Aggregated endpoint returns empty breakdown.entities; fall back to
grouping breakdown.api_keys by team_id.

* Revert "fix: langfuse trace leak key on model params"

* fix: support served_model_name for Baseten dedicated deployments

Baseten dedicated deployments use an 8-char deployment ID for URL routing,
but the vLLM server may expect a different model name in the request body
(e.g. baseten-hosted/zai-org/GLM-5 vs wd1lndkw). Add served_model_name
litellm_param to override the model field in the request body, and declare
it in LiteLLMParamsTypedDict and GenericLiteLLMParams for IDE support.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Harshit Jain <harshitjain0562@gmail.com>
Co-authored-by: Harshit Jain <48647625+Harshit28j@users.noreply.github.com>
Co-authored-by: AlexKer <AlexKer@users.noreply.github.com>
Co-authored-by: joereyna <joseph.reyna@gmail.com>
Co-authored-by: Ryan Crabbe <rcrabbe@berkeley.edu>
Co-authored-by: ryan-crabbe <128659760+ryan-crabbe@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-03-17 12:58:51 -07:00
.circleci fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
.claude Mcp user permissions (#21462) 2026-02-18 18:53:59 -08:00
.devcontainer chore: setting devcontainer for develop 2025-09-27 12:51:44 +09:00
.github Add CodSpeed performance benchmarks (#23676) 2026-03-14 18:44:36 -07:00
.semgrep/rules fix: prompt registry 2026-02-18 00:34:54 +05:30
ci_cd fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
cookbook Revert "[Feature] Add /public/supported_endpoints endpoint" 2026-02-26 17:21:43 -08:00
db_scripts fix(migrate_keys.py): add script for migrating keys to new db 2025-07-16 10:18:36 -07:00
deploy merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
dist build: update dependencies 2025-11-01 12:58:39 -07:00
docker fix: bump PyJWT to 2.12.0 in all Dockerfiles and tar to 7.5.11 2026-03-14 19:54:54 -07:00
docs/my-website fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
enterprise Merge pull request #23718 from BerriAI/litellm_fix_vertex_ai_batch 2026-03-16 19:05:49 +05:30
litellm fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
litellm-js build(deps): bump hono from 4.10.6 to 4.12.7 in /litellm-js/spend-logs (#23312) 2026-03-11 14:13:33 +05:30
litellm-proxy-extras fix: prisma migrate deploy failures on pre-existing instances (#23655) 2026-03-14 16:54:21 -07:00
scripts [Feat] Add Tool Policies for AI Gateway (#22732) 2026-03-03 20:22:20 -08:00
tests fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
ui/litellm-dashboard fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
.dockerignore fix: prompt registry 2026-02-18 00:34:54 +05:30
.env.example Add new model provider Novita AI (#7582) (#9527) 2025-05-12 21:49:30 -07:00
.flake8 chore: list all ignored flake8 rules explicit 2023-12-23 09:07:59 +01:00
.git-blame-ignore-revs Add my commit to .git-blame-ignore-revs 2024-05-12 10:21:10 -07:00
.gitattributes ignore ipynbs 2023-08-31 16:58:54 -07:00
.gitguardian.yaml fix: prompt registry 2026-02-18 00:34:54 +05:30
.gitignore Add observatory test workflow for RC/stable releases 2026-03-01 15:30:09 -03:00
.pre-commit-config.yaml docs(index.md): update release note with rc patch 2025-06-17 22:55:50 -07:00
.trivyignore fix: prompt registry 2026-02-18 00:34:54 +05:30
AGENTS.md [Feat] UI Polish - MCP Servers page - show transport type (#23051) 2026-03-07 13:05:46 -08:00
ARCHITECTURE.md fix: prompt registry 2026-02-18 00:34:54 +05:30
CLAUDE.md fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
codecov.yaml fix comment 2024-10-23 15:44:27 +05:30
CONTRIBUTING.md fix: prompt registry 2026-02-18 00:34:54 +05:30
dev_config.yaml [Feat] UI - Add Open in New Tab on leftnav Bar (#22731) 2026-03-03 19:56:55 -08:00
docker-compose.hardened.yml fix: prompt registry 2026-02-18 00:34:54 +05:30
docker-compose.yml fix: prompt registry 2026-02-18 00:34:54 +05:30
Dockerfile fix: bump PyJWT to 2.12.0 in all Dockerfiles and tar to 7.5.11 2026-03-14 19:54:54 -07:00
GEMINI.md docs: cleanup README and improve agent guides (#17003) 2025-11-23 21:53:53 -08:00
index.yaml add 0.2.3 helm 2024-08-19 23:59:58 +08:00
LICENSE refactor: creating enterprise folder 2024-02-15 12:54:13 -08:00
license_cache.json fix failing tests 2026-02-21 15:48:26 -08:00
Makefile fix: prompt registry 2026-02-18 00:34:54 +05:30
mcp_servers.json fix: prompt registry 2026-02-18 00:34:54 +05:30
model_prices_and_context_window.json fix: auto-fill reasoning_content for moonshot kimi reasoning models in multi-turn tool calling (#23580) 2026-03-13 22:48:45 -07:00
package-lock.json fix pkg lock 2025-11-22 11:51:15 -08:00
package.json bumping tar for security 2026-03-13 10:43:16 -07:00
poetry.lock updating poetry lock 2026-03-14 19:21:50 -07:00
policy_templates.json feat: Add Canadian PII protection (PIPEDA) (#22951) 2026-03-06 18:27:31 -08:00
prometheus.yml build(docker-compose.yml): add prometheus scraper to docker compose 2024-07-24 10:09:23 -07:00
provider_endpoints_support.json merge: resolve conflicts with upstream staging (bedrock + mcp tests) 2026-03-12 13:40:16 -03:00
proxy_server_config.yaml [Fix] Fix test_users_in_team_budget using model with no pricing data 2026-03-13 12:35:20 -07:00
pyproject.toml fix: support served_model_name for Baseten dedicated deployments (#23382) 2026-03-17 12:58:51 -07:00
pyrightconfig.json Agents - support agent registration + discovery (A2A spec) (#16615) 2025-11-14 18:23:30 -08:00
README.md Add CodSpeed performance benchmarks (#23676) 2026-03-14 18:44:36 -07:00
render.yaml build(render.yaml): fix health check route 2024-05-24 09:45:28 -07:00
requirements.txt bumping pyJWT for security 2026-03-14 19:17:19 -07:00
ruff.toml refactor: extract _list_tools_for_single_server helper to fix PLR0915 2026-03-12 20:06:10 +00:00
schema.prisma merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
security.md Corrected docs updates sept 2025 (#14916) 2025-09-25 15:49:19 -07:00
taplo.toml fix: prompt registry 2026-02-18 00:34:54 +05:30
uv.lock fix: prompt registry 2026-02-18 00:34:54 +05:30

🚅 LiteLLM

Call 100+ LLMs in OpenAI format. [Bedrock, Azure, OpenAI, VertexAI, Anthropic, Groq, etc.]

Deploy to Render Deploy on Railway

LiteLLM Proxy Server (AI Gateway) | Hosted Proxy | Enterprise Tier

PyPI Version Y Combinator W23 Whatsapp Discord Slack CodSpeed

Group 7154 (1)

Use LiteLLM for

LLMs - Call 100+ LLMs (Python SDK + AI Gateway)

All Supported Endpoints - /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a, /messages and more.

Python SDK

pip install litellm
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic  
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

AI Gateway (Proxy Server)

Getting Started - E2E Tutorial - Setup virtual keys, make your first request

pip install 'litellm[proxy]'
litellm --model gpt-4o
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

Docs: LLM Providers

Agents - Invoke A2A Agents (Python SDK + AI Gateway)

Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI

Python SDK - A2A Protocol

from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

client = A2AClient(base_url="http://localhost:10001")

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello!"}],
            "messageId": uuid4().hex,
        }
    )
)
response = await client.send_message(request)

AI Gateway (Proxy Server)

Step 1. Add your Agent to the AI Gateway

Step 2. Call Agent via A2A SDK

from a2a.client import A2ACardResolver, A2AClient
from a2a.types import MessageSendParams, SendMessageRequest
from uuid import uuid4
import httpx

base_url = "http://localhost:4000/a2a/my-agent"  # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer sk-1234"}    # LiteLLM Virtual Key

async with httpx.AsyncClient(headers=headers) as httpx_client:
    resolver = A2ACardResolver(httpx_client=httpx_client, base_url=base_url)
    agent_card = await resolver.get_agent_card()
    client = A2AClient(httpx_client=httpx_client, agent_card=agent_card)

    request = SendMessageRequest(
        id=str(uuid4()),
        params=MessageSendParams(
            message={
                "role": "user",
                "parts": [{"kind": "text", "text": "Hello!"}],
                "messageId": uuid4().hex,
            }
        )
    )
    response = await client.send_message(request)

Docs: A2A Agent Gateway

MCP Tools - Connect MCP servers to any LLM (Python SDK + AI Gateway)

Python SDK - MCP Bridge

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm

server_params = StdioServerParameters(command="python", args=["mcp_server.py"])

async with stdio_client(server_params) as (read, write):
    async with ClientSession(read, write) as session:
        await session.initialize()

        # Load MCP tools in OpenAI format
        tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")

        # Use with any LiteLLM model
        response = await litellm.acompletion(
            model="gpt-4o",
            messages=[{"role": "user", "content": "What's 3 + 5?"}],
            tools=tools
        )

AI Gateway - MCP Gateway

Step 1. Add your MCP Server to the AI Gateway

Step 2. Call MCP tools via /chat/completions

curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer sk-1234' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

Use with Cursor IDE

{
  "mcpServers": {
    "LiteLLM": {
      "url": "http://localhost:4000/mcp/",
      "headers": {
        "x-litellm-api-key": "Bearer sk-1234"
      }
    }
  }
}

Docs: MCP Gateway


How to use LiteLLM

You can use LiteLLM through either the Proxy Server or Python SDK. Both gives you a unified interface to access multiple LLMs (100+ LLMs). Choose the option that best fits your needs:

LiteLLM AI Gateway LiteLLM Python SDK
Use Case Central service (LLM Gateway) to access multiple LLMs Use LiteLLM directly in your Python code
Who Uses It? Gen AI Enablement / ML Platform Teams Developers building LLM projects
Key Features Centralized API gateway with authentication and authorization, multi-tenant cost tracking and spend management per project/user, per-project customization (logging, guardrails, caching), virtual keys for secure access control, admin dashboard UI for monitoring and management Direct Python library integration in your codebase, Router with retry/fallback logic across multiple deployments (e.g. Azure/OpenAI) - Router, application-level load balancing and cost tracking, exception handling with OpenAI-compatible errors, observability callbacks (Lunary, MLflow, Langfuse, etc.)

LiteLLM Performance: 8ms P95 latency at 1k RPS (See benchmarks here)

Jump to LiteLLM Proxy (LLM Gateway) Docs
Jump to Supported LLM Providers

Stable Release: Use docker images with the -stable tag. These have undergone 12 hour load tests, before being published. More information about the release cycle here

Support for more providers. Missing a provider or LLM Platform, raise a feature request.

OSS Adopters

Stripe Google ADK Greptile OpenHands

Netflix

OpenAI Agents SDK

Supported Providers (Website Supported Models | Docs)

Provider /chat/completions /messages /responses /embeddings /image/generations /audio/transcriptions /audio/speech /moderations /batches /rerank
Abliteration (abliteration) ✅
AI/ML API (aiml) ✅ ✅ ✅ ✅ ✅
AI21 (ai21) ✅ ✅ ✅
AI21 Chat (ai21_chat) ✅ ✅ ✅
Aleph Alpha ✅ ✅ ✅
Amazon Nova ✅ ✅ ✅
Anthropic (anthropic) ✅ ✅ ✅ ✅
Anthropic Text (anthropic_text) ✅ ✅ ✅ ✅
Anyscale ✅ ✅ ✅
AssemblyAI (assemblyai) ✅ ✅ ✅ ✅
Auto Router (auto_router) ✅ ✅ ✅
AWS - Bedrock (bedrock) ✅ ✅ ✅ ✅ ✅
AWS - Sagemaker (sagemaker) ✅ ✅ ✅ ✅
Azure (azure) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure AI (azure_ai) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure Text (azure_text) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Baseten (baseten) ✅ ✅ ✅
Bytez (bytez) ✅ ✅ ✅
Cerebras (cerebras) ✅ ✅ ✅
Clarifai (clarifai) ✅ ✅ ✅
Cloudflare AI Workers (cloudflare) ✅ ✅ ✅
Codestral (codestral) ✅ ✅ ✅
Cohere (cohere) ✅ ✅ ✅ ✅ ✅
Cohere Chat (cohere_chat) ✅ ✅ ✅
CometAPI (cometapi) ✅ ✅ ✅ ✅
CompactifAI (compactifai) ✅ ✅ ✅
Custom (custom) ✅ ✅ ✅
Custom OpenAI (custom_openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Dashscope (dashscope) ✅ ✅ ✅
Databricks (databricks) ✅ ✅ ✅
DataRobot (datarobot) ✅ ✅ ✅
Deepgram (deepgram) ✅ ✅ ✅ ✅
DeepInfra (deepinfra) ✅ ✅ ✅
Deepseek (deepseek) ✅ ✅ ✅
ElevenLabs (elevenlabs) ✅ ✅ ✅ ✅ ✅
Empower (empower) ✅ ✅ ✅
Fal AI (fal_ai) ✅ ✅ ✅ ✅
Featherless AI (featherless_ai) ✅ ✅ ✅
Fireworks AI (fireworks_ai) ✅ ✅ ✅
FriendliAI (friendliai) ✅ ✅ ✅
Galadriel (galadriel) ✅ ✅ ✅
GitHub Copilot (github_copilot) ✅ ✅ ✅ ✅
GitHub Models (github) ✅ ✅ ✅
Google - PaLM ✅ ✅ ✅
Google - Vertex AI (vertex_ai) ✅ ✅ ✅ ✅ ✅
Google AI Studio - Gemini (gemini) ✅ ✅ ✅
GradientAI (gradient_ai) ✅ ✅ ✅
Groq AI (groq) ✅ ✅ ✅
Heroku (heroku) ✅ ✅ ✅
Hosted VLLM (hosted_vllm) ✅ ✅ ✅
Huggingface (huggingface) ✅ ✅ ✅ ✅ ✅
Hyperbolic (hyperbolic) ✅ ✅ ✅
IBM - Watsonx.ai (watsonx) ✅ ✅ ✅ ✅
Infinity (infinity) ✅
Jina AI (jina_ai) ✅
Lambda AI (lambda_ai) ✅ ✅ ✅
Lemonade (lemonade) ✅ ✅ ✅
LiteLLM Proxy (litellm_proxy) ✅ ✅ ✅ ✅ ✅
Llamafile (llamafile) ✅ ✅ ✅
LM Studio (lm_studio) ✅ ✅ ✅
Maritalk (maritalk) ✅ ✅ ✅
Meta - Llama API (meta_llama) ✅ ✅ ✅
Mistral AI API (mistral) ✅ ✅ ✅ ✅
Moonshot (moonshot) ✅ ✅ ✅
Morph (morph) ✅ ✅ ✅
Nebius AI Studio (nebius) ✅ ✅ ✅ ✅
NLP Cloud (nlp_cloud) ✅ ✅ ✅
Novita AI (novita) ✅ ✅ ✅
Nscale (nscale) ✅ ✅ ✅
Nvidia NIM (nvidia_nim) ✅ ✅ ✅
OCI (oci) ✅ ✅ ✅
Ollama (ollama) ✅ ✅ ✅ ✅
Ollama Chat (ollama_chat) ✅ ✅ ✅
Oobabooga (oobabooga) ✅ ✅ ✅ ✅ ✅ ✅ ✅
OpenAI (openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
OpenAI-like (openai_like) ✅
OpenRouter (openrouter) ✅ ✅ ✅
OVHCloud AI Endpoints (ovhcloud) ✅ ✅ ✅
Perplexity AI (perplexity) ✅ ✅ ✅
Petals (petals) ✅ ✅ ✅
Predibase (predibase) ✅ ✅ ✅
Recraft (recraft) ✅
Replicate (replicate) ✅ ✅ ✅
Sagemaker Chat (sagemaker_chat) ✅ ✅ ✅
Sambanova (sambanova) ✅ ✅ ✅
Snowflake (snowflake) ✅ ✅ ✅
Text Completion Codestral (text-completion-codestral) ✅ ✅ ✅
Text Completion OpenAI (text-completion-openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Together AI (together_ai) ✅ ✅ ✅
Topaz (topaz) ✅ ✅ ✅
Triton (triton) ✅ ✅ ✅
V0 (v0) ✅ ✅ ✅
Vercel AI Gateway (vercel_ai_gateway) ✅ ✅ ✅
VLLM (vllm) ✅ ✅ ✅
Volcengine (volcengine) ✅ ✅ ✅
Voyage AI (voyage) ✅
WandB Inference (wandb) ✅ ✅ ✅
Watsonx Text (watsonx_text) ✅ ✅ ✅
xAI (xai) ✅ ✅ ✅
Xinference (xinference) ✅

Read the Docs

Run in Developer mode

Services

  1. Setup .env file in root
  2. Run dependant services docker-compose up db prometheus

Backend

  1. (In root) create virtual environment python -m venv .venv
  2. Activate virtual environment source .venv/bin/activate
  3. Install dependencies pip install -e ".[all]"
  4. pip install prisma
  5. prisma generate
  6. Start proxy backend python litellm/proxy/proxy_cli.py

Frontend

  1. Navigate to ui/litellm-dashboard
  2. Install dependencies npm install
  3. Run npm run dev to start the dashboard

Enterprise

For companies that need better security, user management and professional support

Talk to founders

This covers:

  • ✅ Features under the LiteLLM Commercial License:
  • ✅ Feature Prioritization
  • ✅ Custom Integrations
  • ✅ Professional Support - Dedicated discord + slack
  • ✅ Custom SLAs
  • ✅ Secure access with Single Sign-On

Contributing

We welcome contributions to LiteLLM! Whether you're fixing bugs, adding features, or improving documentation, we appreciate your help.

Quick Start for Contributors

This requires poetry to be installed.

git clone https://github.com/BerriAI/litellm.git
cd litellm
make install-dev    # Install development dependencies
make format         # Format your code
make lint           # Run all linting checks
make test-unit      # Run unit tests
make format-check   # Check formatting only

For detailed contributing guidelines, see CONTRIBUTING.md.

Code Quality / Linting

LiteLLM follows the Google Python Style Guide.

Our automated checks include:

  • Black for code formatting
  • Ruff for linting and code quality
  • MyPy for type checking
  • Circular import detection
  • Import safety checks

All these checks must pass before your PR can be merged.

Support / talk with founders

Why did we build this

  • Need for simplicity: Our code started to get extremely complicated managing & translating calls between Azure, OpenAI and Cohere.

Contributors