litellm/litellm/proxy
Claude 806b8be451
fix(mcp,jwt): address greptile review concerns
- Cache _get_agent_object_permission via user_api_key_cache (sentinel for
  no-permission rows) so MCP requests from agent keys don't hit the DB on
  every tool-list / tool-call.
- Re-raise HTTPException in handle_sse_mcp so 401 + WWW-Authenticate
  challenges (and other HTTP errors) propagate to SSE clients instead of
  being swallowed as 500.
- Normalise booleans in _validate_token_response so admin rules written as
  JSON-style "true" / "false" match upstream responses that return
  Python True / False.
- Treat configured JWT issuer claim mappings as advisory: when a mapped
  field is absent or empty, leave the normalised claim unset instead of
  raising, matching the global litellm_jwtauth path.

Co-authored-by: Claude <noreply@anthropic.com>
2026-05-20 15:08:54 +00:00
..
_experimental fix(mcp,jwt): address greptile review concerns 2026-05-20 15:08:54 +00:00
agent_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
analytics_endpoints fix(spend): session-TZ-independent date filtering for spend/error log queries 2026-04-10 17:04:52 -07:00
anthropic_endpoints style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
auth fix(mcp,jwt): address greptile review concerns 2026-05-20 15:08:54 +00:00
batches_endpoints fix(proxy/batches): forward model to retrieve_batch for bedrock 2026-04-29 22:48:03 +02:00
client [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
common_utils [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
config_management_endpoints
container_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
credential_endpoints style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
custom_hooks style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
db [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
discovery_endpoints feat: add control plane for multi-proxy worker management 2026-03-19 22:50:19 -07:00
example_config_yaml [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
fine_tuning_endpoints style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
google_endpoints run pre_call_hook on Google generateContent endpoints 2026-04-30 16:43:42 -07:00
guardrails [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
health_check_utils Optimize database query which fetches latest model_id, model_name pairs and dedupes them in memory. 2026-04-15 00:54:37 +00:00
health_endpoints fix(proxy): expose db status on public /health/readiness 2026-05-13 13:18:54 -07:00
hooks [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
image_endpoints fix: tighten file input handling in image edit endpoints 2026-04-22 18:04:39 -07:00
management_endpoints fix(mcp): address oauth passthrough review findings 2026-05-15 21:05:40 +01:00
management_helpers [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
memory Litellm memory improvements v2 (#26541) 2026-04-25 19:03:43 -07:00
middleware fix(proxy): point /metrics 401 at the opt-out flag 2026-05-08 18:42:14 -07:00
ocr_endpoints style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
openai_evals_endpoints style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
openai_files_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
pass_through_endpoints Merge pull request #27801 from stuxf/chore/get-instance-fn-runtime-s3-gate 2026-05-13 20:57:24 -07:00
policy_engine [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
prompts fix: harden /key/update authorization checks (#27878) 2026-05-13 21:33:14 -07:00
public_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
rag_endpoints Merge remote-tracking branch 'origin/litellm_1.84.0rc2' into backport/27878-litellm_1.84.0rc2 2026-05-13 21:39:17 -07:00
realtime_endpoints [Fix] Use type:ignore instead of Union return type for realtime endpoint 2026-03-13 11:42:26 -07:00
rerank_endpoints
response_api_endpoints fix: address Greptile review comments 2026-03-19 14:10:58 +05:30
response_polling style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
search_endpoints feat(proxy): move search tool access to object permissions 2026-04-29 12:29:20 +05:30
spend_tracking fix(proxy): gate image-gen reservation strictly on model mode 2026-05-09 10:29:55 -07:00
swagger
test_prompts
types_utils Merge pull request #27801 from stuxf/chore/get-instance-fn-runtime-s3-gate 2026-05-13 20:57:24 -07:00
ui_crud_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
vector_store_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
vector_store_files_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
vertex_ai_endpoints [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
video_endpoints [Staging] - Ishaan March 17th (#23903) 2026-03-18 15:09:01 -07:00
workflows feat(proxy): durable agent workflow run tracking via /v1/workflows/runs (#26793) 2026-04-29 17:12:18 -07:00
.gitignore
__init__.py
_lazy_features.py Cache normalized SERVER_ROOT_PATH at middleware init 2026-05-12 23:15:19 -07:00
_lazy_openapi_snapshot.json [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
_lazy_openapi_snapshot.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
_logging.py
_new_new_secret_config.yaml
_new_secret_config.yaml Litellm krrish staging 04 20 2026 (#26138) 2026-04-20 16:22:12 -07:00
_super_secret_config.yaml fix: prompt registry 2026-02-18 00:34:54 +05:30
_types.py feat(proxy): support issuer-scoped JWT auth 2026-05-15 17:49:29 +01:00
cached_logo.jpg fix: prompt registry 2026-02-18 00:34:54 +05:30
caching_routes.py
common_request_processing.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
compliance_checks.py Add compliance checker endpoints + UI panel (#21432) 2026-02-17 18:22:26 -08:00
custom_auth_auto.py
custom_prompt_management.py fix: prompt registry 2026-02-18 00:34:54 +05:30
custom_sso.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
custom_validate.py
dd_span_tagger.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
enterprise
health_check.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
lambda.py
litellm_pre_call_utils.py docs(proxy): refresh stale comments referencing removed tag strip 2026-05-12 16:34:42 -07:00
llamaguard_prompt.txt
logo.jpg
mcp_registry.json feat(ui): group MCP tools by CRUD risk category in allowlist panels (#23403) 2026-03-11 21:15:25 -07:00
mcp_tools.py
model_config.yaml
openapi.json
openapi_registry.json feat(ui): OpenAPI MCP server support with popular API quick-picker (#23200) 2026-03-10 13:59:52 -07:00
post_call_rules.py
prisma_migration.py
prometheus_cleanup.py Add Prometheus child_exit cleanup for gunicorn workers 2026-02-27 16:11:15 -08:00
proxy_cli.py feat(proxy): add --timeout_worker_healthcheck flag for uvicorn worker triage 2026-04-27 11:06:56 -07:00
proxy_config.yaml fix: prompt registry 2026-02-18 00:34:54 +05:30
proxy_server.py style(mcp): satisfy black formatting 2026-05-15 17:49:29 +01:00
README.md build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
route_llm_request.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
schema.prisma [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
start.sh
utils.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00

litellm-proxy

A local, fast, and lightweight OpenAI-compatible server to call 100+ LLM APIs.

usage

$ uv tool install litellm
$ litellm --model ollama/codellama 

#INFO: Ollama running on http://0.0.0.0:8000

replace openai base

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

See how to call Huggingface,Bedrock,TogetherAI,Anthropic, etc.


Folder Structure

Routes

  • proxy_server.py - all openai-compatible routes - /v1/chat/completion, /v1/embedding + model info routes - /v1/models, /v1/model/info, /v1/model_group_info routes.
  • health_endpoints/ - /health, /health/liveliness, /health/readiness
  • management_endpoints/key_management_endpoints.py - all /key/* routes
  • management_endpoints/team_endpoints.py - all /team/* routes
  • management_endpoints/internal_user_endpoints.py - all /user/* routes
  • management_endpoints/ui_sso.py - all /sso/* routes