litellm/litellm/proxy
yucheng-berri 4301ac4940
Merge pull request #41986 from BerriAI/litellm_revert_post_call_guardrail_context
revert(guardrails): drop the scoped request conversation and tools from post-call scans (#41220)
2026-09-19 12:06:30 -07:00
..
_experimental fix(mcp): preserve session expiry signals and scope dependency CI 2026-09-18 22:52:10 -07:00
a2a
agent_endpoints refactor(a2a): resolve the relay's Entra hop bearer inside the a2a provider helper 2026-09-16 17:56:42 -07:00
analytics_endpoints fix(proxy): attribute gate-rejected requests to their endpoint in cache analytics (#40824) 2026-09-12 16:03:38 -07:00
anthropic_endpoints fix(claude_code_gateway): mint the bearer before consuming the device code so a signing failure never spends the login 2026-09-18 18:25:55 -07:00
auth Merge pull request #41924 from BerriAI/litellm_role_permissions_normalization 2026-09-18 22:15:20 -07:00
batches_endpoints fix(batches): price model-encoded batch retrievals by their deployment 2026-09-19 00:30:15 -07:00
client refactor(proxy): resolve config and DB settings precedence in one SettingsStore 2026-09-17 23:36:27 -07:00
common_utils Merge pull request #40243 from zoroyihan7/fix-responses-stream-error-events 2026-09-18 17:29:57 -07:00
config_management_endpoints
config_resolvers feat(proxy): say when a stored setting is ignored because the config file owns it 2026-09-19 10:20:31 -07:00
container_endpoints chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
credential_endpoints fix(credentials): answer 409 on a name collision, let PATCH resolve values from model_id 2026-09-15 10:41:43 -07:00
custom_hooks
db Merge pull request #41878 from BerriAI/litellm_requeue_daily_spend_without_redis_buffer 2026-09-18 17:27:36 -07:00
discovery_endpoints fix(proxy): move credentials hint helper into a leaf html_forms module to clear CodeQL cyclic import 2026-09-14 20:34:58 +00:00
enterprise_billing
example_config_yaml
fine_tuning_endpoints
google_endpoints
guardrails Merge pull request #41986 from BerriAI/litellm_revert_post_call_guardrail_context 2026-09-19 12:06:30 -07:00
health_check_utils fix(health): keep team public names to the owning team and let model_id win over model 2026-09-11 19:49:37 -07:00
health_endpoints fix(health): skip background health check DB writes when the latest-row read fails 2026-09-14 23:03:30 +00:00
hooks fix(proxy): keep request metadata out of the cost tracking failure alert 2026-09-19 03:21:28 -07:00
image_endpoints Merge remote-tracking branch 'origin/main' into litellm_fix_image_edits_bracketed_alias 2026-09-17 14:49:54 -07:00
list_api Merge branch 'main' into litellm_bulk_user_delete 2026-09-15 18:03:15 +00:00
logging_endpoints fix(usage): recover aliases for v1.99 double-hashed spend keys 2026-09-03 14:21:25 +00:00
management_endpoints fix(ui): stop Top Virtual Keys from opening keys that are not in the database 2026-09-19 11:08:21 -07:00
management_helpers Merge remote-tracking branch 'origin/main' into litellm_team_member_budget_link_default 2026-09-18 23:05:03 +00:00
memory feat(proxy): add a search param to key, memory, audit, and spend log listings 2026-09-03 15:20:13 -07:00
middleware fix(claude_code_gateway): scope the protobuf body skip to the OTLP routes and match the metrics middleware on the route path 2026-09-18 17:23:59 -07:00
ocr_endpoints merge: origin/main into litellm_rust_bridge_declarative_route_catalog 2026-09-16 22:54:51 +00:00
openai_evals_endpoints
openai_files_endpoints Merge remote-tracking branch 'origin/main' into litellm_unknown_model_spend_logs_outside_router 2026-09-19 04:43:19 -07:00
pass_through_endpoints fix(proxy): keep the raw client model out of spend logs for rejections outside the router 2026-09-19 02:01:28 -07:00
policy_engine feat(policy_engine): explicit priority for policy attachment execution order 2026-09-17 05:59:01 +00:00
prompts Merge pull request #38440 from BerriAI/litellm_prompt_registry_env 2026-09-03 14:09:26 -07:00
public_endpoints feat(rust): add Amazon Textract to litellm.ocr and sign provider requests after host hooks 2026-09-19 09:11:52 -07:00
rag_endpoints fix(rag): read a registered S3 Vectors store's bucket and index from its id 2026-09-19 03:38:20 -07:00
realtime_endpoints Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_openai_error_payload_non_llm_routes 2026-09-08 11:09:54 -07:00
rerank_endpoints fix(proxy): seed litellm_call_id into request data before parsing can fail 2026-09-16 20:55:32 +00:00
response_api_endpoints fix(proxy): validate responses input after prompt template expansion 2026-09-19 16:34:53 +00:00
response_polling fix(proxy): record response.failed frames in background polling 2026-09-18 15:58:15 -07:00
search_endpoints
shutdown
spend_tracking fix(ui): stop Top Virtual Keys from opening keys that are not in the database 2026-09-19 11:08:21 -07:00
swagger
test_prompts
types_utils
ui_crud_endpoints feat(proxy): say when a stored setting is ignored because the config file owns it 2026-09-19 10:20:31 -07:00
vector_store_endpoints fix(proxy): let authorized internal users open vector store details (#40150) 2026-09-07 13:10:09 -07:00
vector_store_files_endpoints fix(vector_stores): run the full model grant check on caller-supplied model hints 2026-09-19 00:54:41 -07:00
vertex_ai_endpoints
video_endpoints
workflows
.gitignore
__init__.py Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…" 2026-09-05 16:07:09 -07:00
_lazy_features.py Merge pull request #34267 from BerriAI/litellm_claude_code_gateway_protocol 2026-09-18 20:49:48 -07:00
_lazy_openapi_snapshot.json fix(ui): stop Top Virtual Keys from opening keys that are not in the database 2026-09-19 11:08:21 -07:00
_lazy_openapi_snapshot.py
_new_new_secret_config.yaml
_new_secret_config.yaml
_super_secret_config.yaml
_types.py Merge pull request #41485 from BerriAI/litellm_jwt_token_exchange_grant 2026-09-18 21:27:20 -07:00
cached_logo.jpg
caching_routes.py refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
collector.py feat(proxy): offload spend tracking to a pod-local collector sidecar (#40545) 2026-09-10 17:14:13 -07:00
common_request_processing.py Merge pull request #41583 from BerriAI/litellm_applied_guardrails_blocker 2026-09-18 16:00:27 -07:00
compliance_checks.py fix(guardrails): record not_run when a skipped role mixes text and images 2026-09-15 05:20:56 +00:00
custom_auth_auto.py
custom_prompt_management.py
custom_sso.py
custom_validate.py
dd_span_tagger.py
dev_config.yaml
enterprise
health_check.py fix(health): resolve a model name the way a request routes before matching the provider model string 2026-09-11 20:40:00 -07:00
lambda.py
litellm_pre_call_utils.py fix(timing): use epoch math for detailed pre-processing and drop client-supplied timing windows 2026-09-19 01:06:53 +00:00
llamaguard_prompt.txt
logo.jpg
logo_dark.png
mcp_registry.json fix(ui): restore MCP catalog provider logos (#40781) 2026-09-12 11:56:59 -07:00
mcp_tools.py
model_config.yaml
openapi.json
openapi_registry.json
plugin_routes.py refactor(proxy): resolve config and DB settings precedence in one SettingsStore 2026-09-17 23:36:27 -07:00
post_call_rules.py
prisma_migration.py fix(proxy): run migrations through python -m prisma when the prisma console script is not on PATH 2026-09-09 18:17:12 -07:00
prometheus_cleanup.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
prometheus_metrics_server.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
proxy_cli.py chore(proxy): drop a comment that restated the NUM_WORKERS assignment 2026-09-17 01:56:05 +00:00
proxy_config.yaml
proxy_server.py Merge pull request #41920 from BerriAI/litellm_claude_auto_cache_providers 2026-09-19 12:54:30 -05:00
read_model_list.py
README.md
route_llm_request.py fix(proxy): return 400 instead of 500 for /v1/responses without input 2026-09-19 08:01:33 +00:00
route_priority.py perf(proxy): register liveness and core inference routes first (#40687) 2026-09-11 17:10:38 +00:00
schema.prisma Merge branch 'litellm_usage_key_free_aggregate_split' into litellm_daily_global_spend_table 2026-09-18 20:06:47 +00:00
start.sh
utils.py fix(proxy): keep queued moderation running past a V1 pre_call guardrail 2026-09-18 23:33:29 +00:00
wildcard_config.yaml

litellm-proxy

A local, fast, and lightweight OpenAI-compatible server to call 100+ LLM APIs.

usage

$ uv tool install litellm
$ litellm --model ollama/codellama 

#INFO: Ollama running on http://0.0.0.0:8000

replace openai base

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

See how to call Huggingface,Bedrock,TogetherAI,Anthropic, etc.


Folder Structure

Routes

  • proxy_server.py - all openai-compatible routes - /v1/chat/completion, /v1/embedding + model info routes - /v1/models, /v1/model/info, /v1/model_group_info routes.
  • health_endpoints/ - /health, /health/liveliness, /health/readiness
  • management_endpoints/key_management_endpoints.py - all /key/* routes
  • management_endpoints/team_endpoints.py - all /team/* routes
  • management_endpoints/internal_user_endpoints.py - all /user/* routes
  • management_endpoints/ui_sso.py - all /sso/* routes