litellm/litellm/proxy
haydster7 ef284bbaca
fix(proxy): resolve model aliases to their target deployment for listing metadata
A model alias (a key's `aliases`, a team's `model_aliases`, or the router's
`model_group_alias`) is rewritten at request time and owns no deployment row,
so `get_model_listing_info` returns None for it and the alias name is not a
cost-map key either. Both metadata candidate sources come up empty and the
listing entry carries no `mode`, `max_input_tokens`, or `max_output_tokens`,
while the deployment it routes to reports them on the same response.

Resolve a listed alias to the deployment it routes to for the metadata lookup,
keeping the alias as the response id. Maps are applied in order, each against
the result of the last, mirroring how `litellm_pre_call_utils` applies the team
map and then the key map, so a chained alias reports the deployment the caller
actually reaches rather than an intermediate one. An alias shadowing a
same-named deployment follows the alias, because the request rewrite does too.

Fixes the model_aliases half of #40553.
2026-09-18 12:35:56 +10:00
..
_experimental fix(mcp): retain selected guardrails for virtual REST calls 2026-09-17 10:40:33 -07:00
a2a refactor(typing): replace Any with proven types in 65 backend files 2026-09-02 09:11:36 +00:00
agent_endpoints fix(a2a): forward caller identity headers on message/send and message/stream (#40305) 2026-09-08 15:56:52 -07:00
analytics_endpoints fix(proxy): attribute gate-rejected requests to their endpoint in cache analytics (#40824) 2026-09-12 16:03:38 -07:00
anthropic_endpoints fix(proxy): keep litellm_call_id on shaped errors and list_batches failure hook 2026-09-16 03:02:50 +00:00
auth Merge pull request #41684 from BerriAI/litellm_wildcard_license_auto_router 2026-09-17 16:11:40 -07:00
batches_endpoints fix(proxy): build failure headers immutably to keep LIT002 within budget 2026-09-16 18:51:22 +00:00
client Merge pull request #41672 from BerriAI/litellm_autoroute_start_stop 2026-09-17 16:07:00 -07:00
common_utils fix(proxy): resolve model aliases to their target deployment for listing metadata 2026-09-18 12:35:56 +10:00
config_management_endpoints
config_resolvers feat(alerting): add native Microsoft Teams alerting destination (#38367) 2026-08-27 16:19:22 -07:00
container_endpoints chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
credential_endpoints fix(credentials): answer 409 on a name collision, let PATCH resolve values from model_id 2026-09-15 10:41:43 -07:00
custom_hooks
db refactor(spend_tracking): type the get_logging_payload parameters 2026-09-17 01:02:03 +00:00
discovery_endpoints fix(proxy): move credentials hint helper into a leaf html_forms module to clear CodeQL cyclic import 2026-09-14 20:34:58 +00:00
enterprise_billing
example_config_yaml
fine_tuning_endpoints
google_endpoints
guardrails Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context 2026-09-17 00:31:46 -07:00
health_check_utils fix(health): keep team public names to the owning team and let model_id win over model 2026-09-11 19:49:37 -07:00
health_endpoints fix(health): skip background health check DB writes when the latest-row read fails 2026-09-14 23:03:30 +00:00
hooks fix(responses): keep the addressed response id off bridged provider requests 2026-09-17 15:35:11 -07:00
image_endpoints Merge remote-tracking branch 'origin/main' into litellm_fix_image_edits_bracketed_alias 2026-09-17 14:49:54 -07:00
list_api Merge branch 'main' into litellm_bulk_user_delete 2026-09-15 18:03:15 +00:00
logging_endpoints fix(usage): recover aliases for v1.99 double-hashed spend keys 2026-09-03 14:21:25 +00:00
management_endpoints Merge pull request #41694 from BerriAI/litellm_rename_model_sync_allowlists 2026-09-17 17:42:06 -07:00
management_helpers Merge pull request #41694 from BerriAI/litellm_rename_model_sync_allowlists 2026-09-17 17:42:06 -07:00
memory feat(proxy): add a search param to key, memory, audit, and spend log listings 2026-09-03 15:20:13 -07:00
middleware feat(proxy): resolve root_path per request from a configured prefix list (SERVER_ROOT_PATHS) (#35935) 2026-09-07 11:54:59 -07:00
ocr_endpoints merge: origin/main into litellm_rust_bridge_declarative_route_catalog 2026-09-16 22:54:51 +00:00
openai_evals_endpoints refactor(typing): replace Any with proven types in 65 backend files 2026-09-02 09:11:36 +00:00
openai_files_endpoints refactor(proxy): drop redundant docstrings from upload allowlist helpers and tests 2026-09-14 19:00:49 +00:00
pass_through_endpoints Merge remote-tracking branch 'origin/main' into litellm_/circleci-specific-sha-0cf414 2026-09-17 18:05:39 -07:00
policy_engine feat(policy_engine): explicit priority for policy attachment execution order 2026-09-17 05:59:01 +00:00
prompts Merge pull request #38440 from BerriAI/litellm_prompt_registry_env 2026-09-03 14:09:26 -07:00
public_endpoints feat(auto-router): refresh family reasoning presets (#40341) 2026-09-08 18:28:46 -07:00
rag_endpoints fix(proxy): surface the upstream status code when a RAG query fails 2026-09-15 17:53:57 -07:00
realtime_endpoints Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_openai_error_payload_non_llm_routes 2026-09-08 11:09:54 -07:00
rerank_endpoints fix(proxy): seed litellm_call_id into request data before parsing can fail 2026-09-16 20:55:32 +00:00
response_api_endpoints refactor(types): replace Any with precise types across 73 modules 2026-09-01 11:05:02 +00:00
response_polling fix(responses): keep background polling alive after the client disconnects 2026-09-07 11:51:35 +00:00
search_endpoints Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5 2026-08-29 06:03:33 -07:00
shutdown
spend_tracking Merge pull request #41569 from BerriAI/litellm_azure_ptu_spillover_cost 2026-09-17 17:09:40 -07:00
swagger
test_prompts
types_utils
ui_crud_endpoints fix(proxy): re-read UI settings on every config reload 2026-09-16 15:38:06 -07:00
vector_store_endpoints fix(proxy): let authorized internal users open vector store details (#40150) 2026-09-07 13:10:09 -07:00
vector_store_files_endpoints
vertex_ai_endpoints
video_endpoints refactor(types): replace Any with precise types across 73 modules 2026-09-01 11:05:02 +00:00
workflows
.gitignore
__init__.py Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…" 2026-09-05 16:07:09 -07:00
_lazy_features.py feat(proxy): add TypeSafe Jev passthrough spend tracking 2026-09-17 15:53:19 +00:00
_lazy_openapi_snapshot.json chore(proxy): regenerate the OpenAPI snapshot with the CI Python version 2026-09-17 18:15:55 -07:00
_lazy_openapi_snapshot.py chore(techdebt): type new signatures and drop slop comments from the last 24h 2026-08-28 07:59:32 +00:00
_new_new_secret_config.yaml
_new_secret_config.yaml
_super_secret_config.yaml
_types.py Merge pull request #41569 from BerriAI/litellm_azure_ptu_spillover_cost 2026-09-17 17:09:40 -07:00
cached_logo.jpg
caching_routes.py refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
collector.py feat(proxy): offload spend tracking to a pod-local collector sidecar (#40545) 2026-09-10 17:14:13 -07:00
common_request_processing.py fix(proxy): reject non-string model with 400 and log its spend as unknown-model 2026-09-17 19:37:14 +00:00
compliance_checks.py fix(guardrails): record not_run when a skipped role mixes text and images 2026-09-15 05:20:56 +00:00
custom_auth_auto.py
custom_prompt_management.py
custom_sso.py
custom_validate.py
dd_span_tagger.py
dev_config.yaml
enterprise
health_check.py fix(health): resolve a model name the way a request routes before matching the provider model string 2026-09-11 20:40:00 -07:00
lambda.py
litellm_pre_call_utils.py Merge pull request #41330 from BerriAI/litellm_team_model_max_budget_v2 2026-09-16 14:48:29 -07:00
llamaguard_prompt.txt
logo.jpg
logo_dark.png
mcp_registry.json fix(ui): restore MCP catalog provider logos (#40781) 2026-09-12 11:56:59 -07:00
mcp_tools.py
model_config.yaml
openapi.json
openapi_registry.json
plugin_routes.py
post_call_rules.py
prisma_migration.py fix(proxy): run migrations through python -m prisma when the prisma console script is not on PATH 2026-09-09 18:17:12 -07:00
prometheus_cleanup.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
prometheus_metrics_server.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
proxy_cli.py fix(proxy): resolve LITELLM_LOG for uvicorn at startup and restore logger state in tests 2026-09-15 22:27:24 +00:00
proxy_config.yaml
proxy_server.py fix(proxy): resolve model aliases to their target deployment for listing metadata 2026-09-18 12:35:56 +10:00
read_model_list.py
README.md
route_llm_request.py fix(proxy): keep the raw model string out of the unknown-model spend-log error message 2026-09-11 18:42:08 -07:00
route_priority.py perf(proxy): register liveness and core inference routes first (#40687) 2026-09-11 17:10:38 +00:00
schema.prisma Merge pull request #41571 from BerriAI/litellm_policy_attachment_priority 2026-09-17 17:27:55 -07:00
start.sh
utils.py fix(mcp): preserve request-selected guardrails during tool execution 2026-09-17 09:57:08 -07:00
wildcard_config.yaml

litellm-proxy

A local, fast, and lightweight OpenAI-compatible server to call 100+ LLM APIs.

usage

$ uv tool install litellm
$ litellm --model ollama/codellama 

#INFO: Ollama running on http://0.0.0.0:8000

replace openai base

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

See how to call Huggingface,Bedrock,TogetherAI,Anthropic, etc.


Folder Structure

Routes

  • proxy_server.py - all openai-compatible routes - /v1/chat/completion, /v1/embedding + model info routes - /v1/models, /v1/model/info, /v1/model_group_info routes.
  • health_endpoints/ - /health, /health/liveliness, /health/readiness
  • management_endpoints/key_management_endpoints.py - all /key/* routes
  • management_endpoints/team_endpoints.py - all /team/* routes
  • management_endpoints/internal_user_endpoints.py - all /user/* routes
  • management_endpoints/ui_sso.py - all /sso/* routes