litellm/litellm/proxy
ishaan-berri 118ce3cc91
feat: add model leaderboard page (#43649)
* feat(proxy): add model leaderboard analytics

* feat: add model insights task and range constants

* feat: record task type from task tags in model usage rollup

* feat: serve 365 days of model insights by UTC date

* test: cover task tag resolution in model usage rollup

* test: update model insights range limit test to 365 days

* chore: regenerate dashboard api types for model insights

* feat: add model insights aggregation helpers

* test: cover model insights aggregation helpers

* feat: redesign model leaderboard with stacked bars, treemap and ranking

* test: update model leaderboard view test

* feat: mark model leaderboard as beta in sidebar

* chore: sync schema.prisma copies from root

* fix: only treat task: prefixed tags as model insight tasks

* feat: add metric type for model insights ranking

* fix: rank model insights by selected metric and scope detail queries to ranked deployments

* test: plain tags are not model insight tasks

* test: cover metric ranking, deployment scoping and rollup round trip

* fix: build model insights weeks and halves from the requested date range

* test: cover empty weeks and range-based change comparison

* fix: refetch by metric, show load errors and ignore stale responses

* test: cover metric refetch and error state

* feat: define model insight tasks in a JSON file

* feat: return task labels and categories from model insights

* feat: load model insight tasks from JSON

* refactor: validate rollup task tags against the JSON task list

* feat: serve the task list with model insights

* refactor: drop hardcoded task list from constants

* build: ship model insight tasks JSON in the wheel

* test: cover model insight task JSON

* refactor: take task labels and categories from the API

* test: pass task info to task tile builder

* refactor: color treemap by API-provided category

* test: include tasks in model leaderboard fixture

* fix: make daily model usage migration idempotent

* feat: bound the model insights task query size

* fix: compute task breakdown independent of the chart metric

* test: task breakdown is stable across chart metrics

* chore: regenerate lazy openapi snapshot for model insights

* chore: regenerate dashboard api types for model insights

* fix: keep previous ranking dimmed while a new metric loads

* test: cover stale metric state in model leaderboard

* refactor: drop task row cap constant

* fix: return the full task breakdown instead of a truncated one

* test: task query is not truncated

* feat: add task summary types for model insights

* feat: summarise tasks server-side on a separate model insights endpoint

* test: cover the model insights tasks endpoint

* chore: regenerate lazy openapi snapshot for model insights tasks

* chore: regenerate dashboard api types for model insights tasks

* refactor: drop client-side task aggregation

* test: remove client-side task aggregation tests

* feat: load task breakdown separately from the chart metric

* test: task breakdown is not refetched on chart metric change

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-28 19:40:04 -07:00
..
_experimental feat(mcp): scan and pin upstream tool descriptions (#43283) 2026-09-28 18:38:49 -07:00
a2a refactor(types): replace Any with proven types in 32 files 2026-09-21 10:51:46 +00:00
agent_endpoints feat(agents): add optional per-agent kill switch webhook (#42841) 2026-09-24 18:26:50 -05:00
analytics_endpoints Merge remote-tracking branch 'origin/main' into litellm_decrease_anys_opus5_r5 2026-09-14 07:27:54 +00:00
anthropic_endpoints refactor(anthropic): rename experimental_pass_through to pass_through (#43329) 2026-09-26 13:00:50 -07:00
auth fix(proxy): log key owner identity on expired key auth failures (#43105) 2026-09-28 08:18:48 -07:00
batches_endpoints fix(spend): attribute CLI session spend to the per-user cli-session alias instead of the hashed session token (#40541) 2026-09-24 18:21:47 -07:00
client feat(cli): reuse saved agent setup and add reconfigure (#43392) 2026-09-26 18:46:36 -07:00
common_utils feat(router): opt in to prompt-cache cost routing (#43232) 2026-09-26 18:13:13 -07:00
config_management_endpoints
config_resolvers fix: enforce disable_custom_api_keys from general_settings (#42437) 2026-09-22 13:28:31 -07:00
container_endpoints refactor(types): replace Any with real types across 54 more backend files 2026-09-08 12:41:07 +00:00
credential_endpoints fix(credentials): answer 409 on a name collision, let PATCH resolve values from model_id 2026-09-15 10:41:43 -07:00
custom_hooks
db feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
discovery_endpoints fix(proxy): move credentials hint helper into a leaf html_forms module to clear CodeQL cyclic import 2026-09-14 20:34:58 +00:00
enterprise_billing
example_config_yaml Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
fine_tuning_endpoints
google_endpoints
guardrails feat(mcp): scan and pin upstream tool descriptions (#43283) 2026-09-28 18:38:49 -07:00
health_check_utils fix(health): keep team public names to the owning team and let model_id win over model 2026-09-11 19:49:37 -07:00
health_endpoints feat(otel): add SigNoz preset for OpenTelemetry v2 (#43296) 2026-09-26 18:15:45 -07:00
hooks feat(proxy): add fail_closed_rate_limit_enforcement to reject requests with 503 while Redis rate limit counters are unreachable (#43251) 2026-09-26 12:11:26 -07:00
image_endpoints Merge remote-tracking branch 'origin/main' into litellm_fix_image_edits_bracketed_alias 2026-09-17 14:49:54 -07:00
list_api fix(proxy): keep the submitted body out of 422 validation errors (#43231) 2026-09-25 17:05:53 -07:00
logging_endpoints refactor(types): replace Any with proven types in 32 files 2026-09-21 10:51:46 +00:00
management_endpoints feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
management_helpers fix(key_management): count team unified access group MCP servers when validating key MCP grants (#41231) 2026-09-24 23:34:15 -07:00
memory feat(proxy): add a search param to key, memory, audit, and spend log listings 2026-09-03 15:20:13 -07:00
middleware fix(proxy): release unclaimed budget reservations at request end (#42304) 2026-09-21 19:51:12 -07:00
ocr_endpoints refactor(ocr): remove the Python OCR execution path and require the Rust route (#43081) 2026-09-24 18:18:50 -07:00
openai_evals_endpoints refactor(typing): replace Any with proven types in 65 backend files 2026-09-02 09:11:36 +00:00
openai_files_endpoints feat(vertex): native batch JSONL passthrough with cost tracking (#42810) 2026-09-24 12:35:34 -07:00
pass_through_endpoints refactor(anthropic): rename experimental_pass_through to pass_through (#43329) 2026-09-26 13:00:50 -07:00
policy_engine feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
prompts Merge pull request #38440 from BerriAI/litellm_prompt_registry_env 2026-09-03 14:09:26 -07:00
public_endpoints feat(providers): add Prism provider (internal copy of #40914) (#41961) 2026-09-28 16:07:09 -07:00
rag_endpoints feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
realtime_endpoints Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_openai_error_payload_non_llm_routes 2026-09-08 11:09:54 -07:00
rerank_endpoints fix(proxy): seed litellm_call_id into request data before parsing can fail 2026-09-16 20:55:32 +00:00
response_api_endpoints fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507) 2026-09-26 18:05:21 -07:00
response_polling fix(proxy): record response.failed frames in background polling 2026-09-18 15:58:15 -07:00
search_endpoints
shutdown style(proxy): format scheduled job timeout configuration 2026-09-21 19:43:11 +00:00
spend_tracking revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377) 2026-09-28 21:47:46 +00:00
swagger
test_prompts
types_utils refactor(types): replace Any with proven types in 5 files (#42722) 2026-09-23 03:14:45 -07:00
ui_crud_endpoints fix(proxy): keep the submitted body out of 422 validation errors (#43231) 2026-09-25 17:05:53 -07:00
vector_store_endpoints fix(vector_stores): keep config-defined vector stores listed and read-only (#42574) 2026-09-23 04:02:16 +00:00
vector_store_files_endpoints Merge remote-tracking branch 'origin/main' into litellm_decrease_anys_opus5_r5 2026-09-21 12:57:38 -07:00
vertex_ai_endpoints
video_endpoints refactor: replace Any with precise types across 54 modules 2026-09-14 11:13:39 +00:00
workflows docs: stop advertising sk-1234 as the master key in shipped configs and examples 2026-09-19 12:59:48 -07:00
.gitignore
__init__.py Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…" 2026-09-05 16:07:09 -07:00
_lazy_features.py feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
_lazy_openapi_snapshot.json feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
_lazy_openapi_snapshot.py
_new_new_secret_config.yaml
_new_secret_config.yaml docs: stop advertising sk-1234 as the master key in shipped configs and examples 2026-09-19 12:59:48 -07:00
_super_secret_config.yaml
_types.py revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377) 2026-09-28 21:47:46 +00:00
bug_report_config.py feat(proxy): admin-only /debug/report sharing the bug report environment (#42440) 2026-09-22 12:05:23 -07:00
cached_logo.jpg
caching_routes.py refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
collector.py feat(proxy): offload spend tracking to a pod-local collector sidecar (#40545) 2026-09-10 17:14:13 -07:00
common_request_processing.py refactor(types): replace Any with proven types in 8 files (#43551) 2026-09-28 04:26:41 -07:00
compliance_checks.py fix(guardrails): record not_run when a skipped role mixes text and images 2026-09-15 05:20:56 +00:00
custom_auth_auto.py
custom_prompt_management.py
custom_sso.py
custom_validate.py
dd_span_tagger.py
dev_config.yaml Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
enterprise
health_check.py feat(auto-router): integrate JEV context and usage accounting 2026-09-18 21:38:15 +00:00
lambda.py
litellm_pre_call_utils.py security(proxy): keep team callback credentials out of the stored request body (#43217) 2026-09-28 11:16:58 -07:00
llamaguard_prompt.txt
logo.jpg
logo_dark.png
mcp_registry.json fix(ui): restore MCP catalog provider logos (#40781) 2026-09-12 11:56:59 -07:00
mcp_tools.py
model_config.yaml
model_insights_tasks.json feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
native_compaction.py feat(router): native compact-to-fit across conversation APIs (#42074) 2026-09-21 22:52:29 -07:00
openapi.json
openapi_registry.json
plugin_routes.py refactor(proxy): resolve config and DB settings precedence in one SettingsStore 2026-09-17 23:36:27 -07:00
post_call_rules.py
prisma_migration.py fix(proxy): run migrations through python -m prisma when the prisma console script is not on PATH 2026-09-09 18:17:12 -07:00
prometheus_cleanup.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
prometheus_metrics_server.py feat(deploy): metrics sidecar and separate metrics port in Helm and Terraform (#40163) 2026-09-08 13:31:07 -07:00
proxy_cli.py feat(proxy_cli): add --validate_config dry-run flag (#41705) 2026-09-24 17:07:10 -05:00
proxy_config.yaml docs: stop advertising sk-1234 as the master key in shipped configs and examples 2026-09-19 12:59:48 -07:00
proxy_server.py revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377) 2026-09-28 21:47:46 +00:00
read_model_list.py
README.md
route_llm_request.py fix(proxy): return 400 instead of 500 for /v1/responses without input 2026-09-19 08:01:33 +00:00
route_priority.py perf(proxy): register liveness and core inference routes first (#40687) 2026-09-11 17:10:38 +00:00
schema.prisma feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
start.sh
utils.py feat(mcp): scan and pin upstream tool descriptions (#43283) 2026-09-28 18:38:49 -07:00
wildcard_config.yaml Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00

litellm-proxy

A local, fast, and lightweight OpenAI-compatible server to call 100+ LLM APIs.

usage

$ uv tool install litellm
$ litellm --model ollama/codellama 

#INFO: Ollama running on http://0.0.0.0:8000

replace openai base

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

See how to call Huggingface,Bedrock,TogetherAI,Anthropic, etc.


Folder Structure

Routes

  • proxy_server.py - all openai-compatible routes - /v1/chat/completion, /v1/embedding + model info routes - /v1/models, /v1/model/info, /v1/model_group_info routes.
  • health_endpoints/ - /health, /health/liveliness, /health/readiness
  • management_endpoints/key_management_endpoints.py - all /key/* routes
  • management_endpoints/team_endpoints.py - all /team/* routes
  • management_endpoints/internal_user_endpoints.py - all /user/* routes
  • management_endpoints/ui_sso.py - all /sso/* routes