Commit graph

51406 commits

Author SHA1 Message Date
kerry
2586b21893 test: keep tests that survive correct cost-map updates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:16:08 +00:00
yassin
9261a72eb7 test(passthrough): cover Vertex router resolution without a matching deployment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:14:44 +00:00
yassin
c05af7a12f feat(ui): add custom request headers to the API Playground
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:12:06 +00:00
yassin
6c8b9a7b05 feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info
Persist and expose the lifecycle of API keys so spend, audit and FinOps
workflows can still resolve a key after it is revoked, expires or is
deleted.

/key/list?status= now accepts active, expired and revoked next to the
existing deleted value. revoked means blocked=true, expired means not
blocked with a past expiry, active is the rest, so the three values
partition the live key table. deleted keeps reading the
LiteLLM_DeletedVerificationToken archive.

/key/info falls back to that archive when the key is no longer in the
live table, running the same owner/team/org authorization check, and
every response now carries a derived status field. The hashed token is
still stripped.

The Virtual Keys page gets a Status filter (URL-persisted) and a Deleted
badge that shows when and by whom the key was deleted.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:12:01 +00:00
yassin
c3e937b845 docs(proxy): describe the single UNION ALL aggregate statement in the endpoint docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:11:20 +00:00
yassin
69be041da7 feat(proxy): add /nvidia_nim passthrough route for NIM object detection and OCR /v1/infer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:11:08 +00:00
yassin
af4a0b4bc3 fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay
Resolves LIT-2053

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:10:42 +00:00
yassin
460a0128c9 refactor(proxy): drop docstrings from key model limit helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:08:39 +00:00
yassin
bae2bf003e fix(proxy): resolve router_settings.model_group_alias before key/team model auth
Key and team router_settings.model_group_alias aliases were resolved only after the key/team model allowlist checks ran, so a key allowed the alias target was denied when it requested the alias. Resolve the alias during auth and rewrite the request body to the target before the allowlist checks. The alias the client sent is kept in the request scope so the response model still echoes it.

Resolves LIT-3054

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:06:55 +00:00
Yassin Kortam
226b1e1bb9
Merge pull request #41288 from BerriAI/litellm_langsmith_preserve_events_during_flush
fix(langsmith): keep events appended during an in-flight flush instead of clearing them
2026-09-15 15:06:36 -07:00
yassin
168b5bc4fb feat(proxy): configurable client-facing model access denied message
Add litellm_settings.model_access_denied_message, a template ({model} placeholder) returned to clients instead of the detailed "can only access models=[...]" text on key/team/user/org/project and team-member model access denials. The full denial reason is still written to the proxy logs at WARNING. Unset keeps the existing detailed message, status codes and error types are unchanged.

Expose the new setting and the existing expose_router_debug_in_errors flag in the Admin UI general settings (String editor, Boolean toggle with an explicit True default) and allow both as safe DB overrides so they persist and propagate across workers.

Resolves LIT-5283

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:05:08 +00:00
Yassin Kortam
e36626174b
Merge pull request #41297 from BerriAI/litellm_return_400_on_lone_surrogate_input
fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body
2026-09-15 15:04:08 -07:00
Yassin Kortam
579a30de83
Merge pull request #41300 from BerriAI/litellm_lit_1742_custom_provider_map_router
fix(router): accept custom_provider_map providers before the first completion call
2026-09-15 15:03:42 -07:00
Yassin Kortam
ee51be4db2
Merge pull request #41289 from BerriAI/litellm_fix_clientside_credential_deployment_scope
fix(router): stop registering a caller-supplied credential as a router deployment
2026-09-15 15:00:50 -07:00
yassin
c49fb1dd9d fix(openai): drop env var read for openai_system_messages_first, config and Admin UI set it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:58:35 +00:00
yassin
5237fe4df3 refactor(passthrough): read the deployment model_info request state in two steps
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:57:12 +00:00
Yassin Kortam
1b8daf20e0
Merge pull request #38268 from BerriAI/litellm_fix_xai_web_search_nested_filters
fix(xai): honor nested web_search filters on the xAI Responses API
2026-09-15 14:56:36 -07:00
yassin
3cb64978d4 fix(proxy): honor LITELLM_LOG for uvicorn and proxy extras loggers
LITELLM_LOG=ERROR still printed INFO lines from uvicorn (startup and access log) and from the litellm_proxy_extras migration logger, because neither read the variable. Forward the resolved level to uvicorn when LITELLM_LOG is set and no explicit log_config or JSON logging is in use, and let the extras logger take its level from LITELLM_LOG

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:55:37 +00:00
yassin
99fa38504a fix(ui): bound Models table page, page size and sort_by read from the URL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:55:07 +00:00
yassin
77d913958d feat(openai): add openai_system_messages_first to put system messages first for prompt caching
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:53:29 +00:00
yassin
f94a40f841 perf(proxy): serve key-free rollups and top-N keys from one UNION ALL statement
Both arms now run in a single query_raw call so totals and per-key
breakdowns come from the same snapshot. USAGE_TOP_API_KEYS_LIMIT can be
raised via env for deployments that need every key in the response.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:52:37 +00:00
yassin
51a4cb9fdd fix(proxy): key model rpm/tpm override takes precedence over team model limit
A key inside a team with model_rpm_limit / model_tpm_limit in team metadata could not
override those limits for itself: the v3 limiter always added the team's per-model
descriptor next to the key's, so the tighter team limit won. The docs already say the
resolution order is key metadata > key model_max_budget > team metadata

get_key_own_model_rate_limit returns only what the key sets on itself, and the team
descriptor now carries only the metrics the key does not override, so an rpm-only
override still leaves the team tpm pool enforced

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:48:06 +00:00
Yassin Kortam
fa09de9e45
Merge pull request #41279 from BerriAI/litellm_reset_budget_decrement_spend
fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows
2026-09-15 14:45:03 -07:00
kerry
a428cc8d84 test: restore runtime-derived tests dropped by mistake
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:44:39 +00:00
Yassin Kortam
c1e1d62903
Merge pull request #41291 from BerriAI/litellm_lit1643_auth_failure_user_agent
fix(proxy): keep client User-Agent on auth failure spend logs
2026-09-15 14:44:32 -07:00
yassin
0cc6968495 fix(router): accept custom_provider_map providers before the first completion call
get_llm_provider() and Router._add_deployment() only knew the built-in
provider_list and JSON providers, so a provider registered through
litellm.custom_provider_map was rejected until custom_llm_setup() had
run inside the first completion() call. Both now check the map directly.

Resolves LIT-1742

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:43:47 +00:00
tin-berri
ce1a4f896a
Merge pull request #41272 from BerriAI/litellm_fuse_v2_classifier_pr
feat(router): add Fuse V2 classifier after capability forecasting
2026-09-15 14:43:44 -07:00
yassin
e62ff9ebee fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:40:53 +00:00
yassin
96bffb1290 fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment
The Vertex passthrough route resolved a router deployment only to rewrite the
upstream URL and dropped its model_info, so the standard logging payload and
the Prometheus litellm_deployment_success_responses_total counter carried
model_id="". Carry the deployment's model_info through request.state into the
passthrough logging metadata, where it overrides any client-supplied model_info.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:39:39 +00:00
yassin
868d3855ab feat(ui): persist Models table search, filters, sort and page in the URL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:37:09 +00:00
kerry
62aa21e810 test: drop remaining tests that pin cost-map vendor facts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:34:00 +00:00
Yassin Kortam
140229bc4a
Merge pull request #41230 from BerriAI/litellm_client_timeout_408_skips_cooldown
fix(router): stop counting caller-set timeout 408s toward deployment cooldown
2026-09-15 14:31:11 -07:00
yassin
a0311dddf7 chore(proxy): regenerate lazy OpenAPI snapshot for api_key_limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:28:09 +00:00
yassin
afb6f8be65 fix(xai): treat an explicit empty web_search filters object as unrestricted
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:24:32 +00:00
Yassin Kortam
a6127d2363
Merge pull request #35180 from BerriAI/litellm_lit4995_vertex_rerank_search_units
fix(rerank): bill Vertex search_units from input records and give every rerank response a unique id
2026-09-15 14:23:23 -07:00
Yassin Kortam
5df127d483 fix(router): stop registering a caller-supplied credential as a router deployment
_handle_clientside_credential registered the per-request Deployment it built for
a client-supplied api_key/api_base via upsert_deployment, which added it to
self.model_list under the shared model_name. That made a request-scoped
credential a permanent, load-balanced deployment that any later caller of the
same model group could be routed onto, reaching the provider with someone
else's forwarded credential.

The per-request Deployment still gets its own stable id for cooldown and
logging identity; it is just never registered with the router.

Resolves LIT-7811
2026-09-15 14:18:03 -07:00
mrinal
4c179f2f59 test(langsmith): type the flush race test and cancel its periodic task
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:17:36 +00:00
yassin
92e55b3b22 perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:17:00 +00:00
Yassin Kortam
d35af8d302
Merge pull request #38682 from BerriAI/devin_ai_1787937848_tf_team_member_fields
feat(terraform): add tpm_limit, rpm_limit, budget_duration, allowed_models to litellm_team_member_add
2026-09-15 14:12:25 -07:00
yassin
595bec46ff fix(router): time fallback-hop 408s against now, not the previous hop's end_time
The failure logger skips fallback hops (has_logged_async_failure is already set), so
model_call_details.end_time still belongs to the previous hop and predates this hop's
api_call_start_time. The fallback cooldown guard measured a negative elapsed time and
cooled down deployments for caller-set timeouts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:11:40 +00:00
yassin
2f719fec52 test(router): run the fallback provider-408 cooldown regression inside an event loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:11:40 +00:00
yassin
5da497f4ac fix(router): only exempt 408s that arrive after the caller's timeout from cooldown
client_side_timeout records that the caller configured a timeout, not that
the timeout fired. A 408 the provider returns before that deadline is a
deployment failure and must still count toward cooldown.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:11:40 +00:00
ryan-crabbe-berri
9b35954347 fix(ui): block usage export and flag the range when a spend page fails
The Usage page drains the daily activity endpoint page by page. A page that
threw was only logged to the console: the loading banner disappeared, the
partial totals stayed on screen looking final, and Export Data stayed
clickable, so the CSV handed to finance was silently short.

The hook now reports `failed`, PaginationStatusAlerts renders it as an error
banner naming how many pages actually loaded, and the export is blocked with
the reason on hover while the data on screen does not cover the range.
2026-09-15 14:10:23 -07:00
yassin
67c17b68fa refactor(jwt-auth): extract issuer-scoped mapping lookup to satisfy C901 budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:09:01 +00:00
Yassin Kortam
e54b93017b fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions 2026-09-15 21:02:17 +00:00
mateo-berri
a9c422735f fix(router): stop counting caller-set timeout 408s toward deployment cooldown
A 408 produced by a timeout the caller set (a timeout body field or an
x-litellm-timeout header, which the proxy marks as client_side_timeout)
says nothing about the deployment's health, yet the router's primary
failure callback counted it toward allowed_fails and cooled the
deployment down. The fallback path already skipped it.

The marker never reached that callback because get_litellm_params drops
kwargs outside OPTIONAL_KWARGS_KEYS, so it is listed there now, and
deployment_callback_on_failure returns before the failure counter when
is_caller_timeout_408 holds. A 408 from a timeout the deployment or the
provider set still counts and still cools the deployment down.
2026-09-15 21:02:15 +00:00
yassin
0746cdbf2c refactor(terraform): drop explanatory comments from team_member_add
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:02:12 +00:00
yassin
5645e17b4f feat(terraform): add tpm_limit, rpm_limit, budget_duration, allowed_models to litellm_team_member_add
budget_duration and allowed_models ride on /team/member_add. tpm_limit and rpm_limit are sent through /team/member_update, the only endpoint that accepts them. Removing any of the four from config sends an explicit clear (null, or an empty list for allowed_models) since member_update is a merge-patch. The resource ID is set before the post-add limits call so a failure there taints the resource instead of orphaning the memberships

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:02:12 +00:00
Devin AI
d415c2856f style: ruff format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:02:10 +00:00
Devin AI
3cb5ceb98c fix(xai): honor nested web_search filters on the xAI Responses API
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:02:10 +00:00