Commit graph

5924 commits

Author SHA1 Message Date
yassin
09fa0833b4 feat(ui): surface top-key truncation on the team usage view
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:17:46 +00:00
yassin
38c1139377 feat(ui): note on the Key Activity tab when only the top-spend keys were loaded
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:13:05 +00:00
yassin
87894f2e7f fix(proxy): report total_api_keys so exact-limit key sets are not treated as truncated
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:02:11 +00:00
yassin
6cf35ed71b feat(proxy): expose lifetime total_spend on virtual keys
Adds a persistent total_spend column to LiteLLM_VerificationToken and LiteLLM_DeletedVerificationToken, incremented in the same write as spend and left alone by budget resets. Surfaces it on /key/info, /key/list and the Admin UI Virtual Keys table and key detail view

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:42:20 +00:00
yassin
61b3611b8c fix(ui): block the global usage export when the aggregated key cap is reached
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:22:20 +00:00
yassin
b5e2e9a392 Merge remote-tracking branch 'origin/main' into litellm_usage_key_free_aggregate_split 2026-09-16 08:58:03 +00:00
Yuneng Jiang
d67c7894dd
fix(ui): keep recorded order when an untimed guardrail shares a phase 2026-09-15 22:25:30 -07:00
Yuneng Jiang
8bb496154a
test(ui): scope lifecycle assertions with within instead of parentElement
The four .parentElement reads in the new lifecycle tests pushed
testing-library/no-node-access to 712 against a 707 budget, failing
frontend-lint. The rows now carry data-testid="lifecycle-row" and the test
picks a row with within(), which keeps the assertion tied to the specific row
rather than the whole panel and takes the count back to 707.
2026-09-15 22:18:42 -07:00
Yuneng Jiang
d20469b0dd
Merge remote-tracking branch 'origin/main' into litellm_fix_guardrail_lifecycle_untimed_entries 2026-09-15 21:44:22 -07:00
Yuneng Jiang
51a243e3ce
fix(ui): keep untimed guardrail entries on the request lifecycle
#39050 changed RequestLifecycle from sorting every entry with
(a.start_time ?? 0) to filtering on isTimed, which drops any entry whose
start_time/end_time are null. That was the right call for the not_run entries
the PR introduced, but it also drops entries that DID run and simply carry no
timing, and those are pre-existing: add_standard_logging_guardrail_information_to_request_data
defaults start_time, end_time and duration to None, and the conduct guardrail
passes none of them. One such entry used to draw the whole four-row lifecycle
and now draws nothing, so an admin opening that log sees an empty
Request Lifecycle panel.

An entry now stays on the lifecycle when it is timed OR when it ran, so not_run
keeps the exclusion #39050 wanted and every other shape comes back. Offsets are
number | null and render as an em dash rather than a fabricated T+0ms, which is
what a null minus a null used to produce on the base. Entries without timing
sort after the timed ones and the base time comes from the timed entries, so
real offsets are unchanged.

The two new tests fail on the base component and pass here; #39050's own
not_run tests keep passing untouched, which is what makes this additive rather
than a revert.
2026-09-15 21:44:19 -07:00
Tin Chi Lo
9d640b86ed fix(ui): simplify Capability and Fuse routing options 2026-09-15 20:58:38 -07:00
tin-berri
3b14631f06
Merge pull request #41315 from BerriAI/litellm_capability_fuse_classifier_ui
feat(ui): configure capability and Fuse v2 classifiers
2026-09-15 20:13:57 -07:00
Tin Chi Lo
39b7812810 feat(ui): configure capability and Fuse v2 classifiers 2026-09-15 19:30:50 -07:00
yucheng-berri
41eb2dbfeb
Merge pull request #41128 from BerriAI/litellm_llm_judge_pre_call
feat(guardrails): support pre_call and during_call modes for llm_as_a_judge
2026-09-15 18:58:21 -07:00
yassin
3d9a30500c Merge remote-tracking branch 'origin/main' into litellm_usage_key_free_aggregate_split
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/management_endpoints/common_daily_activity.py
2026-09-16 01:48:39 +00:00
Yassin Kortam
67cb0bc089
Merge pull request #41313 from BerriAI/litellm_model_activity_response_time
feat(ui): show average response time per model in usage model activity
2026-09-15 18:43:48 -07:00
ryan
9f990c4f86 fix(router): validate routing_groups at save time and keep invalid DB groups from blocking SSO load
Overlapping routing_groups persisted from the Admin UI raised inside
Router._init_routing_groups during the DB config reconcile, which skipped
loading SSO, guardrails and the other DB-backed settings while leaving the
proxy healthy. /config/update now returns 400 for overlapping models,
duplicate names, the reserved default name and unknown strategies before
writing, the Router builds every group selector before replacing its state
so a rejected update keeps the previous groups routing, and the proxy applies
routing_groups separately from the other router settings so an already
persisted invalid value is logged and skipped instead of aborting the
reconcile. The Admin UI modal blocks picking a model another group owns.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:25:55 +00:00
yassin
d8ef940232 chore(ui): regenerate schema.d.ts for organization_id on archived key records
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:06:36 +00:00
yucheng
77a6327675 feat(guardrails): support pre_call and during_call modes for llm_as_a_judge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
ryan
b3432abef7 refactor(ui): derive the ssh skill name from the bare repo path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:58:03 +00:00
ryan
b0ed6c8d89 merge main into litellm_skills_ssh_sources and re-apply ssh source parsing on the shadcn skill form
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:57:30 +00:00
yucheng
9bd3f7b885 Merge remote-tracking branch 'origin/main' into litellm_agent365_mcp_guardrail 2026-09-16 00:52:27 +00:00
yassin
260ff5f491 feat(team): team-level model_max_budget with key-level overrides
A team can now carry a per-model budget map that every key on the team
inherits. A key's own model_max_budget entry for the same model takes
precedence, so it is gated on and billed to the key alone.

Backend: NewTeamRequest/UpdateTeamRequest accept model_max_budget (validated
like the key-level field, enterprise gated); the value is hydrated onto
UserAPIKeyAuth via the token view, TeamGrants and the carried budget state;
_check_team_model_budget enforces it in the centralized common checks; the
limiter meters spend under team_model_spend:<team>:<model>:<duration> and
skips the team counter when the key overrides; /team/update lets only a
proxy admin raise, re-window or drop a cap; /team/info exposes usage.
The Anthropic context-management compaction summary subrequest runs the
same team gate. Both fallback token-view SQL definitions project the column.

UI: team create and edit forms reuse the key-level ModelMaxBudgetEditor,
premium gated, sending {} to clear and omitting unchanged fields.

A key entry overrides the team cap only when it spend-gates the model
(non-negative max_budget); a row that only carries tpm/rpm limits or a
negative cap leaves the team cap in force.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:40:58 +00:00
Mateo Wang
aa710dc6a9
Merge pull request #35388 from BerriAI/litellm_session_total_duration
fix(spend): sum multi-round session duration in logs UI
2026-09-15 17:31:41 -07:00
Yassin Kortam
7eeba69016
Merge pull request #41316 from BerriAI/litellm_nvidia_nim_infer_passthrough
feat(proxy): add /nvidia_nim passthrough route for NIM object detection and OCR /v1/infer
2026-09-15 17:28:25 -07:00
ryan-crabbe-berri
f34a6eda92
Merge pull request #41294 from BerriAI/litellm_usage_export_gating
fix(ui): block usage export and flag the range when a spend page fails
2026-09-15 17:19:38 -07:00
yassin
251aeea97d fix(ui): show the S3 label when editing the s3_v2 callback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:56:31 +00:00
yucheng
aa7f1e16b8 feat(proxy): throttle failed Admin UI sign-ins per source and source/username
Replace the username-global lockout with counters keyed by source address and by
source/username pair. Each has a fixed counting window (60s) and a separate
block TTL (300s). Blocks are soft: a correct password still signs in, wrong
passwords from a blocked key take one of 5 held slots per worker and are held
30s before a 429. Once a pair is blocked its failures stop counting against the
source. The source scope runs only when trusted_proxy_ranges is set, IPv6 is
grouped by /64, and per-source limits accept IP and CIDR overrides with
longest-prefix matching. Redis is authoritative through one Lua script per
failure, with bounded per-worker fallback when Redis raises.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:53:14 +00:00
yucheng
7c24c8bc1d Merge remote-tracking branch 'origin/main' into litellm_lit5285_login_rate_limit_v2 2026-09-15 23:31:07 +00:00
mateo-berri
790c9ab77d test(ui): name the session column fixtures so the inline-object lint budget holds
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
2026-09-15 16:25:01 -07:00
ryan-crabbe-berri
69c3212220 fix(ui): gate the usage export on range coverage, not on a fetch being in flight
A loading flag only flips once the fetch effect runs, so the render right after a
date or filter change still reported the previous range as loaded and let an export
read its rows. Stamp the completed range on the hook and compare it during render
instead, the way the tiles already do.

Also stop the failure banner claiming a page loaded when the first request is what
failed, which left it reading 1/1.
2026-09-15 16:21:37 -07:00
mateo-berri
3441f73118 Merge remote-tracking branch 'origin/main' into litellm_session_total_duration
# Conflicts:
#	litellm/proxy/spend_tracking/spend_management_endpoints.py
#	tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
#	ui/litellm-dashboard/src/components/view_logs/RequestLogsTableColumns.tsx
#	ui/litellm-dashboard/src/components/view_logs/columns.tsx
2026-09-15 16:17:18 -07:00
Louis Vauterin
3f7a344337 feat(jwt-key-mapping): accept token_id as an alternative to the plaintext key
A JWT key mapping can now name its virtual key by the SHA-256 hash the proxy
already stores, instead of only by the plaintext key.

litellm_key makes its generated key write-only so raw keys stay out of Terraform
state, and write-only attributes cannot be referenced at all, so the natural
wiring fails while planning, in every apply ordering:

  Error: Missing required argument
    with litellm_jwt_key_mapping.example
    key = litellm_key.example.key
    The argument "key" is required, but no definition was found.

The only way out today is supplying the plaintext from a variable or a secret
manager, which means the mapped key cannot be one the proxy generated and the
configuration has to carry a credential. The value the mapping stores is
hash_token(key), which is the same hash litellm_key already exports as
token_id, and a hash is not a credential, so accepting it closes the gap:

  resource "litellm_jwt_key_mapping" "service" {
    jwt_claim_name  = "client_id"
    jwt_claim_value = "reporting-service"
    token_id        = litellm_key.service.token_id
  }

CreateJWTKeyMappingRequest and UpdateJWTKeyMappingRequest gain an optional
token. Create requires exactly one of key or token, update accepts at most one,
and omitting both still leaves the mapped key alone. A supplied token must be 64
lowercase hex characters, because hash_token() hashes unconditionally and a
plaintext key sent as token would be stored as a hash of a hash, then silently
match nothing at auth time. Both rejections are 400s raised before the row is
written.

On the provider side, key becomes Optional with ExactlyOneOf{key, token_id} and
token_id is added next to it. token_id is not marked sensitive since a hash is
not a credential, both fields are omitempty on the wire so the proxy receives
only the one that was configured, and a failed update reverts token_id for the
same reason it already reverts key.

key keeps working unchanged and existing state is untouched. The only change to
it is Required to Optional, which no existing configuration can violate.
2026-09-15 23:15:20 +00:00
yassin
f6f782ff67 feat(s3): add s3_log_prompts_only option to log prompts without responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:11:44 +00:00
tin-berri
8cd00d2d6e
Merge pull request #41282 from BerriAI/litellm_fast_mode_toggle_0915
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
ai-gateway image / ai-gateway release image (push) Has been cancelled
feat(auto-router): add per-model Fast mode toggle
2026-09-15 16:10:06 -07:00
Yassin Kortam
0bb95e298c
Merge pull request #41309 from BerriAI/litellm_playground_custom_request_headers 2026-09-15 15:56:22 -07:00
Yassin Kortam
9d75cdd502
Merge pull request #41304 from BerriAI/litellm_lit1698_openai_system_messages_first 2026-09-15 15:56:13 -07:00
Yassin Kortam
9bb83fdaea
Merge pull request #41281 from BerriAI/litellm_lit_7417_jwt_key_mapping_issuer_scope
fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions
2026-09-15 15:51:50 -07:00
Yassin Kortam
eec954bf9a
Merge pull request #41296 from BerriAI/litellm_models_table_url_state
feat(ui): persist Models table search, filters, sort and page in the URL
2026-09-15 15:46:27 -07:00
yassin
be082013f7 fix(ui): label the response time chart tooltip with the series name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:38:55 +00:00
yassin
ebb9a3ceeb chore(ui): regenerate schema.d.ts for the /nvidia_nim passthrough route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:34:59 +00:00
yassin
754e87fa67 fix(ui): keep gateway auth and content-type headers ahead of playground custom headers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:24:13 +00:00
yassin
5b0fa89056 feat(ui): show average response time per model in usage model activity
Roll request_duration_ms of successful, non-internal requests into the
daily spend tables as total_response_time_ms plus timed_requests, expose
both through the daily activity endpoints, and derive the average in the
Usage -> Model Activity view of the Admin UI

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:16:12 +00:00
yassin
c05af7a12f feat(ui): add custom request headers to the API Playground
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:12:06 +00:00
yassin
6c8b9a7b05 feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info
Persist and expose the lifecycle of API keys so spend, audit and FinOps
workflows can still resolve a key after it is revoked, expires or is
deleted.

/key/list?status= now accepts active, expired and revoked next to the
existing deleted value. revoked means blocked=true, expired means not
blocked with a past expiry, active is the rest, so the three values
partition the live key table. deleted keeps reading the
LiteLLM_DeletedVerificationToken archive.

/key/info falls back to that archive when the key is no longer in the
live table, running the same owner/team/org authorization check, and
every response now carries a derived status field. The hashed token is
still stripped.

The Virtual Keys page gets a Status filter (URL-persisted) and a Deleted
badge that shows when and by whom the key was deleted.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:12:01 +00:00
yassin
99fa38504a fix(ui): bound Models table page, page size and sort_by read from the URL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:55:07 +00:00
yassin
77d913958d feat(openai): add openai_system_messages_first to put system messages first for prompt caching
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:53:29 +00:00
tin-berri
ce1a4f896a
Merge pull request #41272 from BerriAI/litellm_fuse_v2_classifier_pr
feat(router): add Fuse V2 classifier after capability forecasting
2026-09-15 14:43:44 -07:00
yassin
868d3855ab feat(ui): persist Models table search, filters, sort and page in the URL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:37:09 +00:00
yassin
92e55b3b22 perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 21:17:00 +00:00