* feat(proxy): add model leaderboard analytics
* feat: add model insights task and range constants
* feat: record task type from task tags in model usage rollup
* feat: serve 365 days of model insights by UTC date
* test: cover task tag resolution in model usage rollup
* test: update model insights range limit test to 365 days
* chore: regenerate dashboard api types for model insights
* feat: add model insights aggregation helpers
* test: cover model insights aggregation helpers
* feat: redesign model leaderboard with stacked bars, treemap and ranking
* test: update model leaderboard view test
* feat: mark model leaderboard as beta in sidebar
* chore: sync schema.prisma copies from root
* fix: only treat task: prefixed tags as model insight tasks
* feat: add metric type for model insights ranking
* fix: rank model insights by selected metric and scope detail queries to ranked deployments
* test: plain tags are not model insight tasks
* test: cover metric ranking, deployment scoping and rollup round trip
* fix: build model insights weeks and halves from the requested date range
* test: cover empty weeks and range-based change comparison
* fix: refetch by metric, show load errors and ignore stale responses
* test: cover metric refetch and error state
* feat: define model insight tasks in a JSON file
* feat: return task labels and categories from model insights
* feat: load model insight tasks from JSON
* refactor: validate rollup task tags against the JSON task list
* feat: serve the task list with model insights
* refactor: drop hardcoded task list from constants
* build: ship model insight tasks JSON in the wheel
* test: cover model insight task JSON
* refactor: take task labels and categories from the API
* test: pass task info to task tile builder
* refactor: color treemap by API-provided category
* test: include tasks in model leaderboard fixture
* fix: make daily model usage migration idempotent
* feat: bound the model insights task query size
* fix: compute task breakdown independent of the chart metric
* test: task breakdown is stable across chart metrics
* chore: regenerate lazy openapi snapshot for model insights
* chore: regenerate dashboard api types for model insights
* fix: keep previous ranking dimmed while a new metric loads
* test: cover stale metric state in model leaderboard
* refactor: drop task row cap constant
* fix: return the full task breakdown instead of a truncated one
* test: task query is not truncated
* feat: add task summary types for model insights
* feat: summarise tasks server-side on a separate model insights endpoint
* test: cover the model insights tasks endpoint
* chore: regenerate lazy openapi snapshot for model insights tasks
* chore: regenerate dashboard api types for model insights tasks
* refactor: drop client-side task aggregation
* test: remove client-side task aggregation tests
* feat: load task breakdown separately from the chart metric
* test: task breakdown is not refetched on chart metric change
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Cherry-pick of merge commit b3882d8e43 (PRs #39321, #39562, #40107), which landed on litellm_internal_staging instead of main.
Co-authored-by: ojensen-berri <ojensen@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A policy attachment with default: true applies only when no non-default
attachment matches the request, so an opt-in guardrail policy replaces the
fallback one instead of running alongside it. Supported in config.yaml,
/policies/attachments, the Admin UI Attachments tab and the resolver
(matched_via is prefixed with default:).
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
vLLM serves no /v1/files or /v1/batches, so a hosted_vllm deployment can never
host a batch. Batch inputs for such a deployment now land in a LiteLLM-owned
storage backend, the batch is executed line by line through the deployment's
own chat, completion, embedding, or responses route, and the batch plus its
output and error files are served back from the database under the creating key
Adds a persistent total_spend column to LiteLLM_VerificationToken and LiteLLM_DeletedVerificationToken, incremented in the same write as spend and left alone by budget resets. Surfaces it on /key/info, /key/list and the Admin UI Virtual Keys table and key detail view
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds a daily spend table without api_key or user_id, written atomically alongside
LiteLLM_DailyUserSpend from the batched writer, reconciled from history by a
scheduled job that advances a marker in LiteLLM_Config, and read by the key-free
arm of the aggregated usage query once the marker covers the requested range.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Roll request_duration_ms of successful, non-internal requests into the
daily spend tables as total_response_time_ms plus timed_requests, expose
both through the daily activity endpoints, and derive the average in the
Usage -> Model Activity view of the Admin UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
LITELLM_LOG=ERROR still printed INFO lines from uvicorn (startup and access log) and from the litellm_proxy_extras migration logger, because neither read the variable. Forward the resolved level to uvicorn when LITELLM_LOG is set and no explicit log_config or JSON logging is in use, and let the extras logger take its level from LITELLM_LOG
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds a nullable tpd_limit column and field to keys, teams, budgets and end users. The batch submission limiter swaps the per-minute RPM/TPM descriptor of any scope that has a tpd_limit for a token-only 24h descriptor, so batch traffic is budgeted per day while online traffic keeps the existing per-minute limits. The Admin UI exposes the field on key, team and budget create/edit forms
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Keeps the base's rule that a non-admin id lookup matching no spend-log row answers 403, so the detail route never consults cold storage without an owner row