* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* perf(lens): claim worker jobs from an indexed due queue instead of scanning every lens
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): apply the due_at index concurrently on its own and default legacy rows to due
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass repository to claim lifecycle tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): page past unsupported due lenses and declare the full due_at index
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(lens): claim due lenses in a loop instead of recursion
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* feat(lens): add dataset case and size limits
* feat(lens): add dataset, case and build models
* feat(lens): build dataset cases from traces, findings and text
* feat(lens): store dataset revisions insert-only
* feat(lens): add dataset routes for build, save, export and eval cases
* feat(lens): mount the dataset router before lens routes
* feat(lens): add LiteLLM_LensDataset table
* feat(lens): add LiteLLM_LensDataset table to proxy schema
* feat(lens): add LiteLLM_LensDataset table to extras schema
* feat(lens): add migration that creates the dataset table
* test(lens): cover dataset case building, dedupe and limits
* test(lens): cover dataset revisions, conflicts and eval cases
* chore(ui): regenerate API types for lens datasets
* feat(lens): add dataset UI types
* feat(lens): add datasets API client
* feat(lens): add dataset query and mutation hooks
* feat(lens): add case selection and expected edit logic
* test(lens): cover case selection and expected edits
* feat(lens): add the add to dataset dialog
* test(lens): cover saving picked cases from the dialog
* feat(lens): add datasets list
* feat(lens): add dataset detail with revisions and export
* test(lens): cover editing, revisions and export in datasets tab
* feat(lens): expose datasets on the lens API
* feat(lens): add in-memory datasets for demo mode
* feat(lens): wire demo datasets into the demo lens API
* feat(lens): add optional lens API hook
* feat(lens): add optional onboarding hook
* feat(lens): add datasets tab and dataset routing
* feat(lens): show the datasets tab
* feat(lens): add to dataset from the trace header
* feat(lens): add a single turn to a dataset from a step
* feat(lens): add finding evidence to a dataset
* docs(lens): add datasets screenshots for the PR
* test(lens): cover dataset revision storage against Postgres
* test(lens): cover dataset trace paging, findings and route errors
* test(lens): cover dataset build fallbacks and no_content skips
* fix(lens): register the dataset table for postgres span names
* feat(lens): add invalid skip reason for unparseable case lines
* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap
* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap
* chore(ui): regenerate API types for the invalid skip reason
* feat(lens): show text for the invalid skip reason
* feat(lens): render datasets in the same inspector table as investigations
* test(lens): open a dataset by clicking its table row
* docs(lens): update the datasets list screenshot
* style(lens): format the datasets table
* fix(lens): normalize JSON text span input and output into messages when building cases
* test(lens): cover JSON text span normalization for dataset cases
* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"
This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.
* fix(traces): normalize agent assistant summaries into UI messages
* feat(lens): add dataset case view helpers built on the trace parsers
* test(lens): cover dataset case view helpers
* feat(lens): show dataset cases in an inspector table
* feat(lens): open a dataset case in a side panel with trace message cards
* feat(lens): rebuild the dataset page header and layout
* feat(lens): keep the open dataset case in the URL
* test(lens): drive dataset edits through the case table and panel
* docs(lens): update dataset view screenshots
* Revert "test(lens): cover JSON text span normalization for dataset cases"
This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.
* fix(lens): import StateMessage from shared in dataset detail
* fix(lens): import StateMessage from shared in datasets list
* fix(lens): give the datasets migration a unique timestamp after review checkpoints
* feat(proxy): add models column to the end user table
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): enforce the end user models allowlist in model access checks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): accept and return models on the customer endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the customer models allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): resolve team aliases before the end user model check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(interactions): durable cross-pod settlement for background interaction billing
Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.
* fix(interactions): survive a settlement install failure at boot and stop carrying request headers
* fix(interactions): drop the stored request context once a settlement row is settled
* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline
* fix(interactions): carry a missing model through the settlement context for agent-only background creates
An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.
* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere
The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.
* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill
* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed
A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.
A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.
* fix(interactions): settle an unverified registration through the durable claim
A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.
* test(proxy): keep the settlement test where the proxy-infra shard collects it
The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.
* fix(interactions): raise on a non-2xx Gemini interaction fetch
AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises
* test(integration): audit durable background interaction settlement across replicas
Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart
* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate
* chore(ui): regenerate dashboard API types after merging main
* test(integration): accept the 422 budget refusal and a respawned worker's resume
The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged
* test(integration): delete the pinned key's interaction with a second key
A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The auto-router usage view summed the lifetime savings of every session
overlapping the date range, so it disagreed with the Overall savings view,
which sums daily rollups by request day.
Record auto-routed money per UTC request day and router in one new table,
written in the same statement as the session rollup and corrected in the
same transaction as late baseline estimates. The all-router headline reads
the same daily rows and filters as Overall; savings no router day row
accounts for are reported as unattributed and void the baseline comparison.
Session shape and caching stay whole-session and are labelled so; the
savings-per-session tile is removed.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* chore(lens): remove deployment screenshots
* refactor(lens)!: rename internal engine package and API
* fix(lens): pin worker image for renamed API
* test(lens): cover fresh and populated rename migrations
* fix(lens): protect db-push upgrades and restore routing and CI
* fix(lens): resolve migration tables across schemas and include database driver
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* feat(proxy): add model leaderboard analytics
* feat: add model insights task and range constants
* feat: record task type from task tags in model usage rollup
* feat: serve 365 days of model insights by UTC date
* test: cover task tag resolution in model usage rollup
* test: update model insights range limit test to 365 days
* chore: regenerate dashboard api types for model insights
* feat: add model insights aggregation helpers
* test: cover model insights aggregation helpers
* feat: redesign model leaderboard with stacked bars, treemap and ranking
* test: update model leaderboard view test
* feat: mark model leaderboard as beta in sidebar
* chore: sync schema.prisma copies from root
* fix: only treat task: prefixed tags as model insight tasks
* feat: add metric type for model insights ranking
* fix: rank model insights by selected metric and scope detail queries to ranked deployments
* test: plain tags are not model insight tasks
* test: cover metric ranking, deployment scoping and rollup round trip
* fix: build model insights weeks and halves from the requested date range
* test: cover empty weeks and range-based change comparison
* fix: refetch by metric, show load errors and ignore stale responses
* test: cover metric refetch and error state
* feat: define model insight tasks in a JSON file
* feat: return task labels and categories from model insights
* feat: load model insight tasks from JSON
* refactor: validate rollup task tags against the JSON task list
* feat: serve the task list with model insights
* refactor: drop hardcoded task list from constants
* build: ship model insight tasks JSON in the wheel
* test: cover model insight task JSON
* refactor: take task labels and categories from the API
* test: pass task info to task tile builder
* refactor: color treemap by API-provided category
* test: include tasks in model leaderboard fixture
* fix: make daily model usage migration idempotent
* feat: bound the model insights task query size
* fix: compute task breakdown independent of the chart metric
* test: task breakdown is stable across chart metrics
* chore: regenerate lazy openapi snapshot for model insights
* chore: regenerate dashboard api types for model insights
* fix: keep previous ranking dimmed while a new metric loads
* test: cover stale metric state in model leaderboard
* refactor: drop task row cap constant
* fix: return the full task breakdown instead of a truncated one
* test: task query is not truncated
* feat: add task summary types for model insights
* feat: summarise tasks server-side on a separate model insights endpoint
* test: cover the model insights tasks endpoint
* chore: regenerate lazy openapi snapshot for model insights tasks
* chore: regenerate dashboard api types for model insights tasks
* refactor: drop client-side task aggregation
* test: remove client-side task aggregation tests
* feat: load task breakdown separately from the chart metric
* test: task breakdown is not refetched on chart metric change
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Cherry-pick of merge commit b3882d8e43 (PRs #39321, #39562, #40107), which landed on litellm_internal_staging instead of main.
Co-authored-by: ojensen-berri <ojensen@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A policy attachment with default: true applies only when no non-default
attachment matches the request, so an opt-in guardrail policy replaces the
fallback one instead of running alongside it. Supported in config.yaml,
/policies/attachments, the Admin UI Attachments tab and the resolver
(matched_via is prefixed with default:).
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
vLLM serves no /v1/files or /v1/batches, so a hosted_vllm deployment can never
host a batch. Batch inputs for such a deployment now land in a LiteLLM-owned
storage backend, the batch is executed line by line through the deployment's
own chat, completion, embedding, or responses route, and the batch plus its
output and error files are served back from the database under the creating key
Adds a persistent total_spend column to LiteLLM_VerificationToken and LiteLLM_DeletedVerificationToken, incremented in the same write as spend and left alone by budget resets. Surfaces it on /key/info, /key/list and the Admin UI Virtual Keys table and key detail view
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds a daily spend table without api_key or user_id, written atomically alongside
LiteLLM_DailyUserSpend from the batched writer, reconciled from history by a
scheduled job that advances a marker in LiteLLM_Config, and read by the key-free
arm of the aggregated usage query once the marker covers the requested range.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Roll request_duration_ms of successful, non-internal requests into the
daily spend tables as total_response_time_ms plus timed_requests, expose
both through the daily activity endpoints, and derive the average in the
Usage -> Model Activity view of the Admin UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Adds a nullable tpd_limit column and field to keys, teams, budgets and end users. The batch submission limiter swaps the per-minute RPM/TPM descriptor of any scope that has a tpd_limit for a token-only 24h descriptor, so batch traffic is budgeted per day while online traffic keeps the existing per-minute limits. The Admin UI exposes the field on key, team and budget create/edit forms
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Keeps the base's rule that a non-admin id lookup matching no spend-log row answers 403, so the detail route never consults cold storage without an owner row
Adds object_permission.skills to keys and teams, enforces it on
/claude-code/marketplace.json?key=, /claude-code/plugins and
/claude-code/plugins/{name}, and exposes an Allowed Skills selector in
the key and team create/edit forms of the Admin UI
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>