Commit graph

6100 commits

Author SHA1 Message Date
Ishaan Jaff
bea84c323c
perf(lens): tick fast only while reasoning is typing 2026-10-04 18:10:57 -07:00
Ishaan Jaff
f0c2baf254
fix(lens): keep polling a finished run until its last reviews arrive 2026-10-04 18:10:57 -07:00
Ishaan Jaff
e50f53965d
refactor(lens): remove dead live helpers and use generated in-flight types 2026-10-04 18:10:57 -07:00
Ishaan Jaff
d58c7a2187
chore(lens): drop the unused review fixture 2026-10-04 18:10:57 -07:00
Ishaan Jaff
68ca72041e
style(lens): format live run files 2026-10-03 18:19:40 -07:00
Ishaan Jaff
8b9e799a19
refactor(lens): name inline objects in the live run 2026-10-03 18:19:40 -07:00
Ishaan Jaff
775a808003
refactor(lens): name now reading conditions 2026-10-03 18:19:40 -07:00
Ishaan Jaff
2c05090b1a
refactor(lens): name the run now handler in investigations view 2026-10-03 18:19:40 -07:00
Ishaan Jaff
ca9db83e17
Merge remote-tracking branch 'upstream/main' into litellm_lens_live_review
# Conflicts:
#	litellm/proxy/lens/worker.py
#	tests/unit/proxy/lens/test_worker.py
#	ui/litellm-dashboard/src/components/lens/investigations/InvestigationsView.tsx
#	ui/litellm-dashboard/src/components/lens/investigations/detail/InvestigationDetail.tsx
#	ui/litellm-dashboard/src/lib/http/schema.d.ts
2026-10-03 18:15:36 -07:00
Ishaan Jaff
4db3c3df22
fix(lens): give demo jobs an empty in-flight list 2026-10-03 18:09:50 -07:00
Ishaan Jaff
3c8309272f
fix(lens): resolve the analysis provider logo from the model catalog 2026-10-03 18:09:50 -07:00
Ishaan Jaff
6fb7e8124f
feat(lens): put the now reading stage above the trace list in View run 2026-10-03 18:09:50 -07:00
Ishaan Jaff
3577f3bebf
feat(lens): show a now reading stage that types each trace's reasoning 2026-10-03 18:09:47 -07:00
Ishaan Jaff
69b5265a17
feat(lens): model live reading lanes from in-flight runs and reviews 2026-10-03 18:09:47 -07:00
devin-ai-integration[bot]
62fb808d4b
feat: improve Lens dev seeding and live UI (#44468)
* feat: seed Lens dev with configurable load profiles

* feat: run Lens UI live through the dev launcher

* fix: verify Lens UI startup before seeding

* fix(ui): render run timestamps on one compact line

The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): use the root query client for the Lens demo

Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: drop unused synthetic_spend reference from query help

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(traces): drop the deeplite fixtures and map every SDK to a logo

The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata

The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): inject the Lens sample through data sources instead of demo checks

Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): infinite scroll for the agent runs list

Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): keep the whole Lens view in the URL

Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs

RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(traces): cargo fmt captures tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 01:03:11 +00:00
Ishaan Jaff
96e30ad045
chore(ui): regenerate api types for lens in-flight runs 2026-10-03 18:02:02 -07:00
Ishaan Jaff
28f70ea41d
feat(lens): slide one model rectangle over the traces being read
A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.
2026-10-03 17:58:18 -07:00
Ishaan Jaff
5724190869
feat(lens): sum up a finished live run with time taken
doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.
2026-10-03 17:58:16 -07:00
Ishaan Jaff
b174232308
feat(lens): show what the worker is reading and make View run obvious
Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.
2026-10-03 17:55:12 -07:00
Ishaan Jaff
0376a8d7c6
refactor(lens): drop client-side replay in favour of real in-flight rows
Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.
2026-10-03 17:55:10 -07:00
Ishaan Jaff
b00cb8d733
feat(lens): simplify View run to traces and conclusions
The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.
2026-10-03 17:47:02 -07:00
Ishaan Jaff
55d9425a82
fix(lens): group live conclusions by check with short labels
There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.
2026-10-03 17:46:52 -07:00
Ishaan Jaff
10260da169
fix(lens): split live conclusions into issues and patterns
A check could show up twice with the same label, once as an issue and once as a pattern.
2026-10-03 17:41:20 -07:00
Ishaan Jaff
7ef76eabff
fix(lens): feed the live run from the reviews endpoint and keep View run open
LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.
2026-10-03 17:40:35 -07:00
Ishaan Jaff
2f5fb3ef44
feat(ui): add a lens reviews query that polls the index cursor while live 2026-10-03 17:40:32 -07:00
Ishaan Jaff
db0e4df4f5
feat(lens): page job reviews by index cursor
Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.
2026-10-03 17:40:30 -07:00
Ishaan Jaff
73ad24f7de
chore(ui): regenerate api types for the lens reviews endpoint 2026-10-03 17:39:38 -07:00
Ishaan Jaff
a75dab2f41
feat(lens): show the queue reason and what the worker is doing in the live strip
The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.
2026-10-03 17:36:52 -07:00
Ishaan Jaff
12738b0d79
feat(lens): explain why a queued investigation is waiting
Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.
2026-10-03 17:36:49 -07:00
Ishaan Jaff
c4fcd4b457
feat(lens): show the live run as a two-pane trace and conclusions view
Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.
2026-10-03 17:31:34 -07:00
Ishaan Jaff
1dfcdf223e
feat(lens): keep the live run ambient until View run is clicked
The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.
2026-10-03 17:31:32 -07:00
Ishaan Jaff
7494885093
feat(lens): pace live playback so each trace stays readable
Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.
2026-10-03 17:28:28 -07:00
tin-berri
6370104c53
feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444)
* feat(enterprise): bundle LiteAdmin Slack worker with native gateway login

* fix(enterprise): preserve gateway prefixes during native Slack linking

* fix(enterprise): retain native Slack linking on the admin backend

* fix(enterprise): reuse shared native Slack connection services
2026-10-03 17:28:18 -07:00
moe-berri
4b67a2b845
feat(roi): default people and branch lists to matched accounts (#44465)
* feat(roi): show matched people by default in contributor lists

* fix(roi): keep matched filter tabs readable on narrow screens

* fix(roi): retain spend-only users and support older browsers
2026-10-03 17:24:18 -07:00
Ishaan Jaff
2cffbdbc95
fix(lens): list recorded agents in the run now dialog
The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.
2026-10-03 17:24:04 -07:00
Ishaan Jaff
7de044b71e
feat(lens): read review spans as a conversation timeline
Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.
2026-10-03 17:22:13 -07:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
f0eda6d2a6
fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API (#44419)
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API

Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded

Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode

* fix(health): resolve the test connection mode from the deployment when the request omits it

The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.

* fix(health): resolve an omitted ahealth_check mode the way the proxy does

* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode

A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.

* fix(health): shape test connection probe params for the model the request probes

A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden

* fix(health): report an early ahealth_check failure as itself, not as a missing mode

With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:05:47 -07:00
Ishaan Jaff
68cc5484d1
feat(lens): add an ambient live strip that opens the drawer 2026-10-03 17:04:54 -07:00
Ishaan Jaff
8e473ade78
feat(lens): put the live strip under the progress bar and drop the inline panel 2026-10-03 17:04:45 -07:00
Ishaan Jaff
09ac1663b8
feat(lens): add a live trace results drawer with readable spans 2026-10-03 17:04:40 -07:00
Ishaan Jaff
b4ef33b7a1
feat(lens): track active jobs before their first review 2026-10-03 17:04:37 -07:00
Ishaan Jaff
d16b0d7132
feat(lens): derive strip status, honest issue counts and drawer focus from a job 2026-10-03 17:00:31 -07:00
Ishaan Jaff
af1c2f1a95
feat(lens): format review span previews as readable messages 2026-10-03 17:00:29 -07:00
Ishaan Jaff
616c59c8d0
fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size 2026-10-03 16:49:55 -07:00
Ishaan Jaff
1fb70ff6be
feat(lens): collapse the live run to an ambient line with show work 2026-10-03 16:49:53 -07:00
Ishaan Jaff
d2e96c41f2
feat(lens): add a reading ticker line and replay for finished runs 2026-10-03 16:49:50 -07:00
Ishaan Jaff
ac4d11abbc
refactor(lens): restyle the live run as the native progress panel 2026-10-03 16:44:49 -07:00
Ishaan Jaff
25c417d8ba
fix(lens): show the live run only for real reviews and keep fixtures test-only 2026-10-03 16:44:45 -07:00
Ishaan Jaff
cac908783a
feat(lens): stream large review backlogs at 150ms or less and list newest first 2026-10-03 16:42:51 -07:00