Commit graph

53626 commits

Author SHA1 Message Date
Yujong Lee
a0a275e486 fix(ui): share one centered loading state across Lens traces and investigations 2026-10-04 16:11:21 -07:00
Yujong Lee
b99c6d40d2 feat(ui): search and plot agent runs on the server
The runs list sends its search and zoom window to /v1/traces, and the
timeline plots /v1/traces/histogram, so every matching run is counted and
listed instead of only the pages already loaded. The demo data answers the
same calls locally

An empty result from a search or a zoom keeps the runs controls. Before,
the empty page looked like a proxy that never received a trace and
replaced the list with onboarding
2026-10-04 15:52:23 -07:00
Yujong Lee
180b5ee482 refactor(traces): read traces through a five-method storage port
Trace storage now answers five storage-neutral reads: runs, run_counts,
spans, span_text and calls. Query and row types live in
litellm_traces::store with no ClickHouse encodings, and traces-clickhouse
only adapts them, so another backend implements five queries instead of
one SQL file per UI widget

The runs list, histogram, field values, trace graph, span detail, error
paging and cost lookup all go through the port. Paging, the graph budget,
excerpting and the agent output fallback moved out of SQL into
TraceReader. QueryScope is the only access wire, and every SQL file reads
owned_spans, owned_runs and owned_calls CTEs built from one ownership rule
per table that the console row policy also renders

Cursors are tagged with their kind, and run search terms reach storage as
raw text and globs. This also carries the run histogram and values
endpoints, the gateway-traces router and the agent label rollup
2026-10-04 15:39:45 -07:00
Yujong Lee
60e7726116 fix(ui): always show the Lens Settings tab
The tab only rendered once GET /lens succeeded for a configuring session, so a
slow or failing list, or a read-only session, silently hid it. Settings now
always renders, and the worker section shows only where it can be used

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 13:54:19 -07:00
devin-ai-integration[bot]
461a58c40a
refactor(lens): storage-independent trace reads, shared keyset pager, typed read failures (#44422)
* feat(lens): own trace reads behind a cached TraceStore port

Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.

Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.

Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(lens): seed large profiles with server-side copies and long sessions

Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* style(lens): float the investigation setup badge on the tab edge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens): restructure the trace drawer and polish its layout

Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.

Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* typesafety

* wip

* fmt

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:39:47 -07:00
devin-ai-integration[bot]
e1d16f51d1
refactor(ui): share CopyButton between Lens traces and logs (#44513)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:49 -07:00
devin-ai-integration[bot]
1459e00430
refactor(ui): rename view_logs to logs and split request, audit and detail (#44505)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:49 -07:00
devin-ai-integration[bot]
05f1c73a3c
refactor(ui): move TraceView into components/lens/traces (#44501)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:48 -07:00
devin-ai-integration[bot]
6532dcb73b
test(ci): pin the ROI estimator flag and the prompt-cache counter in two drifted tests (#44499)
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 08:40:23 -07:00
devin-ai-integration[bot]
a2bf67a037
test(mcp): keep the SSO assertion round trip from matching its refresh token inside random ciphertext (#44494)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 03:18:54 -07:00
devin-ai-integration[bot]
b4fcc5c1bc
fix(otel): read the registered v2 logger without importing the proxy (#44485)
* fix(otel): read the registered v2 logger without importing the proxy

phase_span, which the router enters on every deployment pick since #44150, looked up the
proxy's OTel logger by importing litellm.proxy.proxy_server. In an SDK process that import
loads the whole proxy synchronously on the caller's event loop during its first request,
which stalled litellm_router_unit_testing past its 5s wait. Read the module from
sys.modules instead: when the proxy was never imported it has no registered logger.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): pin the no-proxy-import guarantee in-process

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 03:17:07 -07:00
devin-ai-integration[bot]
bb4f7211d7
refactor: clean up fresh tech debt from 2026-10-03 (#44484)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 01:31:23 -07:00
devin-ai-integration[bot]
0b74ae9c5c
feat(mcp): hand listed-tool description and input schema to pre-call hooks per caller (#41162)
* feat(mcp): hand listed-tool metadata to pre-call hooks with per-caller catalog identity

Track the tools each MCP server listed per caller identity so pre_mcp_call and during_mcp_call hooks receive the tool description and input schema the client saw. Servers with no caller-dependent inputs share one slot; user identity, forwarded headers, stdio env, relayed bearers, and server-specific auth get their own. Local registry and OpenAPI paths pass the registered metadata and admin description overrides. The Agent 365 guardrail reads the new fields into its evaluate payload.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): drop the listed-tools empty sentinel and routine test docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): mark the listed-tools cache digest as a non-security hash

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): key the listed-tools cache by the OBO subject token

token_exchange servers list upstream with the caller's own Entra bearer, so two callers on one
LiteLLM key with different subjects were sharing a catalog slot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): resolve the BYOK credential before keying the listed-tools slot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): drop the OAuth discovery cache when a server definition changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): drop a diff-narrating comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): never validate a supplied header on the tools/list BYOK path

The pre-listing resolver ran the tool-call byok_auth_required check even
when the caller already supplied x-mcp-auth, and it ran outside the
per-server error boundary, so a single deprecated-header caller dropped
the server from the aggregate list. Listing now returns a supplied header
unchanged and falls back to the stored credential without raising

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): assert the BYOK listing lands in the caller's listed-tool slot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover the deprecated string x-mcp-auth header on a BYOK tools/list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): key the per-caller listed-tool slot by the hashed token

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): key discovery cache by the hashed token

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates

Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): derive the listed-tool caller identity from the discovery cache key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys

An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.

The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.

Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep OAuth metadata generations only while a fetch is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): return one masked text per scanned string in the selected-guardrail REST test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): fold the signed caller into the discovery digest instead of a second key hash

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): satisfy type discipline gate on listed-tool identity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): hand tools/call hooks the exact catalog entry tools/list served

get_listed_tool re-applied the admin description override on top of the cached listing, so a
guardrail-masked description was restored to its original wording at call time, and the OpenAPI /
local-registry call path built its metadata from the registry instead of the guarded caller catalog.
Both paths now return the cached entry as served, falling back to the registry only when no listing
was recorded

Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two
regressions, red on the merge base for the feature, green on this head)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): align listed-tool slot tests with per-caller keying

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret

The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that
credential to the upstream client, which on an oauth2 server short-circuited the client_credentials
mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so
tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request
header, letting the M2M mint or signed JWT proceed as on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy

Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential,
one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the
pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list

Listing used the resolved stored credential both to key the caller's catalog slot and as the
upstream transport header, so REST api_key and bearer_token listings sent the user's secret
instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client
and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly
as before the catalog existed, and the stored credential only names the slot tools/call reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): type the listed-tool metadata read from pre-call kwargs for the basedpyright gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): hand never-listed tools/call hooks name and arguments only

The local-registry call path fell back to the registry entry with the admin description override
when no tools/list had been recorded for the caller, so a pre_mcp_call guardrail scanned a
description the caller was never served and blocked OpenAPI calls that passed before, and base's
own selected-guardrail REST test failed on the two-text redaction. _registered_tool_metadata now
returns the listed entry or None, so a tools/call with no prior listing sends name and arguments
only as promised, and that REST test double goes back to its base shape

* fix(mcp): keep during_mcp_call hooks on name and arguments only

call_tool handed the caller's listed entry to the during-hook task as well, so during_mcp_call
guardrails scanned the description line and schema leaves of any listed tool after the upstream
call had already run, blocking calls that passed before whenever the policy matched the
description, returned a fixed-length texts list, or hit the depth guard on a deep schema. The
listed entry is only disclosed for pre_mcp_call, so the during task no longer receives it and its
request object carries no description or schema, as before

* fix(mcp): key the BYOK catalog slot by the client's header, not the stored credential

tools/list resolved the stored BYOK credential to pick the caller's catalog slot, which read the
credential store before the classified try block. With Postgres down and a cold per-worker cache
that made every REST tools/list on an is_byok server fail with tools=[] and no upstream call, and
the read seeded the per-worker cache (including a negative entry), so a tools/call on another
worker after a store, rotate or revoke on this one kept using the stale value.

The slot is now keyed by what the client supplied plus the caller's hashed key, on both sides.
_get_tools_from_server and call_tool take a keyword-only catalog_auth_header that defaults to
mcp_auth_header as received (the default is the builtin Ellipsis so it survives a module reload).
The /mcp fan-out and execute_mcp_tool, which swap the resolved credential into mcp_auth_header,
pass the client's value explicitly. What goes upstream is unchanged. _byok_catalog_auth_header is
gone.

* fix(mcp): drop a listed catalog recorded across a server save

_record_listed_tools ran after the awaited upstream fetch, so a PUT /v1/mcp/server that landed
mid-fetch had its invalidation undone when the fetch completed: hooks then saw the pre-save
description next to the post-save definition until the next listing, instead of name and
arguments only.

The manager now keeps a per-server listed-tools generation, bumped by
_invalidate_server_definition_caches. _get_tools_from_server reads it before the fetch and
_record_listed_tools skips the write when it moved; the next listing records normally.

* fix(mcp): drop the catalog again once a saved OpenAPI server's registry is rebuilt

add_server and update_server publish the saved definition before the OpenAPI registry entries are
rebuilt from the spec, so a listing recorded during that fetch held the pre-save entries under the
new generation. The generation is bumped a second time after the registry refresh.

The during-hook task no longer accepts a listed entry, the one-line wrapper over get_listed_tool is
inlined at its two call sites, and the per-server generation map is a plain dict.

* fix(mcp): keep discovery and OAuth metadata caches across an OpenAPI spec re-read

add_server and update_server ran the full server-definition invalidation a
second time after the awaited OpenAPI spec fetch, which also dropped the
prompts/resources/templates discovery entries and the OAuth protected-resource
metadata filled under the already-published definition, so the next request
went upstream again. Only the listed-tool catalog recorded during the fetch
holds pre-save entries, so the post-fetch pass now drops just that catalog
and bumps its generation via the new _drop_listed_tools helper, which the
full invalidation also calls.

* fix(mcp): look a called tool up in the listed catalog by its bare name only

get_listed_tool stripped the server prefix a second time when the exact
name was absent from the caller's listing, so a never-listed upstream tool
whose bare name starts with the server prefix resolved to the listed
sibling and that sibling's description and input schema reached the
pre-call hooks for a call to a different tool. Every caller already passes
the once-stripped bare name, so the lookup is now exact.

Tests that looked the catalog up by a prefixed name now use the bare name
the callers pass; two new tests pin the never-listed sibling case at the
manager and at the tools/call path.

* fix(mcp): record a listed-tool catalog only for a listing the caller is served

_get_tools_from_server now records the catalog into the caller's
listed-tools slot only when asked (record_listing=True), which the
served listings pass: the /mcp and Responses API tools/list handlers via
_get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST
listing via _list_server_tools. Four internal listings stop recording,
so a later tools/call hands pre_mcp_call hooks name and arguments only,
as on main:

- _list_tools_before_first_call, the implicit listing inside tools/call
  when this worker does not yet expose the tool
- fetch_pinnable_tool_catalog, the admin pin snapshot listed without the
  catalog guard and without description overrides
- _initialize_tool_name_to_mcp_server_name_mapping, the startup fill
- get_tools_for_server, used by the semantic tool filter

_create_prefixed_tools returns to its tool-name mapping job only; the
record follows it in _get_tools_from_server.

* fix(mcp): opt every listing out of catalog recording unless it is served

The aggregate listing and _list_mcp_tools now default to record_listing=False,
so a catalog fetched inside a tools/call no longer fills the caller's
listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools,
get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy
serves only the meta-tools and the search serves only its hits, so a later
pre_mcp_call hook was reading a description the caller never listed.

The tools/list handler, the Responses MCP handler and the /v1/mcp/tools
management listing opt in with record_listing=True, since each serves the
catalog to the caller.

* fix(mcp): key the listed-tool slot by the caller's admission identity and forwarded bearer

The slot a tools/list records for a later tools/call was keyed by (user_id, api_key)
only, so every team-only JWT caller shared one slot and one JWT user acting in two
teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another
caller was served. The slot is now keyed by the hashed key, user, team and organization,
plus the admission credential of a caller admitted with neither a key nor a user.

The caller bearer split the slot only on client-forwarded-token and token-exchange
servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client
credentials) also forwards it upstream and served a different catalog per bearer into
one slot. The bearer now splits the slot on every server whose egress forwards it
(_consumes_caller_authorization) or exchanges it.

* fix(mcp): record only tools served by the bridge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep bridge tool metadata request-local

* refactor(mcp): centralize listed catalog recording guard

* fix(mcp): preserve base TPM reservations for listed tool calls

* fix(mcp): preserve project token reservations for listed calls

* fix(mcp): record served catalogs and preserve call message bytes

* test(mcp): align listing expectation with deferred recording

* test(mcp): audit listed metadata across callers and bridge lifecycles

* test(mcp): preserve guardrail fixture worker affinity

* refactor(mcp): expose listed catalog recording API

* feat(mcp): pass served_tools through the anthropic messages bridge

/v1/messages auto-execution now hands this request's resolved tool
definitions to _execute_tool_calls, matching the Responses and chat
completions bridges: the pre_mcp_call hook receives the description
and input schema the model was shown for that call. Request-local
only; the shared listed-tools catalog is untouched.

* test(mcp): pin served_tools handoff on the anthropic messages bridge

Mirrors the credentials-forwarding test: the request's resolved tool
definitions must reach _execute_tool_calls under served_tools so
pre_mcp_call hooks judge the call on the description and input schema
the model was shown. Fails without the previous commit's one-liner.

* style(mcp): sort the local import block ruff flagged

* fix(mcp): keep the admin include_disabled_tools view off the listed-tools catalog

GET /mcp-rest/tools/list?include_disabled_tools=true is the admin-only
configuration view: apply_tool_filters is False, so it serves the full
server catalog. Recording that response into the caller's listed-tools
slot warmed tools/call metadata no runtime listing ever served,
breaking the only-a-served-listing-records invariant (Bugbot).

The record is now gated on apply_tool_filters; disabled tools stay
unreachable (the call-time allowlist 403 fires before hooks), so the
observable fix is the slot no longer warming from a settings view.
Verified live: the new test fails on the unfixed head and passes here,
and the rest of the listed-tool-metadata suite is unchanged.

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 00:53:58 -07:00
devin-ai-integration[bot]
9a5e828310
feat(ui): inline Lens settings tab and investigation editor (#44479)
* feat(ui): move worker status into the Lens notch and New investigation into the list toolbar

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): make the Lens notch entry a settings gear that houses the worker section

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): replace the Lens worker modal with an inline Settings tab

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): section the Lens settings tab with tracing status and worker cards

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): let Lens settings sections span the full card width

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): replace the investigation setup modal with an inline side-by-side editor

New, edit, and duplicate now take over the Investigations tab body: matching
activity on the left, every setting on the right, with no wizard steps or modal

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): step the inline investigation setup vertically with traces alongside

Setup now sits on the left as three progressive steps (activity, criteria,
run) that collapse to a summary once done and reopen on click. Matching
activity stays on the right for every step. The editor gets a back control
and the Investigations notch shows a New, Editing, or Duplicate badge while
the editor is open

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): page the matching activity preview with useInfiniteQuery as it scrolls

Replace the Previous/Next offset buttons with the same infinite query and
near-tail prefetch the traces list uses, so the preview keeps loaded runs
and its title while the next page arrives

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): make the investigation step field map exhaustive over the form schema

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): always show the Settings tab label in the Lens mode switch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): collapse Lens worker cards into compact status rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the Lens Settings tab icon-only in every state

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* wip

* fix(ui): tick the Lens worker health dot so an expired heartbeat goes stale

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): regroup Lens settings, model and api layers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): dedupe Lens formatting helpers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): poll Lens once, drive the interval from data, settle mutations before invalidating

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): let Lens leaves fetch their own data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): bind drawer and trace shortcuts through react-hotkeys-hook

One useShortcut hook replaces the three hand-rolled keydown listeners in SidePanel,
the trace step tree and the log drawer, with a layer option deciding which keys a
pane claims from the panel around it. The span tree footer now renders ShortcutHints
from what is actually bound instead of hand-typed kbd text.

* refactor(ui): extract Inspector from SidePanel

Inspector.Root owns the open item, J/K stepping, Escape and full screen;
Inspector.Row marks a list entry with aria-selected and data-state and toggles
it on click or Enter/Space; Inspector.Panel is the resizable side panel with the
exit animation and click-outside rules. The runs table and section compose these
parts directly, so RunDrawer and the SidePanel prop bag go away.

* refactor(ui): model the Lens worker screen as a tagged union and slot in its ready action

workerScreen() decides between list, form and install from the worker rows,
the registration result and the edit target, so WorkerSettings switches on
one value and each card owns its own copy. The post-install CTA is now a
ReactNode slot that LensWorkspace fills instead of an onReady callback
threaded through LensSettings and WorkerSettings. Clipboard copy state lives
in WorkerInstall as mutations, and the styled settings leaves export Props
types, set data-slot and accept native element props.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): compose the Lens setup stepper from SetupStep children

Each step's heading, summary and fields now live together in one
SetupStep instead of four parallel structures keyed by index, and the
last-step spacing comes from CSS rather than a passed index. The mode
prop is now required since InvestigationsView always passes it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the Lens Settings panel mounted so a pending worker install survives tab switches

The settings TabsContent unmounted WorkerSettings whenever another tab was
active, dropping the one-time worker token shown during install. The panel
now uses keepMounted, and the workspace test registers a worker, switches
tabs and back, then follows the connected worker into the first
investigation through the slotted CTA.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover the Lens worker install waiting-to-connected transition

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): give the Lens activity preview a grouped contract and a structural debounce

useMatchingActivity now owns its return types (scope options, preview
status, page and optional manual selection) instead of borrowing them
from the components it feeds, and the preview takes those groups plus
the section attributes. The clear-selection action moves into the
preview footer, RunList becomes a RunRow leaf, and ScopeFields drops
the unused nameField and id props now that MetadataFilters calls useId
itself. The preview scope settles through a hashKey-based
useDebouncedValue instead of JSON round-tripping into state, and the
loading title follows isPlaceholderData since the query keeps previous
data. RunFields and AnalysisModelField take the analysis models and the
model gate as two objects instead of seven flat props.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): select the Lens investigations screen with a pure tagged union

investigationScreen maps the list query and the route to one of loading,
failed, welcome, list, detail, setup or missing, so the view can switch
instead of juggling mutually exclusive booleans. The status model gains
activeJob, carries connected inside Readiness and folds the activity probe
into one ActivityCheck value for the welcome page

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover the Lens preview footer clear action and the preview debounce

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): let Lens investigation leaves own their URL slice and express intent

InvestigationsView renders the screen union and owns every write through
useInvestigationActions, so leaves receive on* handlers instead of the API
writer. useInvestigationResults becomes useRunSnapshot; FindingsTab,
HistoryTab, RunPicker and RequestEvidenceSheet read their own nuqs slice
and run their own queries. The finding sheet becomes an Inspector side
panel (FindingDetails) keyed per finding, with the trace and request
evidence sheets grouped in EvidenceSheets. WatchAllBanner owns its
mutation, the welcome page takes the readiness and activity values, the
progress sampler records on the wall clock outside render, and run
history invalidates when the list reports a scheduler-started job

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover Lens investigation intents, pause, cancel, history refresh and the finding panel

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the inspector open behind sheet overlays and mount HotkeysProvider from a client wrapper

* feat(ui): open Lens runs and evidence in the inspector panel

A quote's original trace step or logged request now stacks inside the finding
panel behind a back link, keeping the finding and its feedback draft mounted.
The detail Runs tab and the setup activity preview open runs in the same panel
with J/K stepping, so the TraceSheet and RequestEvidenceSheet modals are gone.
Picking a different finding or run clears any stacked evidence from the URL.

* feat(ui): open Lens investigations in the inspector panel beside the list

The investigations list stays on screen and a row opens its investigation in the
side panel, so J/K walk investigations and their open findings in display order
and the selected row carries the same highlight as runs. The panel body is the
former detail page; findings and runs opened inside it nest their own inspector,
which claims the keys from the one around it while open. Opening an investigation
and peeking at a finding now replace each other in the URL.

* fix(ui): run the Lens notch border along the tab pill and flag only a disconnected worker

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): show the shortcut hints in every inspector panel

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): extract the Lens dot field into composable DotFieldRoot and DotFieldCanvas

Move the dot grid model and canvas painter out of TracesTimeline into
components/lens/dotField so other Lens surfaces can reuse it. The root
owns layout and context; overlays compose as children. agoLabel moves to
lens/model/format and the unused columnTop helper is dropped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): drop the manual refresh button from the runs time controls

Live polls and range changes refetch, so the button only cleared the zoom, which Escape already does

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): extract the Lens run search into a composable SearchBox primitive

Move the query parser, glob matcher and autocomplete out of runSearch into
components/lens/search, generic over a QueryLanguage (field specs plus
what free text searches). SearchBox.Root owns the ProseMirror state, menu
and keyboard; SearchBox.Input and SearchBox.Suggestions compose under it.
Clause highlighting becomes a ProseMirror plugin built from the language.
RunSearch now only declares the run fields and composes the parts, so the
investigations tab can define its own language and reuse the same box.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): label the runs range by preset while Live and pin it once paused

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): split the Lens search language from where its data lives

QueryLanguage is now pure vocabulary (keys, groups, icons). Reading
fields off loaded items moves to a ClientIndex consumed by a separate
evaluator, and value suggestions come from an injectable ValueSource, so
a server-backed runs list can plug in a facet lookup while the
investigations tab keeps filtering in memory. The parsed query serializes
to a typed SearchQuery (text terms plus eq/neq/glob/nglob filters) that
the client evaluator consumes today and a server can consume later. The
suggestion menu shows a loading row while a source is still answering.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(ui): format the Lens SearchBox and its test with the project prettier config

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): mirror the Lens run filters as trace SQL with a copyable curl in the search footer

The suggestions footer gains a slot, and SearchBox.ApiHint fills it with the
API equivalent of the typed query: a dialect chip, a one-line preview and a
Copy as curl button. The runs box translates each filter to a predicate over
the agent_traces_by_key rollup, bounded to the range the list shows, and
copies a POST to /v1/traces/query. Any other list can plug its own translate
into the same part.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): show the Lens introduction as a first-visit dialog with a typed don't-show-again

The guided setup no longer replaces the Lens tabs. It opens in a dialog on
the first visit of a session or from ?setup=lens, with a close and a
"Don't show this again" checkbox in its top-right corner. The header
"Set up Lens" button is gone. Dismissal state lives in a new schema-validated
web storage helper (src/lib/storage.ts) that reads through
useSyncExternalStore, so server renders see the fallback and other tabs stay
in sync; only Lens uses it for now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): keep only Copy as curl in the Lens run search footer

Drop the SQL chip and predicate preview; SearchBox.ApiHint becomes
SearchBox.CopyCommand, which takes the command for the current query.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): align the Lens investigations list with the traces list

Use the shared query SearchBox with investigation fields (name, agent, status, schedule), match the traces toolbar, and drop the count footer and inner padding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): let Lens settings bring back the introduction after don't show again

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): share one InspectorTable between the traces list and the Lens investigations tree

Compose TanStack Table, react-virtual and the shadcn table cells into InspectorTable parts (Root, Grid, Header, Body, Row, Indent). Investigations get findings as real sub-rows with TanStack expansion instead of a hand-rolled flattener.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): break Lens import cycles and move shared pieces out of lens

Search and the dot field go to components/shared, run search and the preview
button go to view_logs where they are consumed. Lens api, services and demo
live under data/, all URL state in route.ts, storage keys in storage.ts, and
the session frame styles become cva variants.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): one Lens readiness source and one onboarding flow

Readiness is computed once in model/readiness and read through
useLensReadiness, replacing useLensSetup, status.readiness and the welcome
screen's own checks. The Investigations welcome now renders the same
onboarding steps as the introduction dialog, with permissions and actions
coming from an OnboardingProvider instead of props passed down four levels.
StepIndicator and StateMessage are shared lens components, and the step
panels are labelled accordion regions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): read the Lens token from services and use semantic status colors

Lens services carry the access token they were built for, so trace evidence,
readiness and onboarding read it from context instead of a prop threaded
through six components. List and history invalidation lives in one data
hook. Status colors use the success, warning and destructive tokens, and
template-literal class names go through cn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): split Lens demo fixtures from the fake demo APIs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): show Lens check history as a dot timeline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* fix(ui): clear stale Lens evidence on run change and keep read-only users off Settings

Also names inline option objects to bring local/no-large-inline-object-arg back under budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:20:14 +00:00
devin-ai-integration[bot]
984b4134be
fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them (#44202)
* fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): reuse litellm secret redaction and mask configured DB passwords exactly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy-extras): name the password alternation in _redact_credentials

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy-extras): wrap v1 migration ERROR lines at 120 columns

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): integration cells for v1 migration ERROR logging and password redaction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): integration cells for component DB env vars, JSON logs, migration Job and v2 resolver

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): give every v1 migration integration cell the 900s timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): bind the recovery forwarder before migrating so the retry cannot race it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): drop the slow P3005 integration cell and bound migration subprocesses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): keep command repr and tolerate non-sequence cmd in migration error logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy-extras): double the subprocess boundary in the cmd=None retry test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-03 22:45:16 -07:00
devin-ai-integration[bot]
cf22deb96a
test(ci): fix six CircleCI test regressions on main (#44429)
* test(ci): fix four CircleCI regressions on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ci): stop reloading auth_checks in unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(constants): cover CLI JWT expiry env parsing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(gateway): restore proxy lifespan after importing gateway.main in launch tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 20:55:01 -07:00
moe-berri
2936307661
feat(lens): guide setup through the first investigation (#44475)
* feat(lens): guide setup through the first investigation

* fix(lens): restore the onboarding reference visuals

* fix(lens): compact onboarding and animate gateway flow

* feat(lens): refine onboarding motion and linked examples

* feat(lens): turn the LED swarm into organized dot groups

* fix(lens): make the LED dot flow visibly animate

* feat(lens): refine the swarm scale palette and motion

* fix(lens): preserve investigations during activity refresh errors
2026-10-03 20:16:41 -07:00
devin-ai-integration[bot]
b61376a99e
feat(sdk): add run_tool_loop and arun_tool_loop helpers (#44381)
* feat(sdk): add run_tool_loop and arun_tool_loop helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(tests): allow-list bounded tool-loop recursion in recursive detector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(sdk): harden run_tool_loop per review

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): keep uv.lock at revision 3

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(tests): authorize typing-extensions PSF-2.0 license

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 20:10:30 -07:00
moe-berri
5d513d8053
fix(lens): align source setup with available worker images (#44476)
* fix(lens): document working preview setup before coordinated releases

* docs(lens): keep current setup guidance factual and preserve Helm options

* fix(lens): support explicit worker images on Compose 2
2026-10-03 19:42:38 -07:00
devin-ai-integration[bot]
5ddcc45b3a
feat(ui): share trace drawer as a closable SidePanel and polish Lens (#44473)
Some checks failed
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-infra-root (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(ui): extract trace drawer into a shared SidePanel that closes on outside press

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): center Lens mode switch in a notch joined to the content card

Larger Traces/Investigations switch, a subtle dot when an investigation is running or queued, and a bigger Lens title with a docs link.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): stronger Lens frame border and header spacing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): align Lens notch fillet with the notch border

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): centered Lens loading/error states and proxy JSON calls in dev

The dev server answered GET /lens with the Lens page HTML because the UI route shadowed the proxy fallback rewrite. JSON API requests now go to the proxy before page routes, and a non-JSON success body raises a readable ApiError.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): theme-scale type and shape in trace views, aligned pane bars

Add local/no-arbitrary-design-value, scoped to TraceView and Lens, banning
arbitrary font size, tracking, leading, radius, border and CSS property
values. Map the Figma-export values onto the theme scale and replace hex
colors with info/destructive tokens.

Add PaneBar, a fixed-height bordered row, and build the step tree and span
detail headers from it so their borders line up across the split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): restore AgentTracesSection emptied in e7c3092571

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): move Set up tracing into the empty runs state

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): let the runs table gate the setup CTA on an empty range

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): mark Lens demo mode with a blue toggle and frame instead of a banner

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): draw Lens notch corners with CSS borders and thicken the demo frame

The SVG corner strokes did not snap to the same device pixels as the tab and
card borders, leaving a visible offset at the join.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): drop the redundant Lens timeline header

The status, run counts, truncated agent legend and range span all repeated the Live toggle, runs table and range picker

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): keep drawer shortcuts out of open menus and use usehooks-ts for timers and observers

SidePanel J/K/Esc now yields to menus and listboxes, not just dialogs.
The step tree shortcut footer wraps instead of clipping in narrow columns.
Replace hand-rolled timeout, keydown, media query and ResizeObserver effects
with usehooks-ts, and drop routine doc comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): offer tracing setup when filters hide every run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): replace Lens runs filters with a single ProseMirror query box

Agent and status dropdowns are gone. One query box (react-prosemirror) takes
free text plus key:value clauses (-key:value, key:*glob*) over name, agent,
status, model, input and trace_id, with field and value autocomplete. The
editor emits after a 150ms pause so typing no longer re-renders the runs view
per keystroke, and the URL keeps only q.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): keep the open trace on screen while the next one loads

Switching to an unvisited trace remounted the panel and flashed a loading skeleton.
The drawer now keeps the previous trace visible, dimmed and inert, until the new one arrives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): virtualize the Lens runs table and drop its status footer

The footer's run count only tracked how many pages had loaded and
"Updated just now" never changed, so it carried no signal. The zoom
clear button moves onto the timeline. Rows now render through
@tanstack/react-virtual so scrolling deep into a range stays cheap

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): load traces with Suspense and show the previous run via useDeferredValue

Replaces keepPreviousData with the React pattern for showing stale content while fresh content loads.
Load failures go through an error boundary that retries the query on reset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): open investigation details from list rows instead of the edit dialog

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 02:20:14 +00:00
devin-ai-integration[bot]
c80e6e2474
test(ui): wait for step search value to settle in TraceDrawer test (#44470)
* test(ui): wait for step search value to settle in TraceDrawer test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): dispatch search typing and nav keys deterministically in TraceDrawer test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 18:25:01 -07:00
devin-ai-integration[bot]
62fb808d4b
feat: improve Lens dev seeding and live UI (#44468)
* feat: seed Lens dev with configurable load profiles

* feat: run Lens UI live through the dev launcher

* fix: verify Lens UI startup before seeding

* fix(ui): render run timestamps on one compact line

The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): use the root query client for the Lens demo

Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: drop unused synthetic_spend reference from query help

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(traces): drop the deeplite fixtures and map every SDK to a logo

The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata

The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): inject the Lens sample through data sources instead of demo checks

Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): infinite scroll for the agent runs list

Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): keep the whole Lens view in the URL

Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs

RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(traces): cargo fmt captures tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 01:03:11 +00:00
moe-berri
8520626e7a
fix(lens): preserve approved worker digests and harden its image (#44467)
* fix(lens): pin worker dependencies and support approved image digests

* fix(lens): include locked dependencies and release identity in build context

* fix(lens): select the dev worker package for SHA-tagged charts
2026-10-03 17:54:17 -07:00
berriai-litellm-provider-info-sync[bot]
a66adb4ff8
chore(cost-map): sync openrouter prices from the models API (#44466)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 17:34:05 -07:00
devin-ai-integration[bot]
9d16412341
fix(caching): count tool_call cache_control marks in the injection census (#43556)
* fix(caching): count tool_call cache_control marks in the injection census

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): remove cache census casts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): only skip injection on message or content marks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): skip injection on messages whose tool calls carry marks

Reverts 9f08d8aef8. A default 5m mark injected on assistant text lands
before the client's 1h tool_use mark, which Anthropic rejects with a 400
because a 1h breakpoint must not follow a 5m one. Keeping the full census
in the skip check leaves the client's tool_call breakpoint as the only
one on that message.

* fix(caching): count every client tool_call cache mark in the breakpoint census

The census gated tool_call marks on type function and dict shape, so a client mark on a call without a type or with a string cache_control slipped past the count and injection overflowed the 4 breakpoint cap. Count any non-None tool_call mark, keep the server tool exclusion, and add integration cells for the capped surfaces, the yaml stand-down, Bedrock and Gemini, and router affinity

* test(integration): hold the upstream so the worker kill lands mid-burst

* refactor(caching): reuse the transform's server tool lookup in the breakpoint census

The census now calls the same helper the Anthropic transform uses to decide
whether a marked tool call becomes a server tool block, so the two cannot
drift apart. The owned-proxy burst test waits up to 90 seconds for the burst
to reach the wire before it kills a worker

* refactor(anthropic): move the server tool rebuild check under llms/anthropic

* test(integration): audit the tool call mark census across chat, messages, responses, and chaos

Adds the /audit cells for the breakpoint census on assistant tool_calls marks: the
Responses stream bridge, the OpenAI and Anthropic SDK clients, in-process Pydantic
messages, response cache twins, request-level points, a provider 401 on a capped request,
malformed provider_specific_fields and tool_call ids, null or empty points, a points
update mid-burst, a proxy restart mid-burst, and a worker SIGKILL that picks the worker
holding the burst's upstream connections

* test(integration): close the SDK clients the cache census cells open

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:33:48 -07:00
tin-berri
6370104c53
feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444)
* feat(enterprise): bundle LiteAdmin Slack worker with native gateway login

* fix(enterprise): preserve gateway prefixes during native Slack linking

* fix(enterprise): retain native Slack linking on the admin backend

* fix(enterprise): reuse shared native Slack connection services
2026-10-03 17:28:18 -07:00
moe-berri
4b67a2b845
feat(roi): default people and branch lists to matched accounts (#44465)
* feat(roi): show matched people by default in contributor lists

* fix(roi): keep matched filter tabs readable on narrow screens

* fix(roi): retain spend-only users and support older browsers
2026-10-03 17:24:18 -07:00
devin-ai-integration[bot]
f850b2c324
test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment (#44451)
* test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): compare every non-transport provider header and check for late provider requests at session end

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): name TranslationTestCase fields after litellm and provider sides and drop regressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prefix checked TranslationTestCase fields with expected_ and name the fake reply mock_provider_response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(integration): name TranslationTestCase fields in the translation README

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): add a claude-opus-5-5 base case and deployment next to claude-sonnet-4-6

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): name translation cases <MODEL>_TEST_CASE and document the naming rule

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(integration): move translation test rules into tests/integration/translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 00:09:58 +00:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
dfdd496db8
fix(tests): match the OS bind error in the owned-proxy port-race retry (#44462)
* fix(tests): match the OS bind error in the owned-proxy port-race retry

* test(integration): keep the port-race predicate pure so its unit tests stay in-process

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:07:10 -07:00
devin-ai-integration[bot]
f0eda6d2a6
fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API (#44419)
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API

Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded

Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode

* fix(health): resolve the test connection mode from the deployment when the request omits it

The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.

* fix(health): resolve an omitted ahealth_check mode the way the proxy does

* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode

A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.

* fix(health): shape test connection probe params for the model the request probes

A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden

* fix(health): report an early ahealth_check failure as itself, not as a missing mode

With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:05:47 -07:00
berriai-litellm-provider-info-sync[bot]
4d30f8c59b
chore(cost-map): add azure_ai/kimi-k2-thinking retirement date from the Azure retired models page (#44455)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 16:44:13 -07:00
devin-ai-integration[bot]
21ecd0af55
fix(traces): reject conflicting spend aliases and unrelated HTTP siblings (#44456)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-10-03 23:43:10 +00:00
devin-ai-integration[bot]
6d73fa6b49
fix(vertex_ai): forward system and tools to partner model count_tokens (#43900)
* fix(vertex_ai): forward system and tools to partner model count_tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): avoid mutable token request construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): return partner count_tokens provider errors as values so the proxy falls back locally

* test(integration): cover Vertex AI partner count_tokens forwarding and local fallbacks

Add wire-level cells for /v1/messages/count_tokens, /utils/token_counter,
/v1/responses/input_tokens and the Gemini countTokens route on a Vertex AI
Claude deployment: the system prompt and tools reach the partner
count-tokens endpoint verbatim, null fields stay out of the body, malformed
tools are rejected before any peer call, peer, token-endpoint and connection
failures fall back to the local tokenizer unless disable_token_counter is
set, generation on the same deployment keeps working, and concurrent bursts
survive a peer outage, a slow peer and a worker SIGKILL. The sdk cells cover
litellm.acount_tokens the same way.

The _support/process.py and _support/client.py harness files are brought to
main's content so the self-booting cells read INTEGRATION_PROXY_READY_SECONDS
instead of a fixed 70 s boot budget.

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 16:36:14 -07:00
devin-ai-integration[bot]
1ef0fe9790
fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check (#43741)
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): key the zero-cost cache by the resolved model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): keep the zero-cost verdict per requested name and include hidden aliases

* test(integration): audit the zero-cost bypass through hidden model_group_alias names

* test(router): cover the extracted routing strategy switch

* test(integration): record a pre-flip burst before the alias flip

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 16:35:11 -07:00
moe-berri
cb17588276
feat(lens): coordinate worker releases and bundled installs (#44428)
* feat(lens): coordinate worker versions and bundled installs

* test(lens): exercise bundled Compose startup and restart in CI

* fix(lens): refund failed model requests without a response

* test(lens): verify trace persistence in the bundled stack

* fix(lens): align Helm images and isolate Compose storage

* fix(lens): reject worker builds without release identity

* fix(lens): encode Compose credentials and normalize worker versions

* fix(lens): refuse worker recommendations for unidentified builds
2026-10-03 16:33:45 -07:00
yuneng-jiang
427158eb5b
feat(ui): show invitation and reset password links in a copyable field (#44454)
* feat(ui): show invitation and reset password links in a copyable field

Put the link in a read-only input with a Copy button beside it, stack the
User ID and link labels above their values, and focus Copy on open so the
field shows the start of the URL. Copy now goes through the shared
copyToClipboard helper, which falls back to a selection copy where the
Clipboard API is unavailable.

* fix(ui): keep focus on the copy control after a fallback clipboard copy

The execCommand fallback focused a temporary textarea and removed it, so
focus fell to the page body and a second Enter on Copy did nothing.
Restore focus to the element that had it. Move the rendered dialog
tests to the integration tier.
2026-10-03 16:23:13 -07:00
berriai-litellm-provider-info-sync[bot]
5f969982fc
fix(azure): set text-embedding max input to 8192 from the models sold directly page (#44449)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 16:15:02 -07:00
moe-berri
1a7023366f
feat(roi): measure shipping velocity, quality, and recorded spend (#44426)
* feat(ui): prototype observed engineering ROI dashboard

* feat(roi): replace effort estimates with measured repository metrics

* fix(roi): finish connection recovery and generated API contracts

* fix(roi): show merged changes before accounts are linked

* fix(roi): preserve selected report tab across refreshes

* fix(roi): recover app authorization and keep detail values readable

* fix(roi): reuse the shared OAuth HTTP client

* feat(roi): combine providers and compare equal reporting periods

* docs: explain ROI metrics for first-time readers

* fix(roi): preserve connections and scheduled reports during setup

* ci(roi): assign database contracts to the active Postgres shard

* fix(roi): preserve issue counts and normalized connections

* fix(ui): compact ROI dashboard header and metrics

* fix(ui): show ROI repository count with expandable list

* fix(ui): wrap ROI controls within narrow panels

* fix(roi): restore sample report preview and simplify setup
2026-10-03 23:07:34 +00:00
yujonglee
c7e60f03de
fix(tracing): preserve spend identity and gateway correlation (#44421)
* fix(tracing): preserve spend identity and gateway correlation

* test(tracing): refresh real SDK spend captures

* fix(tracing): resolve complete gateway attempt costs across SDKs

* docs: add trace cost screenshot for PR 44421

* update fixtures

* wip

* docs: remove trace cost screenshot from PR evidence
2026-10-03 16:00:49 -07:00
devin-ai-integration[bot]
01b4ffe16b
fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint (#43646)
* fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock_mantle): price Claude chat on the Mantle row and route region-prefixed ids

A Mantle request always carries a region, so a Claude id with no bedrock_mantle/<region>/ row fell through the model-info lookup to the bare Bedrock row, which the bedrock provider family also matches, and billed about 10 percent under the Mantle price. The lookup now tries the provider's region-free row before the bare model. The Claude route test asserts the Mantle row, a region-prefixed Claude id is covered end to end, and the provider config map references the Mantle config directly.

* test(bedrock_mantle): cover supported params for Claude and open-weight Mantle ids

* test(bedrock_mantle): audit the Claude chat bridge on the integration rig

Adds the deterministic cells from the /audit of the Mantle Claude chat
bridge: wire-level translation on every chat, responses, and messages
route, SigV4 and bearer auth, region prefixes, api_base suffixes,
unsupported params with and without drop_params, malformed model ids,
upstream errors, the response-cache hit, the health check, pricing from
the Mantle row for Claude and non-Claude ids with a bare Bedrock twin,
and chaos cells for a mixed burst, an upstream outage, slow streams,
and a worker kill on an owned two-worker proxy

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 15:24:41 -07:00
joshua-berri
0ea166c160
fix(mcp): preserve upstream tool schemas and parameter headers (#44425)
* fix(mcp): preserve tool schemas and modern parameter headers

* fix(mcp): allow bounded cold schema worker startup

* fix(mcp): align catalog deadlines and bound schema traversal

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-03 15:00:43 -07:00
devin-ai-integration[bot]
98337c9334
fix(lens): release budget reservations when the analysis model call fails (#44431)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:39:27 -07:00
devin-ai-integration[bot]
41a3781d4e
fix(proxy): treat Postgres connection exhaustion as backpressure, not poison rows (#44266)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
9a4b1951a2
fix(proxy): treat Postgres connection exhaustion as DB unavailable, not poison spend-log rows (#44270)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
e9bf2cfd01
refactor(types): replace Any with proven types in 9 files (#44389)
* refactor(types): replace Any with proven types in 13 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep provider error paths for malformed prefetch and poll JSON

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's login body parsing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's Copilot auth and budget alert typing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for Any sweep 20261003_2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): load audit video deployments at proxy start

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's video-edit prefetch handling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fix audit chaos Responses cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): probe every route after audit worker kill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 21:03:16 +00:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
Misbah Syed
fd14e51ecf
feat(docker): add a Windows quickstart in PowerShell, and a styled terminal for both quickstarts (#44310)
* feat(docker): Windows quickstart in PowerShell, and a styled terminal for both quickstarts

scripts/quickstart.ps1 does what scripts/quickstart.sh does, for Windows:
  powershell -c "irm https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.ps1 | iex"
It asks the same two questions, prints the same lines, writes the same .env
(UTF-8 without a byte order mark, LF endings, readable only by the current
Windows account), and runs in Windows PowerShell 5.1 and PowerShell 7.

Both scripts now give a person at a terminal colors, a check mark per step, a
spinner while Docker starts, and a framed summary. Agents, CI, log files, and
NO_COLOR get the same lines as plain text.

* feat(docker): support podman and rancher desktop in the quickstart scripts

Both quickstarts hard-required the docker CLI. They now pick the first
available engine among docker, podman, and nerdctl (Rancher Desktop in
containerd mode; its dockerd mode already provides a docker CLI), route
every invocation through it, and tailor the start hint (podman machine
start) and the printed stop/logs/volume commands to that engine

* fix(docker): test port availability by binding instead of connecting

On a WSL2-backed engine (Podman, Rancher Desktop), the Windows localhost
relay swallows connection refusals on closed ports, so every connect
waits out its 2-second timeout and Test-PortFree reported ports 4000 to
4099 all taken on a machine with none of them in use. Binding the port
answers instantly and accurately

* fix(quickstart): address review findings

- The PowerShell script names the project with the same POSIX cksum as the
  shell script, so the earlier-install check sees the database from either
  script.
- Both scripts stop when git tracks .env in the install folder.
- .env is written to a temp file with owner-only permissions and moved into
  place, so a failed write never leaves a partial file.
- Errors stay plain text when stderr is redirected.
- The spinners remove their temp files on exit and Ctrl+C, and the shell
  spinner no longer runs date on every frame.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(quickstart): offer to copy the admin password to the clipboard

After the summary, the quickstarts ask "What next?": copy the admin
password, open the admin UI, or finish. The password is piped to the
clipboard (pbcopy, wl-copy, xclip, xsel, clip.exe, or Set-Clipboard), so
it never appears on screen or in the process list. Over SSH, or with no
clipboard, the option is left out and the summary points to .env as
before. The PowerShell menu now honours Ctrl+C.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(quickstart): write .env through mktemp, let Ctrl+C cancel the PowerShell menu, drop redundant comments

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(quickstart): remove the in-progress .env temp file when interrupted

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Mubashir Osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:32 -07:00
ishaan-berri
dd31692282
feat(lens): always-on investigations with findings and investigations tables (#44418)
* feat(lens): record run steps, trigger and exact run windows on investigations

* feat(lens): scan only traces since the last run and keep a capped step log

* feat(lens): log each analysis model call with its model, tokens and cost

* feat(lens): accept agent and time window on run now and add turn-all-on

* test(lens): cover new-traces-only windows, manual runs and the step cap

* test(lens): cover run now overrides and turning paused investigations on

* chore(ui): regenerate api types for lens run steps and run now options

* feat(lens): group open findings into one row per problem with agent filters

* test(lens): cover the findings table grouping, filters and schedule labels

* feat(lens): add a findings table across all investigations

* feat(lens): show a live step feed with the model behind each call

* feat(lens): offer turning all paused investigations on

* feat(lens): let run now pick an agent and time window

* test(lens): cover run now request building

* feat(lens): show the step feed and run now dialog on an investigation

* feat(lens): open run now choices instead of running immediately

* feat(lens): open findings first and peek a finding without leaving the table

* feat(lens): fold investigation actions into the findings toolbar

* feat(lens): show each investigation's schedule and open findings

* feat(lens): name the agent on a finding

* feat(lens): keep new investigations watching every 15 minutes by default

* feat(lens): show the watch schedule outside advanced options

* feat(lens): send run now options and turn-all-on from the dashboard

* test(lens): give demo runs steps and a trigger

* feat(lens): let the findings table fill the screen

* test(lens): add steps and trigger to progress fixtures

* test(lens): add steps and trigger to status fixtures

* test(lens): cover the default watch schedule in setup

* test(lens): cover run now choices from an investigation

* test(lens): open saved investigations from the manage view

* feat(lens): use one tab bar for traces, findings and investigations

* feat(lens): place page actions on the lens tab row

* feat(lens): drop the nested tabs and edit investigations in place

* feat(lens): show investigations as a table with run now and edit

* feat(lens): name each findings row for screen readers

* test(lens): open saved investigation links on findings

* test(lens): reach findings and investigations from the top tabs

* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan

* fix(lens): keep run now since-last-run windows even with an agent override

* test(lens): cover that manual runs never skip scheduled traces

* test(lens): cover run now windows with agent and lookback overrides

* fix(lens): group findings without Map.groupBy and expose sampled runs

* fix(lens): open older findings and review every merged copy from one row

* fix(lens): hide edit and run now from read-only viewers

* chore(lens): drop restating comments from the findings table

* chore(lens): drop restating comments from the step feed

* chore(lens): drop restating comments from the paused banner

* chore(lens): drop restating comments from header actions

* chore(lens): drop restating comments from run now

* test(lens): cover merged findings and read-only investigation rows

* fix(lens): record a model step even when the response has no usage

* test(lens): cover model steps with and without reported usage

* fix(lens): keep a merged finding open when one of its updates fails

* refactor(lens): accept update results from the investigations view

* refactor(lens): accept update results in investigation actions

* test(lens): cover retrying a merged finding after a failed update
2026-10-03 20:27:02 +00:00
devin-ai-integration[bot]
f445e466b4
refactor(traces): extract snapshot cache (#44424)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 13:03:30 -07:00