Move the timeline renderer to shared/timeline/Timeline taking buckets,
a selected window, and callbacks as props, and TimeRangeControls to
shared/timeline. Lens keeps bucketRuns as the adapter that converts
loaded traces into buckets, preserving behavior without the histogram
endpoint.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(ui): add url-state agent skill
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): migrate dashboard URL state to nuqs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): cover chat URL id sync with the real chat shell provider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(dashboard): route dev API calls on Accept and fail fast on non-JSON 2xx
The next dev rewrite that sends API calls to the proxy keyed on
Content-Type: application/json, which openapi-fetch rightly omits on a
bodyless GET, so GET /lens fell through to the Lens page and returned
HTML. Route on Accept: application/json instead, send it from both HTTP
clients, turn a non-JSON 2xx into a non-retryable ApiError in the typed
client, and stop react-query from retrying ApiError below 500.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(dashboard): accept a JSON body served without a JSON content type
Test fakes and some servers hand back JSON as text/plain, so the typed
client only rejects a 2xx whose body does not parse as JSON. The
system_one request test expects the Accept header the legacy client now
sends.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): own trace reads behind a cached TraceStore port
Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.
Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.
Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): seed large profiles with server-side copies and long sessions
Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chores
* style(lens): float the investigation setup badge on the tab edge
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): restructure the trace drawer and polish its layout
Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.
Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* typesafety
* wip
* fmt
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): move worker status into the Lens notch and New investigation into the list toolbar
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): make the Lens notch entry a settings gear that houses the worker section
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): replace the Lens worker modal with an inline Settings tab
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): section the Lens settings tab with tracing status and worker cards
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): let Lens settings sections span the full card width
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): replace the investigation setup modal with an inline side-by-side editor
New, edit, and duplicate now take over the Investigations tab body: matching
activity on the left, every setting on the right, with no wizard steps or modal
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): step the inline investigation setup vertically with traces alongside
Setup now sits on the left as three progressive steps (activity, criteria,
run) that collapse to a summary once done and reopen on click. Matching
activity stays on the right for every step. The editor gets a back control
and the Investigations notch shows a New, Editing, or Duplicate badge while
the editor is open
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): page the matching activity preview with useInfiniteQuery as it scrolls
Replace the Previous/Next offset buttons with the same infinite query and
near-tail prefetch the traces list uses, so the preview keeps loaded runs
and its title while the next page arrives
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): make the investigation step field map exhaustive over the form schema
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): always show the Settings tab label in the Lens mode switch
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): collapse Lens worker cards into compact status rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): keep the Lens Settings tab icon-only in every state
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* wip
* fix(ui): tick the Lens worker health dot so an expired heartbeat goes stale
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): regroup Lens settings, model and api layers
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): dedupe Lens formatting helpers
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): poll Lens once, drive the interval from data, settle mutations before invalidating
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): let Lens leaves fetch their own data
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): bind drawer and trace shortcuts through react-hotkeys-hook
One useShortcut hook replaces the three hand-rolled keydown listeners in SidePanel,
the trace step tree and the log drawer, with a layer option deciding which keys a
pane claims from the panel around it. The span tree footer now renders ShortcutHints
from what is actually bound instead of hand-typed kbd text.
* refactor(ui): extract Inspector from SidePanel
Inspector.Root owns the open item, J/K stepping, Escape and full screen;
Inspector.Row marks a list entry with aria-selected and data-state and toggles
it on click or Enter/Space; Inspector.Panel is the resizable side panel with the
exit animation and click-outside rules. The runs table and section compose these
parts directly, so RunDrawer and the SidePanel prop bag go away.
* refactor(ui): model the Lens worker screen as a tagged union and slot in its ready action
workerScreen() decides between list, form and install from the worker rows,
the registration result and the edit target, so WorkerSettings switches on
one value and each card owns its own copy. The post-install CTA is now a
ReactNode slot that LensWorkspace fills instead of an onReady callback
threaded through LensSettings and WorkerSettings. Clipboard copy state lives
in WorkerInstall as mutations, and the styled settings leaves export Props
types, set data-slot and accept native element props.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): compose the Lens setup stepper from SetupStep children
Each step's heading, summary and fields now live together in one
SetupStep instead of four parallel structures keyed by index, and the
last-step spacing comes from CSS rather than a passed index. The mode
prop is now required since InvestigationsView always passes it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): keep the Lens Settings panel mounted so a pending worker install survives tab switches
The settings TabsContent unmounted WorkerSettings whenever another tab was
active, dropping the one-time worker token shown during install. The panel
now uses keepMounted, and the workspace test registers a worker, switches
tabs and back, then follows the connected worker into the first
investigation through the slotted CTA.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(ui): cover the Lens worker install waiting-to-connected transition
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): give the Lens activity preview a grouped contract and a structural debounce
useMatchingActivity now owns its return types (scope options, preview
status, page and optional manual selection) instead of borrowing them
from the components it feeds, and the preview takes those groups plus
the section attributes. The clear-selection action moves into the
preview footer, RunList becomes a RunRow leaf, and ScopeFields drops
the unused nameField and id props now that MetadataFilters calls useId
itself. The preview scope settles through a hashKey-based
useDebouncedValue instead of JSON round-tripping into state, and the
loading title follows isPlaceholderData since the query keeps previous
data. RunFields and AnalysisModelField take the analysis models and the
model gate as two objects instead of seven flat props.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): select the Lens investigations screen with a pure tagged union
investigationScreen maps the list query and the route to one of loading,
failed, welcome, list, detail, setup or missing, so the view can switch
instead of juggling mutually exclusive booleans. The status model gains
activeJob, carries connected inside Readiness and folds the activity probe
into one ActivityCheck value for the welcome page
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(ui): cover the Lens preview footer clear action and the preview debounce
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): let Lens investigation leaves own their URL slice and express intent
InvestigationsView renders the screen union and owns every write through
useInvestigationActions, so leaves receive on* handlers instead of the API
writer. useInvestigationResults becomes useRunSnapshot; FindingsTab,
HistoryTab, RunPicker and RequestEvidenceSheet read their own nuqs slice
and run their own queries. The finding sheet becomes an Inspector side
panel (FindingDetails) keyed per finding, with the trace and request
evidence sheets grouped in EvidenceSheets. WatchAllBanner owns its
mutation, the welcome page takes the readiness and activity values, the
progress sampler records on the wall clock outside render, and run
history invalidates when the list reports a scheduler-started job
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(ui): cover Lens investigation intents, pause, cancel, history refresh and the finding panel
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): keep the inspector open behind sheet overlays and mount HotkeysProvider from a client wrapper
* feat(ui): open Lens runs and evidence in the inspector panel
A quote's original trace step or logged request now stacks inside the finding
panel behind a back link, keeping the finding and its feedback draft mounted.
The detail Runs tab and the setup activity preview open runs in the same panel
with J/K stepping, so the TraceSheet and RequestEvidenceSheet modals are gone.
Picking a different finding or run clears any stacked evidence from the URL.
* feat(ui): open Lens investigations in the inspector panel beside the list
The investigations list stays on screen and a row opens its investigation in the
side panel, so J/K walk investigations and their open findings in display order
and the selected row carries the same highlight as runs. The panel body is the
former detail page; findings and runs opened inside it nest their own inspector,
which claims the keys from the one around it while open. Opening an investigation
and peeking at a finding now replace each other in the URL.
* fix(ui): run the Lens notch border along the tab pill and flag only a disconnected worker
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): show the shortcut hints in every inspector panel
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): extract the Lens dot field into composable DotFieldRoot and DotFieldCanvas
Move the dot grid model and canvas painter out of TracesTimeline into
components/lens/dotField so other Lens surfaces can reuse it. The root
owns layout and context; overlays compose as children. agoLabel moves to
lens/model/format and the unused columnTop helper is dropped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): drop the manual refresh button from the runs time controls
Live polls and range changes refetch, so the button only cleared the zoom, which Escape already does
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): extract the Lens run search into a composable SearchBox primitive
Move the query parser, glob matcher and autocomplete out of runSearch into
components/lens/search, generic over a QueryLanguage (field specs plus
what free text searches). SearchBox.Root owns the ProseMirror state, menu
and keyboard; SearchBox.Input and SearchBox.Suggestions compose under it.
Clause highlighting becomes a ProseMirror plugin built from the language.
RunSearch now only declares the run fields and composes the parts, so the
investigations tab can define its own language and reuse the same box.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ui): label the runs range by preset while Live and pin it once paused
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): split the Lens search language from where its data lives
QueryLanguage is now pure vocabulary (keys, groups, icons). Reading
fields off loaded items moves to a ClientIndex consumed by a separate
evaluator, and value suggestions come from an injectable ValueSource, so
a server-backed runs list can plug in a facet lookup while the
investigations tab keeps filtering in memory. The parsed query serializes
to a typed SearchQuery (text terms plus eq/neq/glob/nglob filters) that
the client evaluator consumes today and a server can consume later. The
suggestion menu shows a loading row while a source is still answering.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* style(ui): format the Lens SearchBox and its test with the project prettier config
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): mirror the Lens run filters as trace SQL with a copyable curl in the search footer
The suggestions footer gains a slot, and SearchBox.ApiHint fills it with the
API equivalent of the typed query: a dialect chip, a one-line preview and a
Copy as curl button. The runs box translates each filter to a predicate over
the agent_traces_by_key rollup, bounded to the range the list shows, and
copies a POST to /v1/traces/query. Any other list can plug its own translate
into the same part.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): show the Lens introduction as a first-visit dialog with a typed don't-show-again
The guided setup no longer replaces the Lens tabs. It opens in a dialog on
the first visit of a session or from ?setup=lens, with a close and a
"Don't show this again" checkbox in its top-right corner. The header
"Set up Lens" button is gone. Dismissal state lives in a new schema-validated
web storage helper (src/lib/storage.ts) that reads through
useSyncExternalStore, so server renders see the fallback and other tabs stay
in sync; only Lens uses it for now.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): keep only Copy as curl in the Lens run search footer
Drop the SQL chip and predicate preview; SearchBox.ApiHint becomes
SearchBox.CopyCommand, which takes the command for the current query.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(ui): align the Lens investigations list with the traces list
Use the shared query SearchBox with investigation fields (name, agent, status, schedule), match the traces toolbar, and drop the count footer and inner padding.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): let Lens settings bring back the introduction after don't show again
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): share one InspectorTable between the traces list and the Lens investigations tree
Compose TanStack Table, react-virtual and the shadcn table cells into InspectorTable parts (Root, Grid, Header, Body, Row, Indent). Investigations get findings as real sub-rows with TanStack expansion instead of a hand-rolled flattener.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): break Lens import cycles and move shared pieces out of lens
Search and the dot field go to components/shared, run search and the preview
button go to view_logs where they are consumed. Lens api, services and demo
live under data/, all URL state in route.ts, storage keys in storage.ts, and
the session frame styles become cva variants.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): one Lens readiness source and one onboarding flow
Readiness is computed once in model/readiness and read through
useLensReadiness, replacing useLensSetup, status.readiness and the welcome
screen's own checks. The Investigations welcome now renders the same
onboarding steps as the introduction dialog, with permissions and actions
coming from an OnboardingProvider instead of props passed down four levels.
StepIndicator and StateMessage are shared lens components, and the step
panels are labelled accordion regions.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): read the Lens token from services and use semantic status colors
Lens services carry the access token they were built for, so trace evidence,
readiness and onboarding read it from context instead of a prop threaded
through six components. List and history invalidation lives in one data
hook. Status colors use the success, warning and destructive tokens, and
template-literal class names go through cn.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): split Lens demo fixtures from the fake demo APIs
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): show Lens check history as a dot timeline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chores
* fix(ui): clear stale Lens evidence on run change and keep read-only users off Settings
Also names inline option objects to bring local/no-large-inline-object-arg back under budget.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* feat(lens): guide setup through the first investigation
* fix(lens): restore the onboarding reference visuals
* fix(lens): compact onboarding and animate gateway flow
* feat(lens): refine onboarding motion and linked examples
* feat(lens): turn the LED swarm into organized dot groups
* fix(lens): make the LED dot flow visibly animate
* feat(lens): refine the swarm scale palette and motion
* fix(lens): preserve investigations during activity refresh errors
* feat(ui): extract trace drawer into a shared SidePanel that closes on outside press
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): center Lens mode switch in a notch joined to the content card
Larger Traces/Investigations switch, a subtle dot when an investigation is running or queued, and a bigger Lens title with a docs link.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): stronger Lens frame border and header spacing
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): align Lens notch fillet with the notch border
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): centered Lens loading/error states and proxy JSON calls in dev
The dev server answered GET /lens with the Lens page HTML because the UI route shadowed the proxy fallback rewrite. JSON API requests now go to the proxy before page routes, and a non-JSON success body raises a readable ApiError.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): theme-scale type and shape in trace views, aligned pane bars
Add local/no-arbitrary-design-value, scoped to TraceView and Lens, banning
arbitrary font size, tracking, leading, radius, border and CSS property
values. Map the Figma-export values onto the theme scale and replace hex
colors with info/destructive tokens.
Add PaneBar, a fixed-height bordered row, and build the step tree and span
detail headers from it so their borders line up across the split.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): restore AgentTracesSection emptied in e7c3092571
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): move Set up tracing into the empty runs state
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): let the runs table gate the setup CTA on an empty range
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): mark Lens demo mode with a blue toggle and frame instead of a banner
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): draw Lens notch corners with CSS borders and thicken the demo frame
The SVG corner strokes did not snap to the same device pixels as the tab and
card borders, leaving a visible offset at the join.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): drop the redundant Lens timeline header
The status, run counts, truncated agent legend and range span all repeated the Live toggle, runs table and range picker
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): keep drawer shortcuts out of open menus and use usehooks-ts for timers and observers
SidePanel J/K/Esc now yields to menus and listboxes, not just dialogs.
The step tree shortcut footer wraps instead of clipping in narrow columns.
Replace hand-rolled timeout, keydown, media query and ResizeObserver effects
with usehooks-ts, and drop routine doc comments.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): offer tracing setup when filters hide every run
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): replace Lens runs filters with a single ProseMirror query box
Agent and status dropdowns are gone. One query box (react-prosemirror) takes
free text plus key:value clauses (-key:value, key:*glob*) over name, agent,
status, model, input and trace_id, with field and value autocomplete. The
editor emits after a 150ms pause so typing no longer re-renders the runs view
per keystroke, and the URL keeps only q.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): keep the open trace on screen while the next one loads
Switching to an unvisited trace remounted the panel and flashed a loading skeleton.
The drawer now keeps the previous trace visible, dimmed and inert, until the new one arrives.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): virtualize the Lens runs table and drop its status footer
The footer's run count only tracked how many pages had loaded and
"Updated just now" never changed, so it carried no signal. The zoom
clear button moves onto the timeline. Rows now render through
@tanstack/react-virtual so scrolling deep into a range stays cheap
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): load traces with Suspense and show the previous run via useDeferredValue
Replaces keepPreviousData with the React pattern for showing stale content while fresh content loads.
Load failures go through an error boundary that retries the query on reset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): open investigation details from list rows instead of the edit dialog
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(ui): wait for step search value to settle in TraceDrawer test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): dispatch search typing and nav keys deterministically in TraceDrawer test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat: seed Lens dev with configurable load profiles
* feat: run Lens UI live through the dev launcher
* fix: verify Lens UI startup before seeding
* fix(ui): render run timestamps on one compact line
The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): use the root query client for the Lens demo
Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: drop unused synthetic_spend reference from query help
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(traces): drop the deeplite fixtures and map every SDK to a logo
The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata
The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): inject the Lens sample through data sources instead of demo checks
Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): infinite scroll for the agent runs list
Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): keep the whole Lens view in the URL
Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs
RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* style(traces): cargo fmt captures tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): show matched people by default in contributor lists
* fix(roi): keep matched filter tabs readable on narrow screens
* fix(roi): retain spend-only users and support older browsers
* feat(anthropic): workload identity federation and pluggable identity sources
Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.
Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.
The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.
Fixes#28607
Resolves LIT-6107
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries
The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.
LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.
* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle
* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries
* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly
* fix(proxy): decrypt stored litellm_params before the WIF write gate
* fix(proxy): hide WIF secret references from /health output
* fix(proxy): keep the proxy error shape on credential endpoint refusals
* fix(proxy): hide identity token file paths from /health output
* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working
The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader
* fix(auth): share one exchanged token across workers reading the same assertion
Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it
* fix: keep anthropic federation from being shadowed or leaked
An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.
The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.
The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.
* fix: unlink a staged token file a failed write leaves behind
The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.
* refactor: move anthropic jwks derivation behind a provider-owned tagged union
* fix: unlink the staged token file when its write fails at close
A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.
* fix(anthropic): close the staging descriptor before writing the shared token file
* fix(wif): judge federation writes by what they set, not what is stored
The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500
The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL
* fix(proxy): let a deployment write name a federated credential
reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.
is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.
* refactor(proxy): derive health display policy from the federation key sets
The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places
WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one
* fix(anthropic_wif): treat blank identity-source fields as unset
* test(proxy): classify the federation params in the credential slot registry
main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret
* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics
Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.
* fix(credentials): gate PATCH on WIF fields resolved from model_id
The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path
* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header
Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API
Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded
Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode
* fix(health): resolve the test connection mode from the deployment when the request omits it
The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.
* fix(health): resolve an omitted ahealth_check mode the way the proxy does
* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode
A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.
* fix(health): shape test connection probe params for the model the request probes
A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden
* fix(health): report an early ahealth_check failure as itself, not as a missing mode
With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(ui): show invitation and reset password links in a copyable field
Put the link in a read-only input with a Copy button beside it, stack the
User ID and link labels above their values, and focus Copy on open so the
field shows the start of the URL. Copy now goes through the shared
copyToClipboard helper, which falls back to a selection copy where the
Clipboard API is unavailable.
* fix(ui): keep focus on the copy control after a fallback clipboard copy
The execCommand fallback focused a temporary textarea and removed it, so
focus fell to the page body and a second Enter on Copy did nothing.
Restore focus to the element that had it. Move the rendered dialog
tests to the integration tier.
* feat(ui): prototype observed engineering ROI dashboard
* feat(roi): replace effort estimates with measured repository metrics
* fix(roi): finish connection recovery and generated API contracts
* fix(roi): show merged changes before accounts are linked
* fix(roi): preserve selected report tab across refreshes
* fix(roi): recover app authorization and keep detail values readable
* fix(roi): reuse the shared OAuth HTTP client
* feat(roi): combine providers and compare equal reporting periods
* docs: explain ROI metrics for first-time readers
* fix(roi): preserve connections and scheduled reports during setup
* ci(roi): assign database contracts to the active Postgres shard
* fix(roi): preserve issue counts and normalized connections
* fix(ui): compact ROI dashboard header and metrics
* fix(ui): show ROI repository count with expandable list
* fix(ui): wrap ROI controls within narrow panels
* fix(roi): restore sample report preview and simplify setup
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(lens): record run steps, trigger and exact run windows on investigations
* feat(lens): scan only traces since the last run and keep a capped step log
* feat(lens): log each analysis model call with its model, tokens and cost
* feat(lens): accept agent and time window on run now and add turn-all-on
* test(lens): cover new-traces-only windows, manual runs and the step cap
* test(lens): cover run now overrides and turning paused investigations on
* chore(ui): regenerate api types for lens run steps and run now options
* feat(lens): group open findings into one row per problem with agent filters
* test(lens): cover the findings table grouping, filters and schedule labels
* feat(lens): add a findings table across all investigations
* feat(lens): show a live step feed with the model behind each call
* feat(lens): offer turning all paused investigations on
* feat(lens): let run now pick an agent and time window
* test(lens): cover run now request building
* feat(lens): show the step feed and run now dialog on an investigation
* feat(lens): open run now choices instead of running immediately
* feat(lens): open findings first and peek a finding without leaving the table
* feat(lens): fold investigation actions into the findings toolbar
* feat(lens): show each investigation's schedule and open findings
* feat(lens): name the agent on a finding
* feat(lens): keep new investigations watching every 15 minutes by default
* feat(lens): show the watch schedule outside advanced options
* feat(lens): send run now options and turn-all-on from the dashboard
* test(lens): give demo runs steps and a trigger
* feat(lens): let the findings table fill the screen
* test(lens): add steps and trigger to progress fixtures
* test(lens): add steps and trigger to status fixtures
* test(lens): cover the default watch schedule in setup
* test(lens): cover run now choices from an investigation
* test(lens): open saved investigations from the manage view
* feat(lens): use one tab bar for traces, findings and investigations
* feat(lens): place page actions on the lens tab row
* feat(lens): drop the nested tabs and edit investigations in place
* feat(lens): show investigations as a table with run now and edit
* feat(lens): name each findings row for screen readers
* test(lens): open saved investigation links on findings
* test(lens): reach findings and investigations from the top tabs
* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan
* fix(lens): keep run now since-last-run windows even with an agent override
* test(lens): cover that manual runs never skip scheduled traces
* test(lens): cover run now windows with agent and lookback overrides
* fix(lens): group findings without Map.groupBy and expose sampled runs
* fix(lens): open older findings and review every merged copy from one row
* fix(lens): hide edit and run now from read-only viewers
* chore(lens): drop restating comments from the findings table
* chore(lens): drop restating comments from the step feed
* chore(lens): drop restating comments from the paused banner
* chore(lens): drop restating comments from header actions
* chore(lens): drop restating comments from run now
* test(lens): cover merged findings and read-only investigation rows
* fix(lens): record a model step even when the response has no usage
* test(lens): cover model steps with and without reported usage
* fix(lens): keep a merged finding open when one of its updates fails
* refactor(lens): accept update results from the investigations view
* refactor(lens): accept update results in investigation actions
* test(lens): cover retrying a merged finding after a failed update
* feat(lens): add dot field layout and live status helpers for the traces timeline
* test(lens): cover dot field layout, agent colors and live status
* feat(lens): draw the traces timeline as a live dot field
* feat(lens): add the sweep animation for the live traces timeline
* feat(lens): put the lens tabs in a compact header and fill the screen with traces
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): reorganize Lens dashboard components
* refactor(ui): align Lens forms with react-hook-form
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep Lens submit errors out of form validity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): reduce Lens lint budget usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): use a query key factory for Lens queries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): simplify trace inspection in the gateway drawer
* feat(lens): add a conversation view for full traces
* test(lens): keep normalized message fixtures type safe
* refactor(lens): make conversation view read like a chat
* refactor(lens): use a quiet trace view menu
* refactor(lens): make trace view tabs explicit
* refactor(lens): restore compact trace view switch
* fix(lens): preserve complete conversation history and tool types
* feat(lens): add trace full-screen and close controls
* fix(lens): show forwarded answers and agent errors once
* fix(lens): reset full screen when closing a trace
* fix(roi): estimate linked authors and clarify model selection
* fix: preserve trace errors and ROI results across partial failures
* fix(roi): correct pagination variable typing
* fix(ui): place loaded root failures in conversation order
* fix(roi): read estimator recommendations from model catalog
* revert: remove catalog-driven ROI recommendations
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.
The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(dashboard): migrate to zod 4 and openai 6
Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(dashboard): prettier format zod 4 schema files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(dashboard): restore system one missing state message under zod 4
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* feat(lens): add preset watch-for checks for common agent failures
* feat(lens): add keyboard-driven watch-for picker with lens dot animation
* feat(lens): use the watch-for picker in investigation setup
* feat(lens): show preset checks by name in the criteria tab
* test(lens): cover saving and editing watch-for presets
* feat(lens): shorten watch-for summaries and start with three presets on
* feat(lens): lay out watch-for presets as toggle tiles with a clear add-your-own button
* feat(lens): open a custom check from the watch-for picker
* test(lens): cover watch-for tiles and the add-your-own button
* fix(lens): draw the selected tile border inside the tile so the dialog edge cannot clip it
* refactor(lens): move analysis prompts into markdown files
* feat(lens): ask the investigator for a scoped agent fix brief with two options
* test(lens): cover the agent fix brief through investigation and merges
* chore(ui): regenerate api types for the lens fix brief
* feat(ui): build copyable lens fix prompts
* feat(ui): show the lens fix brief with copy buttons for claude code and codex
* test(ui): cover copying a lens fix option
* refactor(lens): replace the fix options with a plain issue brief
* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes
* test(lens): cover the issue brief through investigation and merges
* chore(ui): regenerate api types for the lens issue brief
* refactor(ui): drop the lens fix prompt builders
* feat(ui): add a lens issue brief panel
* feat(ui): show the lens issue brief in the finding drawer
* test(ui): cover the lens issue brief and the legacy fallback
* feat(ui): render a lens issue brief as a markdown document
* test(ui): pin the lens issue brief markdown layout
* feat(ui): show the issue brief as a copyable file with claude code and codex buttons
* feat(ui): pass the finding title into the issue brief
* test(ui): cover copying the issue brief for claude code and codex
* feat(ui): bold the input and expected labels in lens test cases
* test(ui): pin the bold test case labels in the issue brief
* feat(ui): render the issue brief as formatted markdown
* test(ui): cover the rendered issue brief sections and raw markdown copy
* feat: add Bespoke Nimble gateway and OSS classifier support
* feat: accept Ollama's nimble model name for the Bespoke provider
* test: exempt the POST-only bespoke decisions route from the all-methods check
test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(proxy): extract shared spend log read policy
* test(proxy): use named bindings for spend scope regression
* test(proxy): reuse existing spend log query harness
* test(proxy): cover spend log permission lookup adoption
* chore(proxy): relocate existing spend query baseline
* refactor(proxy): make scope query returns explicit
* refactor(proxy): inject deferred log permission lookup
* test(proxy): cover teamless management compatibility lookup
* refactor(proxy): compose user and team log grants
* refactor(proxy): share generic authorization composition
* refactor(proxy): compose trace read permissions
* refactor(proxy): centralize spend and trace authorization
* refactor(proxy): strengthen spend and trace scope types
* refactor(proxy): flatten log read scope into owned logs
Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.
A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): run spend scope tests through one SQLite emulator
Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.
load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(proxy): unify log and trace ownership permissions
* test(tracing): align fixtures with ownership read scopes
* refactor(tracing): align query scopes with row ownership
* refactor(spend): make ownership SQL predicates explicit
* test(spend): validate ownership SQL against PostgreSQL
* docs(traces): drop key-row visibility from query help guide
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): reach the empty-memberships branch in team lookup test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): compose logs tabs directly in the route
* feat(ui): share composable dashboard page layouts
* refactor(ui): compose page header and logs toolbar from parts
PageHeader drops its icon/title/subtitle/primaryAction/tabs/utilities props and the
leadingControls render prop in favor of PageHeaderTitle, PageHeaderDescription and
PageHeaderControls that each wrap one element and forward native props.
LogsTableToolbar's 15 props collapse into one LogsTimeRange value plus composable
LogsToolbar, LogsTimeRangePicker and LogsToolbarSwitch parts assembled in the panel.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): express DataTable layout classes as cva variants
Replaces the hand-rolled class-pair constants with boolean cva variants,
which also brings DataTable back under the complexity budget.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): compute overall investigation progress, speed and time left
* test(lens): cover overall progress, speed and time left estimates
* feat(lens): show investigation progress as one staged bar with time left
* feat(lens): track per-stage counts and durations for the progress readout
* test(lens): cover stage durations and short time left labels
* feat(lens): restyle investigation progress as a terminal-style readout
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* feat(traces): add Framework column to otel_traces
* feat(traces): pass span events to normalizers and add framework field
* feat(traces): add Claude Code and Agent SDK span normalizer
* feat(traces): decode events before normalizing and apply tool span names
* feat(traces): list distinct frameworks per trace
* feat(traces): return span framework in trace spans query
* test(traces): add scrubbed Claude Agent SDK OTLP fixtures
* test(traces): cover Claude Agent SDK normalization from real exports
* test(traces): assert trace list frameworks stay scoped per trace
* feat(tracing): validate framework in native normalized spans
* feat(tracing): add framework to Span and frameworks to TraceSummary
* feat(tracing): store normalized framework on span rows
* feat(tracing): surface span framework and trace frameworks
* test(tracing): cover framework aggregation in trace summaries
* test(tracing): decode Claude Agent SDK rows with framework and tool args
* chore(ui): regenerate API types for trace frameworks
* feat(ui): add trace framework registry for Claude Agent SDK and Claude Code
* feat(ui): show SDK logo and label in the runs list Agent column
* feat(ui): show SDK logo and label in the run header
* test(ui): cover SDK label and logo in the runs list
* test(ui): cover SDK label and logo in the run header
* feat(tracing): show the agent's final answer as claude agent span output
* feat(tracing): name claude code agents after their otel service
* test(tracing): cover claude code agent naming from the service
* fix(tracing): mark the span row framework field read-only
* test(tracing): scrub host os details from the claude sdk fixture
* test(tracing): scrub host os details from the detailed claude sdk fixture
* fix(ui): hide the decorative sdk logo from screen readers
* feat(ui): show the agent name with the sdk logo in the runs list
* feat(ui): show the agent name with the sdk logo in the run header
* test(ui): cover agent names beside the sdk logo in the runs list
* test(ui): cover the agent name in the run header